Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
56,348 characters · 7 sections · 28 citation commands
Exact Testing of Many Moment Inequalities Against Multiple Violations
\doublespacing
As discussed by chernozhukov2019inference, henceforth CCK, the moments inequalities framework has developed into a powerful tool for inference on causal and structural parameters in partially identified models. In such models, the parameters of interest may be restricted to a subset of the parameter space defined by a collection of moment inequalities. The simultaneous testing of these moment inequalities provides inference about the true underlying parameter values. CCK provide an excellent review of the literature with detailed motivating examples.
They point out that many economic models give rise to problems where the number of moment inequalities $p$ may be much larger than the number of observations $n$. While there exists a large literature on testing moment inequalities, traditional methods are not well equipped for dealing with many moment inequalities.\footnote{For work on unconditional moment inequalities see, e.g., canay2010inference, andrews2012inference, andrews2009validity, chernozhukov2007estimation, rosen2008confidence, romano2008inference. andrews2013inference notice that conditional moment inequalities can be viewed as an infinite number of unconditional moment inequalities. Contributions to conditional moment inequalities are found in chernozhukov2013intersection, lee2013testing, lee2018testing, armstrong2014weighted, armstrong2015asymptotically, armstrong2016multiscale.} In order to test the moment inequalities against the alternative where at least one of them is violated, CCK use the maximum of $p$ $t$ values as a test statistic. They find critical values by using asymptotic theory and bootstrap methods. Additionally, a first-stage inequality selection step is included to improve the power of their tests. allen2018testing suggests a refinement of the selection step. bugni2016inference consider a generalization of the same problem and use Lasso for the first-stage selection.
Our contribution to the framework of testing many moment inequalities is threefold. First, we propose two novel test statistics. Notice that the maximum of $p$ $t$ values is invariant to the size of the second largest $t$ value. If the alternative hypothesis allows for no more than one violation, inference based on the maximum may be powerful, but it would discard power against alternatives where multiple moment inequalities are violated.
In order to retain this power, we propose an enhanced maximum statistic, where we maximize over an artificially expanded set of $t$ values. This expanded set of $t$ values is constructed by adding an extra set of $t$ values to the set of $p$ $t$ values. Such extra $t$ values are created by taking a non-negatively weighted combination of sample means divided by the square-root of its combined sample variance. This results in a test statistic that is also large if such a weighted combination is large, and not just if any of the original $t$ values is large. Interestingly, this expanded set of $t$ values is still a set of $t$ values (though correlated by construction). Therefore, the results of CCK can also be applied to this new test statistic, if the set of extra $t$ values is finite.
In addition, we consider adding all non-negative combinations, so that the set of extra $t$ values is infinite. This results in a test statistic that can be viewed as a one-sided version of Hotelling's $T^2$ statistic. An interesting feature of our statistic is that, unlike the Hotelling's $T^2$ statistic, it can also be used in high dimensional settings. This is because its existence condition is weaker than non-singularity of the sample covariance matrix koning2019directing. This condition is related to the positive eigenvalue condition meinshausen2013sign, slawski2013non or copositivity of the sample covariance matrix. Like the positive eigenvalue condition, it may be satisfied if many elements of the sample covariance matrix are positive. However, it is also satisfied if many of the sample means are negative. In a simulation study, we find substantial increases in power compared to the statistic proposed by CCK, against alternatives where multiple moment inequalities are violated. However, a limitation of the test seems to be that it loses power against alternatives where the number of not satisfied moment inequalities is large compared to $n$ and the underlying covariance matrix exhibits weak correlations, as this interferes with the existence condition.
Secondly, we propose using randomization tests in the moment inequalities framework, by imposing a symmetry assumption on the errors. Randomization tests originate with fisher1935design and have been widely used in the literature.\footnote{See, for example, lehmann2006testing, maritz1995distribution, romano1990behavior, bekker2002exact and bekker2008symmetry} An advantage of randomization tests is that they are exact if all moment inequalities are satisfied. In addition, we provide a sufficient condition to guarantee our tests control size for sufficiently large samples. This result also holds for the statistic that is the maximum over an infinite set of $t$ values.
In our simulation experiments we find that the randomization tests perform similarly or slightly better compared to the empirical bootstrap proposed by CCK in large samples, and better in small samples. While randomization tests work well in small samples, they do require the validity of a symmetry assumption. However, in the simulation experiments where the symmetry assumption does not hold, we find, to our surprise, that the randomization test has better control of size than the empirical bootstrap procedure proposed by CCK, even if the sample is large.
Finally, we describe a first stage selection step in order to improve power by eliminating non-binding moment inequalities. The approach is based on symmetry and it does not depend on a specific selection method. Simulation results show that the tests with selection control size, and have increased power in the presence of non-binding moment inequalities compared to tests without selection.
Let observations in the $n\times p$ matrix $\bm X$ satisfy
where $\bm \iota_n$ is an $n$-vector of ones, $\bm \mu$ is a $p$-vector of parameters and $\bm E$ is a random matrix that satisfies a symmetry assumption. Specifically, we assume that any row of $\bm E$ can be multiplied by $-1$ without affecting the distribution of $\bm E$. This assumption is satisfied if the rows of $\bm E$ are independently symmetrically distributed about zero. Moments need not exist and the rows of $\bm E$ need not be identically distributed. We describe exact inference based on the finite sample.
Following CCK, we are interested in testing the hypothesis H$_0$:\ $\bm \mu \leq \boldsymbol{0}$ against the alternative H$_1$:\ $\bm \mu \not\leq \boldsymbol{0}$. That is, we want to test whether all elements of $\bm \mu$ are non-positive, against the alternative that at least one is positive. We will refer to the null hypothesis as the moment inequalities. The $j$th moment inequality is said to be violated if $\mu_j > 0$, it is binding if $\mu_j = 0$, and it is strictly satisfied if $\mu_j < 0$.
To define the test statistics, let $\bm I_n = (\bm e_1, \dots, \bm e_n)$ be the $n \times n$ identity matrix, and define $\bm P_{\bm \iota_n} = \bm \iota_n\bm \iota_n'/\bm \iota_n'\bm \iota_n$. Define $\widehat{\bm \mu} = \bm X'\bm \iota_n/n$ with elements $\widehat{\mu}_j$, and $\widehat{\bm \varSigma} = \bm X'(\bm I_n - \bm P_{\bm \iota_n})\bm X/n$ with diagonal elements $\widehat{\sigma}_j^2$. We assume $\widehat{\bm \mu}\nleq\boldsymbol{0}$, otherwise we would not reject H$_0$, and consider test statistics of the form
where $\mathcal{U}\subseteq\mathbb{R}^p$ and Condition (ref) holds koning2019directing.
The test statistic has been used before for particular choices of $\mathcal{U}$. Here we propose new choices as well.
In order to test H$_0$ against H$_1$, CCK use the maximum $t$ value, $t_j =n^{1/2} \widehat{\mu}_j/\widehat{\sigma}_j$, for $j =1, \dots, p$. We will denote this statistic by $ t_{\text{max}}=t_{\text{max}}(\bm X) = \max_{1 \leq j \leq p} t_j = T_{\mathcal{C}^p} $, where $\mathcal{C}^p=\{\bm e_1, \dots, \bm e_p\}$. By ignoring the sizes of all but the largest $t$ value, the $t_{\text{max}}$ statistic may perform well against alternatives with one or few violations, but lack power when testing against alternatives with multiple violations. To shift power towards multiple violations, one might expand $\mathcal{U}$ by adding vectors from the non-negative orthant, $\bm a_i\in\mathbb{R}_+^p$, $i=1,\ldots,m$. Let $\mathcal{U}=\mathcal{C}^p\cup\{\bm a_1,\ldots,\bm a_m\}$, then $T_\mathcal{U}$ may outperform $t_{\text{max}}$ in terms of power under alternatives with multiple violations.
Alternatively, we consider infinite sets $\mathcal{U}$. In particular, we use $\mathcal{U}=\mathbb{R}_+^p$ and define $ T_+=T_{\mathbb{R}_+^p}. $\footnote{To handle cases with multiple violations, CCK mention the sum statistic $ \sum_{i = 1}^{p} (\sqrt{n}\max\{\widehat{\mu}_j/\widehat{\sigma}_j, 0\})^2$ and the quasi likelihood ratio test statistic $\min_{\bm \mu \leq \boldsymbol{0}} n (\widehat{\bm \mu} - \bm \mu)'\widehat{\bm \varSigma}^{-1}(\widehat{\bm \mu} - \bm \mu)$ as posibilities. Only if $\widehat\bm \varSigma$ is diagonal, these statistics fit within the $T_\mathcal{U}$ framework. In that case the statistics are equal to $T_+$. However, the sum statistic ignores the covariances, and the quasi likelihood-ratio statistic requires $\widehat{\bm \varSigma}$ to be invertible.} The $T_+$ statistic can be viewed as a one-sided version of the square-root of Hotelling's $T^2$-statistic, since $T_{\mathbb{R}^p}^2=T^2 = n\widehat{\bm \mu}'\widehat{\bm \varSigma}^{-1}\widehat{\bm \mu}$, which requires $\widehat\bm \varSigma$ to be invertible. For the $T_+$ statistic $\widehat\bm \varSigma$ need not be invertible. Instead, a weaker condition suffices, which is a special case of Condition (ref).
This condition is weaker than (strict) copositivity of $\widehat{\bm \varSigma}$, which requires that $\bm \lambda'\widehat{\bm \varSigma}\bm \lambda > 0$ for all $\bm \lambda \geq \boldsymbol{0}$, $\bm \lambda \neq \boldsymbol{0}$ (see e.g. hiriart2010variational), or the equivalent positive eigenvalue condition formulated by meinshausen2013sign.\footnote{The positive eigenvalue condition $\min_{{\bm \lambda \geq \boldsymbol{0},\ \bm \lambda \neq \boldsymbol{0}}} {\bm \lambda'\widehat{\bm \varSigma}\bm \lambda}/{\|\bm \lambda\|_1^2} > 0$ is equivalent to $\bm \lambda'\widehat{\bm \varSigma}\bm \lambda > 0$, for all $\bm \lambda \geq \boldsymbol{0}$, $\bm \lambda \neq \boldsymbol{0}$.}
An example where copositivity holds is if all elements of $\widehat{\bm \varSigma}$ are strictly positive (i.e. $\widehat{\bm \varSigma}$ is positive). meinshausen2013sign also provides an example where $\widehat{\bm \varSigma}$ is allowed to contain some negative entries. In particular, if $\mathcal{A}^c$ contains the indices of the the largest positive principal submatrix of $\widehat{\bm \varSigma}$ and $\mathcal{A}$ is its complement, then the condition only has to hold on the principal submatrix of $\widehat{\bm \varSigma}$ with indices in $\mathcal{A}$. If the number of elements in $\mathcal{A}$ is small compared to $n$, this need not be very restrictive.
The difference between Condition (ref) and copositivity lies in the role played by $\widehat{\bm \mu}$. Condition (ref) is equivalent to copositivity if $\widehat{\bm \mu} > \boldsymbol{0}$. At the other extreme, if $\widehat{\bm \mu} \leq \boldsymbol{0}$ the set $\{\bm \lambda\ |\ \bm \lambda \geq\boldsymbol{0},\ \lambda\neq\boldsymbol{0},\ \bm \mu'\bm \lambda\geq 0\}$ is empty, so that Condition (ref) holds, whether or not copositivity holds. If the elements of $\widehat{\bm \mu}$ have comparable positive or negative sizes, then Condition (ref) tends to become less restrictive the more negative elements there are. So, intuitively, even if $\widehat{\bm \varSigma}$ is not copositive, Condition (ref) may still hold, in particular if many moment inequalities are strictly satisfied. This reasoning is confirmed in the simulation experiments, where the test based on $T_+$ performs well even when $n = 30$ and $p = 1000$, and only 10 moment inequalities are violated.
To compute $T_+$ we use a recent result by koning2019directing, who shows
This problem can be solved by the non-negative least squares algorithm of lawson1995solving.\footnote{This algorithm is implemented in the R package `nnls' and the MATLAB function `lsqnonneg'.} An interesting additional feature is that the resulting maximizing weight vector $\widehat{\bm \lambda}$ may contain zeros, resulting in a selection of moment inequalities that are `suspected' to violate the null.
The randomization tests that we consider are based on the symmetry assumption. For completeness, we include this section explain how the assumption leads to exact randomization tests, and show how they can be applied in practice. For other descriptions of randomization tests, see e.g. lehmann2006testing, maritz1995distribution and bekker2008symmetry.
Let $\mathcal{R} = \{\bm R_1, \bm R_2, \dots, \bm R_N\}$ be the set of $N = 2^n$ diagonal $n \times n$ matrices with diagonal elements in $\{-1, 1\}$, where $\bm R_1$ denotes the identity matrix. This set constitutes a finite reflection group under matrix multiplication. The group $\mathcal{R}$ determines a partitioning $\mathcal{E}$ of $\mathbb{R}^{n \times p}$ into equivalence classes (“orbits") denoted by $\mathcal{R}_{\bm E^*} = \{\bm E^*, \bm R_2\bm E^*, \dots, \bm R_N\bm E^*\}$, where $\bm E^* \in \mathbb{R}^{n \times p}$ acts as a representative of the class. For simplicity, we only consider orbits $\mathcal{R}_{\bm E^*} \in \mathcal{E}$, for which the cardinality satisfies $|\mathcal{R}_{\bm E^*}| = |\mathcal{R}| = N$. This is equivalent to assuming that we only consider the subset of orbits $\mathcal{E}^0 \subset \mathcal{E}$ where the elements of $\mathcal{R}_{\bm E^*} \in \mathcal{E}^0$ have no rows that equal zero.
The basic assumption that permits the construction of randomization tests is a distributional invariance assumption on the error term under group transformations. In particular, we assume that $\bm R\bm E$ and $\bm E$ have the same distribution, for all $\bm R \in \mathcal{R}$. Equivalently, we assume the conditional distribution of $\bm E$, given an orbit $\bm E\in\mathcal{R}_{\bm E^*}$, is uniform for all orbits $\mathcal{R}_{\bm E^*} \in \mathcal{E}^0$.
Let $g: \mathcal{R}_{\bm E^*} \to \mathbb{R}\cup\{\infty\}$ be a mapping. Notice that if $g$ is injective on all orbits $\mathcal{R}_{\bm E^*} \in \mathcal{E}^0$, then the symmetry assumption implies that the conditional distribution of $g(\bm E)$, given an orbit $\mathcal{R}_{\bm E^*} \in \mathcal{E}^0$, is uniform just as well. Furthermore, the function $p(\bm E) = |\{\bm R \in \mathcal{R}\ |\ g(\bm R\bm E) \geq g(\bm E)\}| / N$ has a conditional distribution, given an orbit, that is uniform over $\{1/N, 2/N, \dots, 1\}$. As this holds for all $\mathcal{R}_{\bm E^*} \in \mathcal{E}^0$, the function $p(\bm E)$ has an unconditional uniform distribution over $\{1/N, 2/N, \dots, 1\}$. This leads to the following results.
Following bekker2008symmetry, we will refer to $g$ as an inferential function. As the cardinality of $\mathcal{R}$ grows exponentially in $n$, the computation of $p(\bm E)$ is intractable even in small samples. Therefore, we provide the following result to allow for sampling from $\mathcal{R}$.\footnote{bekker2008symmetry also describe sampling with replacement.}
To test the moment inequalities of model \hyperref[{model}]{\tagform@{\ref*{model}}} , we use randomization tests based on the $T_{\mathcal{U}}$ statistic with $\mathcal{U} \subseteq \mathbb{R}_+^p$, such as $t_{\text{max}}$ and $T_+$. We will use the convention that $T_{\mathcal{U}} = 0$ if $\widehat{\bm \mu} \leq \boldsymbol{0}$ and $T_{\mathcal{U}} = \infty$ if Condition (ref) does not hold. Let the test statistic be generically denoted as $T$. The tests are exact if the moment inequalities are binding. In that case $\bm X = \bm E$ and $T$ may be seen as an inferential function $g(\bm E) = T(\bm E)$. So, Proposition (ref) can be used based on $p_M(\bm E) = p_M(\bm X)=|\{\bm R \in \mathcal{R}^M\ |\ T(\bm R(\bm X))\geq T(\bm X)| / M$. In particular, we use the following testing procedure. \\
\noindentAlgorithm 1 (Symmetry Randomization Test).
If $\bm \mu=\boldsymbol{0}$ and $T$ is injective and $\alpha\in\{1/N,2/N,\ldots,1\}$, this test is exact (note that point accumulation could occur at 0 or $\infty$).
In case some moment inequalities are assumed to be strictly satisfied, the test is no longer exact, but we observe in all our simulations that the size is controlled if the symmetry assumption holds. To prove this formally is another matter. If, under the null hypothesis $\bm \mu\leq\boldsymbol{0}$, it can be verified that $T(\bm R\bm X)\geq T(\bm \iota_n\bm \mu'+\bm R\bm E)$, then the test is conservative.\footnote{Notice that for the numerator of $T_+(\bm R\bm X)$ it holds that $\bm \iota_n'\bm R\bm X\bm \lambda=\bm \iota_n'\bm R\bm \iota_n\bm \mu'\bm \lambda+\bm \iota_n'\bm R\bm E\bm \lambda\geq\bm \iota_n'(\bm \iota_n\bm \mu'+\bm R\bm E)\bm \lambda$.} In that case we find, that $T(\bm X) = T(\bm \iota_n\bm \mu' + \bm E)$ is an inferential function, $g(\bm E) = T(\bm \iota\bm \mu' + \bm E)$, so that
which implies $\mbox{{P}}(p^T(\bm X) \leq \alpha) \leq \mbox{{P}}(p(\bm E) \leq \alpha) \leq \alpha$. Unfortunately, we have not been able to formally prove whether or not the condition holds for our test statistics. However, a condition that the test is conservative for sufficiently large samples is easy.
To verify $T(\bm R\bm X) \geq T(\bm \iota_n\bm \mu' + \bm R\bm E)$, first notice that this inequality holds trivially if $\bm R = \bm I_n$ or $\bm \mu'\bm \lambda = 0$. So, we can restrict ourselves to the case that $\bm R \neq \bm I_n$ and $\bm \mu'\bm \lambda < 0$, as $\bm \mu'\bm \lambda \leq \boldsymbol{0}$ for $\bm \lambda \geq \boldsymbol{0}$. Furthermore, notice that
is a strict monotone transformation of $T$, so that we can instead look at $T^*$. Consider the function of $\bm \lambda \in \mathcal{U} \subseteq \mathbb{R}_+^p$ that is maximized at $T^*(\bm R\bm X)$,
as $\bm \iota_n'\bm R\bm \iota_n < n$. Consequently, due to Condition (ref), as $\bm E$ and $\bm R\bm E$ follow the same distribution,
Furthermore, $T^*(\bm \iota_n\bm \mu' + \bm R\bm E)$ maximizes
so that
Let $\bm \lambda^*$ be the maximizer of $T^*(\bm \iota_n\bm \mu' + \bm R\bm E)$, then
As a result, the condition $T^*(\bm R\bm X) \geq T^*(\bm \iota_n\bm \mu' + \bm R\bm E)$ holds, and therefore $T(\bm R\bm X) \geq T(\bm \iota_n\bm \mu' + \bm R\bm E)$, so the test is conservative if $n$ is sufficiently large.
As the tests are found to be conservative in case some moment equalities are strictly satisfied, there may be room for power improvements. CCK propose an inequality selection step aiming at removal of such strict moment inequalities $j$. In particular, they remove moment inequalities $j$ for which $t_j < c$, where $c$ is some chosen cut-off value. They propose several techniques to select this cut-off value such that the asymptotic testing procedures remain applicable, up to a small correction of the significance level. A similar way to select the cut-off value using Lasso is proposed by bugni2016inference.
Our approach for inference with pre-selection is similar to the approach of the previous section where there was no pre-selection. We use symmetry based inference. Given any inequality selection rule we describe tests that are exact or conservative if the moment inequalities are binding. That is to say, if $\bm \mu=\boldsymbol{0}$, then $T(\bm X)=T(\bm E)$ and $p^T(\bm X)=|\{\bm R \in \mathcal{R}\ |\ T(\bm R\bm X) \geq T(\bm X)\}|\leq\alpha$.
Let $\mathcal{J}(\bm X)\subset\{1,\ldots, p\}$ be an index selection subset, and let $\bm X_{\mathcal{J}(\bm X)}$ be the submatrix of $\bm X$ consisting of columns with indexes in $\mathcal{J}(\bm X)$. As the selection depends on $\bm X$, it is not guaranteed that $p^T(\bm X_{\mathcal{J}(\bm X)})=|\{\bm R \in \mathcal{R}\ |\ T(\bm R\bm X_{\mathcal{J}(\bm X)}) \geq T(\bm X_{\mathcal{J}(\bm X)})\}|\leq\alpha$. However, if $\bm \mu=\boldsymbol{0}$, then $T(\bm X_{\mathcal{J}(\bm X)})=T(\bm E_{\mathcal{J}(\bm E)})=g(\bm E)$ is an inferential function and
which implies $\mbox{\emph{P}}(p_{\text{sel}}^T(\bm X_{\mathcal{J}(\bm X)})\leq \alpha)=\mbox{\emph{P}}(p(\bm E) \leq \alpha) \leq \alpha$, for $\alpha\in[0,1]$ and, if $g$ is injective, then $\mbox{\emph{P}}(p_M(\bm E) \leq \alpha) = \alpha$ for $\alpha \in\{1/M, 2/M, \dots, 1\}$, as in Proposition (ref). Therefore, we use the following testing procedure. \\ \noindentAlgorithm 2 (Symmetry randomization test with inequality selection).
If $\bm \mu\leq\boldsymbol{0}$ and $\bm \mu\neq\boldsymbol{0}$, the test is conservative if $T(\bm R\bm X_{\mathcal{J}(\bm R\bm X)})\geq T\left((\bm \iota\bm \mu'+\bm R\bm E)_{\mathcal{J}(\bm \iota\bm \mu'+\bm R\bm E)}\right)$. The power of the test varies with the value of $\bm \mu$ and the selection method. We follow CCK and use $\mathcal{J}(\bm X)=\left\{j\in\{1,\ldots,p\}\ |\ t_j>c\right\}$, where the selection constant is chosen using their empirical bootstrap procedure. The simulations show that the size of the tests is controlled when $\bm \mu \leq \boldsymbol{0}$.
In this section, we present Monte Carlo simulation results. We use setups based on the simulation experiments presented by CCK and bugni2016inference. The tests described in this paper are compared to the Empirical Bootstrap (EB) tests described in CCK.\footnote{We only compare to the EB tests described in CCK, as they find that the EB tests perform similarly to the Multiplier Bootstrap tests and better than the self-normalized tests.} All experiments were implemented in R and code for the tests will be made available at \url{https://github.com/nickwkoning/}. \\
\noindentData generation \\ The data is created as follows: we generate $n \times p$ matrix $\bm X = \bm \iota_n\bm \mu' + \bm E\bm A$, where $\bm \iota_n$ is an $n$-vector of ones, $n\times p$ matrix $\bm E$ has i.i.d. elements drawn from a distribution $F$, and $\bm A$ is defined such that $\bm \varSigma = \bm A'\bm A$, where $\bm \varSigma$ has elements $\sigma_{ij} = \rho^{|i-j|}$.
The parameters we use are $n \in \{30, 400\}$, $p \in \{200, 500, 1000\}$ and $\rho \in \{0, 0.5, 0.9\}$. For the vector $\bm \mu$, we consider four different designs. The signs of the elements of $\bm \mu$ are chosen to represent four general cases: all moment inequalities are binding (Design 1), most moment inequalities are strictly satisfied and some are binding (Design 2), all moment inequalities are violated (Design 3), and some moment inequalities are violated while some are strictly satisfied (Design 4). The values of $\bm \mu$ for each combination of $n$ and $p$ were selected to ensure that the rejection probabilities are bounded away from 1.
For the errors, we use a symmetric and two asymmetric distributions: one left-skewed and one right-skewed. As symmetric distribution, we use $t(4)/\sqrt{2}$, where $t(4)$ is the Student's $t$ distribution with 4 degrees of freedom and we divide by $\sqrt{2}$ so that the variance is 1. As asymmetric distributions we use the skew-normal distribution with mean 0, variance equal to 1 and two configurations for the skewness: $\gamma = -0.667$ and $\gamma = 0.667$. A density plot comparing these skewed distributions to the standard normal distribution is provided in Figure (ref). For the sake of brevity, the asymmetric error distributions were only considered for Design 1.
This setup is similar to the setup used by an earlier working paper of CCK.\footnote{This earlier version can be found at \url{https://arxiv.org/abs/1312.7614v4}.} The differences are that we do not consider equicorrelated data or uniformly distributed errors, but instead consider small samples ($n = 30$) and asymmetric error distributions. In addition, the positive values of $\bm \mu$ are substantially decreased to account for the higher power of tests based on the $T_+$ statistic. \\
\noindentTests \\ We use symmetry based randomization (SR) tests based on the test statistics $t_{\text{max}}$ and $T_+$, with and without pre-selection. In addition, we also consider the $t_{\text{max}}^{\bm \iota} = T_{\mathcal{C}_p \cup \{\bm \iota_p\}}$ statistic, where the maximum is taken over just a single extra vector compared to $t_{\text{max}}$. As a comparison, we include the EB tests described by CCK with and without pre-selection for both the $t_{\text{max}}$ and $t_{\text{max}}^{\bm \iota}$ statistics.
For each setting, 1000 realizations of $\bm X$ were generated and the proportion of rejections was recorded. For the EB tests 1000 bootstrap samples were used and for the SR tests we used 1000 reflection samples. The significance level was fixed at $\alpha = 0.05$. The selection constant for the tests that include pre-selection is chosen using the EB procedure with $\beta = 0.001$ and 1000 resamples as described in CCK. For the small sample experiments with $n = 30$, we do not use pre-selection as the selection constant depends on the asymptotic properties of the EB test.
In terms of computation time for the largest settings, a single test without pre-selection using the $T_+$ statistic a on single core of a standard 13-inch 2017 MacBook Pro takes approximately 80-270 seconds, depending on the design. For the $t_{\text{max}}$ and $t_{\text{max}}^{\bm \iota}$ statistics this is approximately 5-7 seconds.
The results of the simulation experiments can be found in Tables (ref) to (ref), and are discussed separately for each design. \\
\noindentDesign 1: $\bm \mu = \boldsymbol{0}$. \\ In Design 1, all elements of $\bm \mu$ are equal to zero. Therefore, the rejection probability for the tests should be at most $\alpha = 0.05$. The results for Design 1 are displayed in Table (ref) for the symmetric error distribution, and Table (ref) for the asymmetric error distributions.
In Table (ref), for the case that the number of observations is large ($n = 400$), we see that the EB tests for both $t_{\text{max}}$ and $t_{\text{max}}^{\bm \iota}$ reject with probability close to 0.05, which is in line with Remark (ref). For the symmetry randomization (SR) tests, Proposition (ref) states that all rejection probabilities should be at most 0.05, which is confirmed in the simulations. In addition, if the test statistic is injective on all orbits, then the rejection probability is equal to $\alpha$. One exception found to this in the simulations is the configuration $p = 1000$ and $\rho = 0$, for the statistic $T_+$, where the proportion of rejections is 0. Further inspection shows that for this configuration without pre-selection, the test statistics $T_+(\bm X)$ and critical values, defined as the $1 - \alpha$ quantile of the set $\{T_+(\bm R\bm X)\ |\ \bm R \in \mathcal{R}_M\}$, are infinite, as Condition (ref) is not satisfied. Therefore, the function $T_+$ is not injective on the orbit and the rejection probability can be strictly smaller than $\alpha$, according Proposition (ref). A similar problem of non-injectivity occurs for the case with pre-selection. In line with the relation between Condition (ref) and copositivity, the issue diminishes if $\rho$ is increased.
If the number of observations is small ($n = 30$), then for the $t_{\text{max}}$ and $t_{\text{max}}^{\bm \iota}$ statistics, the rejection rates for the SR tests remain approximately $\alpha$. In contrast, the EB tests over-reject. For the SR tests based on the $T_+$ statistic, the rejection rate is frequently zero due to the non-injectivity phenomenon described above.
Table (ref) shows the rejection rate for Design 1 when the error distributions are left-skewed ($\gamma = -0.667$) and right-skewed ($\gamma = 0.667$). As the symmetry assumption is violated, it is not surprising that the SR tests have a rejection rate different from $\alpha$. In particular, we find that the rejection rate for the left-skewed error distribution is larger than $\alpha$, while the rejection rate for the right-skewed error distribution is smaller than $\alpha$.
Surprisingly, even though the sample size is large ($n = 400$), the EB tests over-reject even more than the SR tests under the left-skewed distribution. This suggests that the EB tests may also benefit from errors being symmetrically distributed in finite samples. Although these simulation results are by no means exhaustive, they suggest that one may sometimes be better off using an SR test than an EB test even if it is known that the true error distribution is not symmetric. \\
\noindentDesign 2: $\bm \mu \leq \boldsymbol{0}$, $\bm \mu \neq \boldsymbol{0}$. \\ In Design 2, for $n = 400$, the first $0.1p$ elements of $\bm \mu$ are equal to 0 and the remaining elements are equal to $-0.8$. For $n = 30$, the first $10$ elements are 0 and the remaining elements are equal to $-5$. As the data for Design 2 is generated under $H_0$, the rejection rates should be smaller than $\alpha$ (up to sampling errors) if no pre-selection is used, and close to $\alpha$ if pre-selection is used. The results for Design 2 can be found in Table (ref).
From Table (ref), it can be observed that the rejection rates of the tests without pre-selection is smaller than $\alpha$. For the SR tests without pre-selection and $n = 400$, this is in line with the conclusions in Section (ref) as $n$ is large. For $n = 30$, the rejection rates are closer to, but still smaller than $\alpha$. In addition, pre-selection lifts the rejection rates close to the nominal size $\alpha$. \\
\noindentDesign 3: $\bm \mu > \boldsymbol{0}$. \\ In Design 3, all elements of $\bm \mu$ are positive. In particular, for $n = 400$ they are equal to 0.01 and for $n = 30$ they are equal to 0.03. As the data is generated under the alternative, the rejection probability should be as large as possible. The results for Design 3 are found in Table (ref).
In Table (ref), it can be seen that for $n = 400$ the power of the EB and SR tests is similar, and inequality selection has no noticeable effect. The $T_+$ statistic generally leads to the highest power. One exception is the case where $p = 1000$ and $\rho = 0$, which is the most challenging case with respect to Condition (ref). The power difference between the $t_{\text{max}}^{\bm \iota}$ and $t_{\text{max}}$ statistics is quite remarkable, as the $t_{\text{max}}^{\bm \iota}$ statistic maximizes over just a single extra vector. However, in Design 4 we will see that the statistic is more sensitive to the alternative than $T_+$. Furthermore, the difference in the power between the statistics is largest under weak correlations, which is in line with Remark (ref).
For $n = 30$, the results for the EB tests should be ignored, as the tests do not control size. Although there are multiple violations, the $T_+$ statistic performs worst. This is expected as Condition (ref) is unlikely to be satisfied here. So the same phenomenon observed under Design 1 occurs, where both the test statistic and critical value are infinite. Therefore, it is recommended that the test statistic and critical value are inspected before the outcome of the test is interpreted. If both are infinite, then the test based on the $T_+$ statistic should not be used. In addition, the $t_{\text{max}}^{\bm \iota}$ statistic generally performs as well as or better than the $t_{\text{max}}$ statistic, especially if $\rho = 0$. \\
\noindentDesign 4: $\bm \mu \not\leq \boldsymbol{0}$, $\bm \mu \not > \boldsymbol{0}$. \\ In Design 4, for $n = 400$, the first $0.1p$ elements of $\bm \mu$ are equal to 0.02 and the remaining elements are equal to $-0.75$. For $n = 30$, the first 10 elements of $\bm \mu$ are equal to 0.3 and the remaining elements are equal to -5. The data is generated under the alternative, so the rejection rates should be as large as possible. The results for Design 4 are presented in Table (ref).
The results in Table (ref) for $n = 400$ show that the tests with pre-selection have substantially higher power than the tests without pre-selection. When considering the tests with pre-selection, tests based on the $T_+$ statistic have much more power than the tests based on the $t_{\text{max}}$ and $t_{\text{max}}^{\bm \iota}$ statistics. The power for the EB and SR tests based on the $t_{\text{max}}$ statistic are similar if pre-selection is used, but the SR tests seem more powerful if pre-selection is not used. In addition, the performance of the $t_{\text{max}}^{\bm \iota}$ and $t_{\text{max}}$ statistics is similar, unlike in Design 3. Further inspection shows that the pre-selection often fails to eliminate a few strictly satisfied moment inequalities, which causes the added weight vector $\bm \iota_p$ to rarely be the maximizing vector.
Even though the EB tests do not control size for $n = 30$, the SR tests seem to be more powerful. Comparing the different statistics for the SR tests, it can be observed that the $T_+$ statistic performs well compared to the $t_{\text{max}}$ statistic if $\rho = 0$, even if $p = 1000$ and $n = 30$. So, the issue regarding Condition (ref) is not observed here. This is in line with the intuitions obtained in Section (ref) which suggest that the condition may still be satisfied if many of the moment inequalities are strictly satisfied. The $t_{\text{max}}^{\bm \iota}$ statistic performs best overall for $n = 30$.
\singlespacing