Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
84,131 characters · 16 sections · 85 citation commands
Simple Adaptive Size-Exact Testing for Full-Vector and Subvector Inference in Moment Inequality Models
\\ \\
{\bf Keywords:} Moment Inequalities, Uniform Inference, Likelihood Ratio, Subvector Inference, Convex Polyhedron, Linear Programming
In the past decade or so, inequality testing has become a mainstream inference method used for models where standard maximum likelihood or method of moments are difficult to use, for reasons including multiple equilibria, incomplete data, or complicated dynamic patterns. In such models, inequalities can often be derived from equilibrium conditions and rational decision making. Inference can then be conducted by inverting tests for these inequalities at each given parameter value.\footnote{An incomplete list of applications that use inequalities as estimation restrictions includes Tamer2003, Uhlig2005, BajariBenkardLevin2007, BGIM2007, CilibertoTamer2009, BMM2011, Holmes2011, BIWY2012, Chetty2012, NevoRosen2012, KawaiWatanabe2013, Eizenberg2014, HuberMellace2015, PPHI2015, MagnolfiRoncoroni2016, Sheng2016, Sullivan2017, He2017, ISS2018, Wollman2018, FackGrenetHe2019, MSZ2019. For recent overview of the literature, see for example HoRosen2017, CanayShaikh2017, and Molinari2020.}
Although conceptually simple, conducting inference via test inversion poses considerable computational challenges to practitioners. This is because in order to get an accurate calculation of the confidence set, one needs to test the inequalities at a set of parameter values that is dense enough in the parameter space. Depending on the application, the number of values needed to be tested can be astronomical and increases exponentially with the dimension of the parameter space. Moreover, existing tests often require simulated critical values that are nontrivial to compute even for a single value of the parameter, let alone repeated for a large number of parameter values.\footnote{Existing tests for general moment inequalities with simulated critical values include CHT2007, RomanoShaikh2008, AndrewsGuggenberger2009, AndrewsSoares2010, Bugni2010, Canay2010, RomanoShaikh2012, and RomanoShaikhWolf2014. See CanayShaikh2017 and Molinari2020 for more references.}
Besides computational challenges, most existing methods for moment inequality models involve tuning parameter sequences that are required to diverge at a certain rate as the sample size increases. The threshold in the generalized moment selection procedures (e.g. Rosen2008 and AndrewsSoares2010) and the subsample size in subsampling-based methods (e.g. CHT2007 and RomanoShaikh2012) are notable examples.\footnote{Arguably, the size of a first stage confidence set or the number of simulation/bootstrap draws are also tuning parameters commonly used to test moment inequalities.} Appropriate choices often depend on data in complicated ways, and an inappropriate choice can threaten the validity of the test.
Clearly, there are two ways to ease the computational burden: one is to make the inequality test easier for each parameter value, and the other is to reduce the number of parameters that need to be tested. We contribute to the literature in both. First, we propose a simple test for general moment inequalities that requires no simulation. It simply uses the $\text{(quasi-)}$ likelihood ratio statistic ($T_n$) and a chi-squared critical value, where the data-dependent degrees of freedom come as a by-product of computing $T_n$. By not requiring simulation, the test saves computation time hundreds-fold compared to tests involving simulated critical values, where a critical value statistic needs to be computed for each simulated sample.
Second, we then specialize to a conditional moment inequality model with nuisance parameters ($\delta$) entering linearly, a common empirical setting, and propose a confidence set for a subvector parameter of interest ($\theta$). The subvector confidence set is based on our full-vector test applied to the nuisance-parameter-eliminated model. By focusing on the parameter of interest, one only needs to consider a grid on the space of $\theta$ which can be much lower dimensional than the space of $(\theta',\delta')'$. Thus, the number of parameter values that need to be tested is drastically reduced. Also, as an added benefit, the subvector test is less conservative and more powerful than the projection of the full vector test, as demonstrated in our second Monte Carlo experiment.
In both contexts, our test is simulation and tuning parameter free. Its critical value is simply the chi-squared critical value with degrees of freedom equal to the rank of the active moment inequalities, where we call a moment inequality {\em active} if it holds with equality at the restricted estimator of the moments.\footnote{Active inequalities are the sample counterpart of binding inequalities, which hold with equality at the population expectation of the moments.} The test is shown to have exact finite sample size in a normal model with known variance and to be uniformly asymptotically valid more generally. Moreover, it automatically adapts to the slackness of the moment inequalities despite the absence of a deliberate moment selection step. In particular, when all but one inequality get increasingly slack, the test asymptotes to one that ignores all the slack inequalities, which coincides with the uniformly most powerful test for the limiting model.
The idea of simple chi-squared critical values for testing inequalities appeared as early as in Bartholomew1961 and Rogers1986 for testing one-sided alternatives against a {\em simple} null, but was only recently proved to be valid for a composite null in MohamadGoemanZwet2020 in a normal model. We move beyond MohamadGoemanZwet2020 in three ways: (a) we allow an intercept in the inequalities defining the null and thus generalize the null space from a cone to a polyhedron. This is important for moment inequality models as the limiting experiment may not be a cone in general when we allow local-to-binding inequalities; (b) we design a simple but novel refinement to make the test size-exact; and (c) we prove the uniform asymptotic validity of the test for moment inequality models.
The idea of eliminating nuisance parameters from linear moment inequalities is first suggested in GuggenbergerHahnKim2008, where they introduce the Fourier-Motzkin elimination to the literature and propose a Wald-type test on the resulting inequalities. Yet two main difficulties hinder the application of this idea: (a) numerical calculation of the Fourier-Motzkin elimination in general is an NP-hard computational problem, and (b) the estimated slopes in front of the nuisance parameters enter the inequalities via a non-differentiable function, and thus undermine the validity of testing procedures applied directly to them. The first difficulty is not present for our test because its special structure only requires us to calculate the {\em rank} of the active inequalities, avoiding the full-blown elimination procedure. The second difficulty is circumvented by considering the conditional distribution of the moments given a vector of instrumental variables.
AndrewsRothPakes2019 has the closest setting with our paper. They propose a test based on the max statistic. In the most basic version, their test uses a conditional critical value from a truncated normal distribution. This basic version involves no simulation or tuning parameter and as a result is easy to compute. However, the basic version has poor power properties that prompt them to recommend a hybrid test. The hybrid test uses a simulated critical value as well as a tuning parameter that determines the size of a first-stage confidence set.
There are a few papers in the literature that propose methods to mitigate the computational challenges described above. KMS2019 cast the problem of finding the bounds of the projection confidence interval of each parameter into a nonlinear nonconvex constrained optimization problem, and provide a novel algorithm to solve this optimization problem more efficiently. Our simple inequality test is complementary to KMS2019's algorithm in that we make testing for each value hundreds-fold easier while their algorithm reduces the number of values that need to be tested. BugniCanayShi2017 propose a profiling method that simplifies computation in the same way as the subvector confidence set proposed in this paper, by reducing the search from the space of the whole parameter vector to that of a low dimensional subvector. The difference is that our subvector test, by taking advantage of the linearity of the model, is much easier to compute than BugniCanayShi2017's test, which applies more generally. CCT2018 propose a quasi-Bayesian method to subvector inference which has similar computational cost as BugniCanayShi2017 when applied to moment inequality models.\footnote{More details about the comparison of the computational aspect of these papers can be found in Section 4.3 of HsiehShiShum2020.} CCT2018 also propose a simple test for a scalar parameter of interest that uses a chi-squared critical value that is valid under an additional assumption.
A couple of other papers aim to reduce the sensitivity of testing to tuning parameters. AndrewsBarwick2012 (AB, hereafter) refines the procedure of AndrewsSoares2010 by computing an optimal moment selection threshold that maximizes a weight average power and a size correction. Using the optimal threshold and the size correction provided in that paper, one no longer needs to choose a tuning parameter. Computationally, it is the same as AndrewsSoares2010 if one has 10 or fewer moment inequalities and can use the tables of optimal tuning and size correction values in the paper. It is much more computationally demanding otherwise. RomanoShaikhWolf2014 (RSW, hereafter) replace the moment selection step of the previous literature with a confidence set for the slackness parameter and employ a Bonferroni correction to take into account the error rate of this confidence set. There is still a tuning parameter, the confidence level of the first step, but this tuning parameter no longer affects the asymptotic size of the test. Computationally, using the same number of critical value simulations, it is slightly more costly than AndrewsSoares2010 due to the first-step confidence set construction. AB and RSW are our points of comparison in our first set of Monte Carlo experiments, where we show that our simple test saves computational cost hundreds-fold, while comparing favorably to AB and RSW in terms of size and power.
The remainder of this paper proceeds as follows. Section 2 covers full-vector inference in moment inequality models. Section 3 covers the extension to subvector inference in conditional moment inequality models with nuisance parameters entering linearly. Section 4 reports the simulation results. Section 5 concludes. An appendix contains the proofs and additional results.
We consider a moment inequality model of the form
where $A$ is a $d_A\times d_m$ matrix, $b$ is a $d_A\times 1$ vector, and $\overline{m}_n(\theta)=n^{-1}\sum_{i=1}^n m(W_i, \theta)$ for a $d_m$-dimensional moment function $m(\cdot,\theta)$ known up to the parameter $\theta$ and the data $\{W_i\}_{i=1}^n$ with joint distribution $F$. Let $\Theta$ be the parameter space for $\theta$. The quantities $A$ and $b$ may depend on $\theta$ and the sample size $n$, a dependence that we keep implicit for simplicity unless otherwise needed. The moment inequality model identifies the true parameter value up to the identified set,\footnote{If $A$ and $b$ depend on $\theta$, the formula for $\Theta_0(F)$ becomes $\{\theta\in \Theta: A(\theta)\mathbb{E}_{F}\overline{m}_n(\theta)\leq b(\theta)\}$.}
Our specification of a moment inequality model slightly differs from that in Andrews and Soares (2010, AS hereafter) by including a coefficient matrix $A$ and an intercept $b$. The model reduces to the setup of AS when we let $b=0$ and
where $d_A=p+2v$, the first $p$ moments are inequalities, and the last $v$ moments are equalities. Introducing $A$ and $b$ is useful because it allows us to maintain an invertible variance matrix assumption while succinctly covering both equalities and inequalities, as well as accommodating upper and lower bounds with a gap between bounds that is deterministic.\footnote{For example, $\mathbb{E}[\bar{W}_n]-1\leq \theta\leq \mathbb{E}[\bar{W}_n]$ can be written in our notation with $m(w,\theta) = \theta-w$, $A = \left(
\right)$ and $b = \left(
\right)$.}
Like most papers in the literature, including AS, AB, and RSW, we conduct inference for the true parameter $\theta_0$ by test inversion. That is, for a given significance level $\alpha\in (0,1)$, one constructs a test $\phi_n(\theta,\alpha)$ for $H_0:\theta=\theta_0$, and obtains the confidence set for $\theta_0$ by calculating
We introduce two simple tests, one being a refinement of the other. Both are easy to compute, requiring no tuning parameters or simulations. Both use the (quasi-) likelihood ratio statistic,
where $\widehat\Sigma_n(\theta)$ denotes an estimator of $\textup{Var}_F(\sqrt{n}\overline{m}_n(\theta))$.
Both tests use data-dependent critical values that are based on the rank of the rows of $A$ corresponding to the inequalities that are active in finite sample. To define them rigorously, let $\hat{\mu}$ be the solution to the minimization problem in ((ref)). This is the restricted estimator for the moments. Let $a'_j$ denote the $j$th row of $A$ and let $b_j$ denote the $j$th element of $b$ for $j = 1,2,\dots,d_A$. Let
which is the set of indices for the active inequalities. For a set $J\subseteq\{1,2,\dots,d_A\}$, let $A_J$ be the submatrix of $A$ formed by the rows of $A$ corresponding to the elements in $J$. Let $\textup{rk}(A_J)$ denote the rank of $A_J$. Let $\hat r = \textup{rk}(A_{\widehat J})$.\footnote{$\hat\mu$, $\hat J$, and $\hat r$ depend on $\theta$, a dependence that we keep implicit for simplicity.}
The critical value of the first simple test is the $100(1-\alpha)\%$ quantile of $\chi^2_{\hat r}$, the chi-squared distribution with $\hat r$ degrees of freedom. Thus, the first simple test is
where CC stands for “conditional chi-squared” indicating that the test uses the chi-squared critical value conditional on the active inequalities. We show the validity of the CC test below. The intuition is that $T_n(\theta)$ follows the $\chi^2_{\hat r}$ distribution conditional on $\hat r$ when all inequalities are binding (that is, $A\mathbb{E}_{F}\overline{m}_n(\theta) = b$), and is stochastically dominated by the $\chi^2_{\hat r}$ distribution when some of the inequalities are slack.
The CC test can be somewhat conservative because when $\hat r=0$ (effectively no inequality is active in finite sample), $T_n(\theta)$ is a point mass at zero and the conditional rejection probability is zero instead of $\alpha$. For this reason, the null rejection probability of the CC test can be as low as $(1-\Pr(\hat r=0))\alpha$, which lies in the interval $[\alpha/2,\alpha]$.
We propose a second simple test that eliminates the conservativeness. We call this the RCC (refined CC) test. We define the RCC test by adjusting the quantile of the $\chi^2_1$ distribution when $\hat r=1$. Instead of the $100(1-\alpha)\%$ quantile, the RCC test uses a $100(1-\hat\beta)\%$ quantile, where $\hat\beta$ varies between $\alpha$ and $2\alpha$ depending on how far from active the additional (inactive) inequalities are. We construct $\hat\beta$ carefully so that the refinement exactly restores the size of the test. When $\hat r=1$, suppose without loss of generality that the first inequality is active and satisfies $a_1\neq 0$.\footnote{In this case, other inequalities may be active too because we have not ruled out the possibility that $A$ contains redundant rows or zero rows. But this is possible only if the other active inequalities are collinear with $a_1$. } Now define a measure of how far from being active the other inequalities are. For each $j=2,...,d_A$, let
where $\|a\|_\Sigma = (a'\Sigma a)^{1/2}$. This is equal to zero when the $j$th inequality is active, and positive when it is inactive. It is scaled using the angle between $\widehat\Sigma^{1/2}_n(\theta)a_1$ and $\widehat\Sigma^{1/2}_n(\theta)a_j$.\footnote{Note that $a'_1\widehat\Sigma_n(\theta)a_j=\|a_1\|_{\widehat\Sigma_n(\theta)} \|a_j\|_{\widehat\Sigma_n(\theta)} \cos\varphi$, where $\varphi$ stands for the angle.} Then let
This quantity is easy to compute and has a nice geometric interpretation that is illustrated in Example (ref) below.
Now we can define
where $\Phi(\cdot)$ is the standard normal cumulative distribution function (cdf).\footnote{$\hat\beta$, $\hat\tau_j$, and $\hat\tau$ depend on $\theta$, a dependence that we keep implicit for simplicity.} When a second inequality is close to being active, $\hat \tau$ is close to $0$ and then $\hat\beta$ is close to $\alpha$. When all the other inequalities are far from active, then $\hat\tau$ is very large and $\hat\beta$ is close to $2\alpha$. We define the RCC test for $H_0:\theta=\theta_0$ to be
Since $\hat\tau\in[0,\infty]$, $\hat\beta\in[\alpha,2\alpha]$. Thus we have the following comparison of the CC and the RCC tests:
Moreover, when an equality is being tested, at least two inequalities are always active, in which case we have $\hat\beta=\alpha$, and the RCC test reduces to the CC test.
We also define a reduced test that only uses a subset of the inequalities. For $J\subseteq\{1,...,d_A\}$, let $\phi^{\textup{RCC}}_{n,J}(\theta,\alpha)$ denote the RCC test defined with $A_J$ and $b_J$ instead of $A$ and $b$, where $b_J$ denotes the subvector of $b$ formed by the elements of $b$ corresponding to the indices in $J$. This test is a useful point of comparison when the inequalities not in $J$ are very slack.
When the moments are normally distributed with known variance, the following theorem states the finite sample properties of the RCC test.
\noindentRemarks. (1) Part (a) shows the finite sample validity of the RCC test when the moments are normally distributed with known variance. Part (b) shows that the RCC test is size-exact when all the inequalities bind. Using ((ref)), part (b) also implies that the finite sample size of the CC test is between $\alpha/2$ and $\alpha$ if all the inequalities may bind. Part (c) shows that, in the presence of very slack inequalities, the RCC test reduces to the test that only uses the not-very-slack inequalities. Put another way, the RCC test adapts to the slackness of the inequalities. We call this property “irrelevance of distant inequalities” or IDI. This is especially useful if all but one inequality is very slack, because the reduced test is the one-sided t-test for the sole binding inequality, which is uniformly most powerful.
(2) Several papers, including Kudo1963 and Wolak1987, propose a classical test for inequalities that can be applied here. The classical test is based on the $1-\alpha$ quantile of the least favorable distribution of $T_n(\theta)$, which is a mixture of $\chi^2_0, \chi^2_1,\dots,\chi^2_{d_A}$ distributions. This test also has exact size, but lacks the IDI property. When many inequalities tested are slack, the power of this test can be very low. Besides, the critical value typically requires simulation, which makes it computationally less attractive than the RCC test.
(3) AS introduced a generalized moment selection procedure that achieves an asymptotic version of the IDI property via a sequence of tuning parameters. AB's asymptotic normality-based test has finite sample exact size. It nearly has the IDI property, but not exactly. The size correction the test uses causes it to respond to very slack inequalities, albeit to a lesser extent than the classical test.
(4) The only other moment inequality test with exact size and the IDI property is the non-hybrid test in AndrewsRothPakes2019. This test is based on the conditional distribution of the maximum standardized element of $\sqrt{n}(A\overline{m}_n(\theta)-b)$ given the second-largest maximum. The test also asymptotes to the one-sided t-test when all but one inequality get increasingly slack. However, the test has undesirable power when multiple inequalities are not well separated, which prompts them to recommend a hybrid test instead.
(5) Theorem 1 and the other results in this paper are stated in terms of hypothesis tests. However, they can be extended to results on the coverage probability of confidence sets defined by test inversion in a standard way. Specifically, under the conditions of Theorem 1(a), we have for all $\theta_0\in\Theta_0(F)$,
where $CS_n^{\text{RCC}}(1-\alpha)=\{\theta\in\Theta: \phi^{\text{RCC}}_n(\theta,\alpha)=0\}$ is the confidence set formed by inverting the RCC test.
(6) The proof of Theorem 1 is challenging. It relies on a partition of the space of realizations of the moments according to which inequalities are active. It then uses a bound on probabilities of translations of sets to bound the rejection probability conditional on each set in the partition. MohamadGoemanZwet2020 prove a special case of part (a) for the CC test when the inequalities define a cone. We extend the result to the RCC test and allow the inequalities to define an arbitrary polyhedron, which are important extensions for moment inequality models. \qed
It helps to demonstrate the CC and RCC tests in a simple two-inequality example.
Now we turn to a model without normality or a known variance matrix. For expositional purposes, we focus on the independent and identically distributed (i.i.d.) data case here, while results in Appendix (ref) cover more general cases.
With i.i.d. data, we can estimate the variance matrix $\Sigma_n(\theta)$ with the usual sample variance matrix:
We show that the RCC test has correct asymptotic size uniformly over a large class of data generating processes.
The following assumption defines the set of data generating processes allowed. Here $|\cdot|$ denotes the matrix determinant, and $\epsilon$ and $M$ are fixed positive constants that do not dependent on $F$ or $\theta$. This set of assumptions is weak and also seen in AndrewsGuggenberger2009 or AS.
Let $D_F(\theta)$ denote the diagonal matrix formed by $\sigma_{F,j}^2(\theta):j=1,\dots,d_m$. For $J\subseteq\{1,\dots,d_A\}$, let $I_J$ denote the rows of the identity matrix corresponding to the indices in $J$.\footnote{Note that $I_JA$ is an alternate notation for $A_J$.} The following theorem states the asymptotic properties of the RCC test.
\noindentRemarks. (1) Part (a) shows that the RCC test is asymptotically uniformly valid. Part (b) shows that when all the inequalities bind, or are sufficiently close to binding, the RCC test does not under-reject asymptotically. Part (c) shows an asymptotic IDI property of the RCC test: if some of the inequalities are very slack, the test reduces to the one based only on the not-very-slack inequalities. Parts (b) and (c) can be combined to show that the RCC test has exact asymptotic size (and hence is asymptotically non-conservative) when there exists at least one fixed triple $(F,\theta,j)$ such that $\theta\in\Theta_0(F)$, $a_j\neq \mathbf{0}$ and $a_j'\mathbb{E}_Fm(W_i,\theta) = b_j$.
(2) Theorem (ref) combines with ((ref)) to imply that the CC test is asymptotically uniformly valid, and when the RCC test is asymptotically non-conservative, it can only be conservative to a limited extent:
(3) The outline of the proof of Theorem (ref)(a) is conceptually simple. The almost sure representation theorem is invoked on the convergence of the moments, and then Theorem (ref) is invoked on the limiting experiment. However, the details are quite complicated. A technical complication that arises is that the rank of the inequalities can be lower in the limit than in the finite sample. This is handled by adding additional inequalities so the limit is an appropriate approximation to the finite sample (see Lemma (ref) in the appendix).
(4) Theorem (ref) is stated for finitely many inequalities, but the test may also work well in high dimensions. The biggest challenge in high dimensions is the covariance matrix estimation. If a good covariance matrix estimator can be found and the moments are approximately normal, then one can appeal to Theorem (ref) as a good approximation. One way to improve covariance matrix estimation is to assume sparsity or use shrinkage as in LedoitWolf2012. \qed
In this section, we apply the CC test and the RCC test to a subvector inference problem in a conditional moment inequality model:
where $Z=\{Z_i\}_{i=1}^n$ is a sample of instrumental variables (each $Z_i$ is taken to be a subvector of $W_i$ without loss of generality), $B_Z$, $C_Z$, and $d_Z$ are $k\times d_m$, $k\times p$, and $k\times 1$ matrices, $\delta$ is an unknown nuisance parameter, $\theta$ is the unknown parameter of interest, and $F_Z$ denotes the conditional distribution of $\{W_i\}_{i=1}^n$ given $\{Z_i\}_{i=1}^n$. The subscript $Z$ is used to denote dependence on $Z_1,...,Z_n$. The quantities $B_Z$, $C_Z$, and $d_Z$ may also depend on $\theta$ and the sample size $n$, a dependence that we keep implicit for simplicity. Similar to the full-vector case, this model can succinctly cover both equalities and inequalities by an appropriate choice of $B_Z$, $C_Z$, and $d_Z$, as well as accommodating upper and lower bounds with a gap between the bounds that depends on $\{Z_i\}_{i=1}^n$.
The model ((ref)) is similar to that considered in AndrewsRothPakes2019, and both are special cases of conditional moment inequality models compared to the setup of AndrewsShi2013. It has two key features: (a) the nuisance parameter $\delta$ enters linearly, and (b) the coefficients on $\delta$ depend only on the exogenous variables $\{Z_i\}_{i=1}^n$. These two features allow us to develop a simple subvector test that is also tuning parameter and simulation free.\footnote{AndrewsRothPakes2019 rely on these features as well.} Special as it is, these features reflect a recurring theme of many empirical models: exogenous covariates are used to incorporate heterogeneity and/or to control for confounders. Here we consider three examples.
Two additional examples that fit into our framework are Katz2007 and Wollman2018 as reviewed in AndrewsRothPakes2019.
To conduct inference on $\theta$, we invert a test for $H_0:\theta=\theta_0$. This is equivalent to the following null hypothesis for a given value $\theta_0$:
In the next few subsections, we proceed to construct a computationally simple, tuning parameter and simulation free test for ((ref)).
Directly testing ((ref)) is difficult because it requires checking the validity of the inequality for all possible values of $\delta$. Instead, we construct a representation of the hypothesis that does not involve $\delta$. The construction makes use of the following lemma that is a corollary of Theorem 4.2 of Kohler1967. Kohler's result is a refinement of the well-known Fourier-Motzkin method for eliminating variables from a system of linear inequalities. The latter was first introduced to the moment inequality literature by GuggenbergerHahnKim2008.
It is immediately implied by Lemma (ref) that the statement in ((ref)) is equivalent to
where $A_Z = A(B_Z,C_Z)$ and $b_Z=b(C_Z,d_Z)$, which, like $B_Z, C_Z$, and $d_Z$, may also depend on $Z_1,...,Z_n$. This is the same as model ((ref)) except that the expectation is conditional on the instrumental variables.
If $H(C_Z)$ and thus $A_Z$ and $b_Z$ can be obtained, one can apply a conditional version of any inequality testing procedures on ((ref)). However, a significant obstacle is that calculating $H(C_Z)$ requires enumerating the vertices of a polyhedron. Vertex enumeration is doable in small dimension, but becomes highly nontrivial as the number of conditional moment inequalities increases because the number of vertices can grow exponentially with $k$. See for example SierksmaZwols2015.
On the other hand, we shall see below that, if one applies the CC test on the hypothesis ((ref)), one only needs to know the {\em rank} of the rows of $A_Z$ corresponding to the active inequalities. And this rank can be obtained by solving a relatively small number of linear programming problems without computing $A_Z$ or $b_Z$. This is the key insight that makes the subvector version of the CC and RCC tests feasible.
Let $\widehat\Sigma_n(\theta)$ denote an estimator of the conditional variance matrix of the moments:
The subvector CC test for $H_0:\theta=\theta_0$ is the CC test defined in ((ref))-((ref)) for the inequalities ((ref)). We denote it by $\phi_n^{\textup{sCC}}(\theta,\alpha)$ to distinguish it from the full-vector CC test.
Now we describe how to compute $T_n(\theta)$ and $\hat r$ without computing $A_Z$ or $b_Z$. For $T_n(\theta)$, it is immediate from Lemma (ref) that it can be equivalently written as
Let $\hat{\mu}$ denote the minimizer. For $\hat r$, we first give the following lemma. In the lemma, ${\cal H} = \{h\in R^k:h\geq 0, C'_Zh=\mathbf{0},h'(B_Z\hat{\mu} - d_Z) = 0\}$, $B'_Z{\cal H} = \{B'_Zh:h\in{\cal H}\}$, and $\text{rk}(S)$ is the maximum number of linearly independent vectors in $S$.
The rank of the polyhedral cone $B'{\cal H}$ for some matrix $B$ is the dimension of $\textup{span}(B'{\cal H})$, the linear span of $B'{\cal H}$. The key to calculate it is to find the set of linear equalities that define $\textup{span}(B'{\cal H})$. This can be done by finding the parametric form of ${\cal H}$ as described in HuynhLassezLassez1992. That is, a representation of ${\cal H}$ of the form
where $J_{co}$ is a subset of $\{1,\dots,k\}$, $J_{co}^c$ is its complement, $h_J$ is the subvector of $h$ selected by the index set $J$, and $G_1$ and $G_0$ are two matrices constructed so that ((ref)) holds and $\{h_{J_{co}}:G_0h_{J_{co}}\geq 0\}$ is a full-dimensional polyhedral cone. In other words, the parametric form divides $h$ into a core subvector $h_{J_{co}}$ and a non-core subvector $h_{J_{co}^c}$, such that $h_{J_{co}^c}$ is the largest subvector that can be written as a linear combination of $h_{J_{co}}$. Based on ((ref)), it is clear that
Accordingly, algebra shows that
where $B_{J}$ is the matrix formed by the rows of $B$ corresponding to the indices $J$, and $k_{co}$ is the number of elements in $J_{co}$. This implies that $\text{rk}(B'{\cal H}) = \text{rk}(G_1'B_{J_{co}^c}+B_{J_{co}})$. Finally, invoking Lemma (ref), we have,
It remains to find the parametric form ((ref)). In particular we need to find the matrix $G_1$.\footnote{The matrix $G_0$ needs not be computed since it does not enter the subsequent rank calculation.} HuynhLassezLassez1992 present a way to do this by solving $k$ linear programming (LP) problems. The first step is to determine if the constraint $h_j\geq 0$ is an implicit (or implied) equality, which holds if
where $h_j$ is the $j$th element of $h$ and $h_{-j}$ is $h$ with its $j$th element removed. To determine if $h_j\geq 0$ is an implicit equality, one simply solves the LP problem:
where $e_j$ is the $j$th column of $I_k$. Then $h_j\geq 0$ is an implicit equality if and only if the minimum value of this linear programming problem is zero.
Let $J_{eq}$ denote the set of all $j\in \{1,\dots,k\}$ such that $h_j\geq 0$ is an implicit equality. Let $I_{J}$ denote the submatrix of $I_{k}$ formed by rows of $I_{k}$ corresponding to indices in $J$. Then ${\cal H}$ is defined by $h\geq 0$ and the following linear equation system
Finding the matrix $G_1$ amounts to solving for the $k-k_{co}$ redundant variables from this equation system where $k_{co} = k-rk\left(
\right)$. This can be done easily via Gauss-Jordan reduction.
In fact, in the special case that $\text{rk}(B_Z) = k$, we have $\text{rk}(B'_Z{\cal H}) = \text{rk}({\cal H})$, and there is no need to even calculate $G_1$. This is because $rk({\cal H}) = k_{co} = k-rk\left(
\right)$ by the definition of the parametric form.
To sum up, we propose the following procedure to implement the sCC test.
In this procedure, the most time consuming step is Step 2 where we solve $k$ LP problems. Fast and accurate algorithms for LP problems are well-developed and widely available, which makes the CC test feasible even for a large number of inequalities and nuisance parameters.
The subvector RCC test for $H_0:\theta=\theta_0$ is the RCC test defined in ((ref)) for the inequalities ((ref)). We denote it by $\phi_n^{\textup{sRCC}}(\theta,\alpha)$ to distinguish it from the full-vector RCC test.
Note that it is only when $\hat r = 1$ {\em and} $T_n(\theta)\in [\chi^2_{1,1-2\alpha},\chi^2_{1,1-\alpha}]$ that the RCC test can possibly differ from the CC test. Thus, the RCC test can be performed via the following steps.
The vertex enumeration part of Step R2 can be difficult for large $k$ and $p$. However, notice that Step R2 is only needed when $\hat r=1$ and $T_n(\theta)\in [\chi^2_{1,1-2\alpha},\chi^2_{1,1-\alpha}]$ in Step R1, which is infrequent for large $k$.
The following result shows the finite sample properties of the sRCC test assuming normally distributed moments and a known conditional variance matrix. The result is a corollary of Theorem (ref). Let $z$ be a realization of $Z$, and let $\Theta_0(F_z) = \{\theta\in\Theta: \exists\delta\textup{ s.t. }C_z\delta\geq B_z\mathbb{E}_{F_z}[\overline{m}_n(\theta)|z]-d_z\}$. Let $A_z=A(B_z,C_z)$ and $b_z=b(C_z,d_z)$ from Lemma (ref). Let $e_j$ denote the $R^{d_A}$-vector with $j$th element one and all other elements zero. For any $J\subseteq \{1,\dots,d_A\}$, let $\phi^{\textup{sRCC}}_{n,J}(\theta,\alpha)$ denote the sRCC test defined using $I_J A_z$ and $I_Jb_z$ in place of $A_z$ and $b_z$.
\noindentRemarks. (1) Part (a) shows the finite sample validity of the sRCC test under normality. Part (b) states mild conditions under which the sRCC test is not conservative. Part (c) states the IDI property of the sRCC test.
(2) Since the sRCC test rejects whenever the sCC test does, the corollary implies the validity of the sCC test: $\mathbb{E}_{F_z}[\phi_n^{\textup{sCC}}(\theta,\alpha)|z]\leq \alpha$.
(3) A result for asymptotically uniform size control of the subvector tests is available in the appendix. It relies on the asymptotic normality of the moments conditional on $Z_1,...,Z_n$ and a consistent estimator for $\Sigma_n(\theta)$. When the data are i.i.d., we provide low level conditions for two cases: discrete $Z_i$ and continuous $Z_i$.
In the first case, suppose $Z_i$ takes on a finite number of values in a set $\mathcal{Z}$. A straightforward estimator of $\textup{Var}_{F_z}(\sqrt{n}\overline{m}_n(\theta))$ is the weighted average of the sample variances of $m(W_i,\theta)$ within each category of $Z_i$:
where $n_\ell = \sum_{i=1}^n 1\{z_i=\ell\}$ and $\overline{m}_n^\ell(\theta) = \frac{1}{n_\ell}\sum_{i=1}^n m(W_i,\theta)1\{z_i=\ell\}$. As we show in Appendix (ref), sufficient conditions for the consistency of this estimator involve boundedness of the fourth moment of $m(W_i,\theta)$ and the assumption that every $z_i$ value occur twice or more in the sample $\{z_i\}_{i=1}^n$ eventually.
When $Z_i$ contains continuous random variables, one can use a nearest neighbor matching estimator similar to that used for the standard error of a regression discontinuity estimator in Abadie, Imbens, and Zheng (2014).\footnote{This is also the estimator used in AndrewsRothPakes2019.} Let $\Sigma_{Z,n} = n^{-1}\sum_{i=1}^n(Z_i-\overline{Z}_n)(Z_i-\overline{Z}_n)'$ where $\overline{Z}_n = n^{-1}\sum_{i=1}^n Z_i$. For each $i$, define the nearest neighbor to be
When the argmin is not unique, picking one randomly does not affect the consistency of the resulting estimator. The estimator of $\Sigma_n(\theta)$ is then given by
As we show in Appendix (ref), sufficient conditions for the consistency of this matching estimator involves the boundedness of $\{z_i\}_{i=1}^\infty$ and the Lipschitz continuity of $\textup{Var}(m(W_i,\theta)|z_i)$ in $z_i$. \qed
We consider two sets of Monte Carlo simulations, one to evaluate the performance of our tests in a general moment inequality model without nuisance parameters, and the second to evaluate the performance of our subvector tests in the interval regression model of GandhiLuShi2019.
Our first set of simulations takes the generic moment inequality design from AB. This design allows a variety of correlation structures across moments and thus can approximate a wide range of applications.
We briefly describe the Monte Carlo design here and refer the readers to Section 6 of AB for further details. Consider the moment inequality model
and the null hypothesis $H_0:\theta=\mathbf{0}$, where $W_i$ is a $p$-dimensional random vector. Let the data $\{W_i\}_{i=1}^n$ be i.i.d. with sample size $n$. Let $W_i \sim N(\mu,\Omega)$, where $\Omega$ is a correlation matrix and $\mu$ is a mean-vector. Three choices of $\Omega$ are considered: $\Omega_{\text{Neg}}$, $\Omega_{\text{Zero}}$, and $\Omega_{\text{Pos}}$, indicating negative, zero, and positive correlation among the moments, respectively. The exact numerical specifications of these matrices for different $p$'s are in Section 4 of AB and Section S7.1 of the Supplemental Material of AB. Also, three choices of $p$: $2$, $4$, and $10$, are considered.
For each combination of $\Omega$ and $p$, we approximate the size of the tests using the maximum null rejection probability (MNRP) over a set of $\mu$ values that satisfies $\mu\geq \mathbf{0}$. These $\mu$ values are taken from AB whose calculations suggest that these points are capable of approximating the size of the tests. We compute a weighted average power (WAP) for easy comparison. The WAP is the simple average of a set of carefully chosen points in the alternative space. We take these points also from AB who design them to reflect cases with various degrees of violation or slackness for each of the inequalities. These $\mu$ values are given in Section 4 of AB and Section S7.1 of the Supplemental Material of AB. Besides WAP, we also report size-corrected WAP, which is obtained by adding a (positive or negative) number to the critical value where the number is set to make the size-corrected MNRP equal to the nominal level.
We report Monte Carlo simulation results to compare the CC and the RCC tests to the recommended tests in AB and RSW, more specifically, the bootstrap-based AQLR (adjusted quasi-likelihood ratio) test in AB and the two-step procedure in RSW.\footnote{We use the AB test for comparison because it is tuning parameter free like ours (in the sense that AB propose and use an optimal choice of the AS tuning parameter), and we use RSW's two-step test for comparison because it should be insensitive to reasonable choices of its tuning parameter.}
Two sets of results are reported. The first set assumes a known $\Omega$ and is reported in Table (ref). In this case, the RCC test should have exact size and the CC test should be somewhat under-sized especially with small $k$. These theoretical predictions are exactly confirmed in the table. The second set of results does not assume a known $\Omega$ and is reported in Table (ref). In this case, the RCC test still has very good MNRP at $k=2$, but has some noticeable over-rejection when $k=10$ with $\Omega = \Omega_{\text{Neg}}$ and $\Omega_{\text{Zero}}$. This may reflect the difficulty in estimating the large dimensional $\Omega$ with a small sample size ($n=100$). In comparison, the AB and the RSW tests exhibit under-rejection in some cases and over-rejection in other cases, even with known $\Omega$. This could be the result of the relatively small number of critical value simulations (1000 for AB and 499 for RSW, as recommended therein). Even with the small number of critical value simulations, the AB test and the RSW test are 200-400 times more costly than the RCC test, as shown in the ReT column in the tables. Increasing this number will increase their computational cost proportionally.
In terms of power, we find that the RCC test compares favorably to the AB and RSW tests. The biggest advantage of the RCC test seems to come from the points with a small number of violated inequalities and a large number of mildly slack inequalities while the magnitude of the advantage varies with $\Omega$.\footnote{For example, with $k=10$, $\Omega=\Omega_{\text{Zero}}$ and estimated $\Omega$, at the point $\mu = (-0.268,0.1, 0.1, 0.1, 0.1, \allowbreak0.1,0.1, 0.1,0.1, 0.1)'$, the size-corrected power of the RCC test is 0.55 while that for AB and RSW are respectively 0.28 and 0.22.} By the uni-dimensional criterion ScWAP, the RCC test is better than or the same as the RSW test in all cases but the case with $k=10, \Omega=\Omega_{\text{Pos}}$ with unknown $\Omega$ where the RSW test under-rejects and the size-correction adjusts its power up. Also according to ScWAP, the RCC test has better or equal power as AB in half of the cases, while the difference in all cases is small.
Overall, when we consider not only size and power from the simulations, but also computational cost, simplicity, simulation error (or lack thereof) of the critical values, and finite-sample properties, we recommend the RCC test over the other tests.
Now consider a special case of Example (ref) above. More specifically, we consider a model where $Y_i^\ast=s^\ast_i $ is the probability of an event of interest, for example, death by homicide for a random person county $i$, or a product being purchased by a random consumer in market $i$. For simplicity, we consider a simple logit model for the probability: $s_i^\ast = \frac{\exp(X'_i\theta_0+Z'_{ci}\delta_0+\varepsilon_i)}{1+\exp(X'_i\theta_0+Z'_{ci}\delta_0+\varepsilon_i)}$, where $\varepsilon_i$ is the country or market level unobservable that satisfies $\mathbb{E}[\varepsilon_i|Z_i]=0$. Then ((ref)) holds with
The variable $s^\ast_i$ is unobserved, but we observe $s_{N,i}$, an empirical estimate of $s^\ast_i$ based on $N$ independent chances for the event of interest to happen: $Ns_{N,i}$ follows a binomial distribution with parameters $(N,s^\ast_i)$. For example, $N$ could be the population of the county and $s_{N,i}$ is the homicide rate of the county. We use the method introduced in GandhiLuShi2019 to construct $\psi_{i}^L(\theta)$ and $\psi_{i}^U(\theta)$ based on $s_{N,i}$. By GandhiLuShi2019, for $N\geq 100$, the following construction satisfies ((ref)):
where $\underline{s}$ is the smaller of $0.05$ and half of the minimum possible value of $\min(s^\ast_i,1-s^\ast_i)$.\footnote{ These bounds are not necessarily sharp, but that is not important for our purpose, which is to investigate the statistical performance of the sCC and sRCC tests.} We assume that $\underline{s}$ is known and refer the reader to GandhiLuShi2019 for practical recommendations regarding $\underline{s}$.\footnote{It is worth pointing out that we deviate from GandhiLuShi2019 by holding $N$ fixed as the sample size (i.e. the number of observations for $(s_{N,i},X_i,Z_{ci})$) increases, and thus do not aim for point identification.}
For this simulation, we consider a scalar endogenous variable $X_i$, a scalar excluded instrument $Z_{ei}$ and a $d_c$-dimensional exogenous covariate $Z_{ci} = (1,Z_{c,2,i},\dots,Z_{c, d_c,i})'$. To generate the data, we let $N=100$, $\varepsilon_i\sim \min\{ \max\{-4,N(0,1)\},4\}$. We let the non-constant elements of $Z_i$ be mutually independent Bernoulli variables with success probability $0.5$. Let $X_i = 1\{Z_{ei}+\varepsilon_i/2>0\}$. Also consider $\theta_0=-1$, $\delta_0=(0,-1, \mathbf{0}_{d_c-2}')'$. Given this data generating process, the lowest and the highest possible values for $s^\ast_i$ are respectively $$ \frac{\exp(-6)}{1+\exp(-6)} = 0.0025\text{ and }\frac{\exp(4)}{1+\exp(4)}=0.982. $$ Thus $\underline{s} = 0.00125$.
We calculate that the identified set of $\theta_0$ is approximately $[-1.203,-0.757]$. Details of the calculation are given in Appendix (ref).
We simulate the rejection rate of the tests using 5000 Monte Carlo repetitions. In each repetition, we generate an i.i.d. data set $\{s_{N,i},X_i, Z_i\}_{i=1}^n$, for two sample sizes $n=500$ and $n=1000$. We consider three cases of $d_c$: $d_c=2,3$ and $4$.
For instrumental functions, we use $I(Z_i) = (I_{j,\ell}(Z_i))_{j, \ell = 2,\dots,d_c+1:j\neq \ell}$, where
where $Z_{ji} $ is the $j$th element of $Z_i$. Thus, when $d_c=2$ (or $3$, $4$), there are $4$ (or $12$, $48$) instrumental functions, which give us $8$ (or $24$, $96$) moment inequalities.
Figure (ref) reports the rejection rates of the sCC test and the sRCC test for $H_0:\theta_0=\theta$ at $\theta$ values in $[-2.5,0.5]$ in the three cases of $d_c$ and two sample sizes. The shaded area indicates the identified set for $\theta_0$. As we can see, both tests reject $H_0$ at rates lower than $5\%$ for $\theta$'s in the identified set. The rejection rates are closer to 5% on the boundary of the identified set than in the interior and the sRCC test is better than the sCC test in all cases. The rejection rate grows steadily toward 1 as $\theta$ moves away from the identified set, and it grows with the sample size as well, as expected.
On the computational side, fixing the sample size at $n=1000$ and computing each test 5000 times, we document the time needed to compute the rejection rate of each test at one $\theta$ value on a computer with Intel X5680 3.33HZ CPU and 12Gb RAM running Matlab 2018. We find that on average, the subvector CC test uses about 13 minutes, 29 minutes, and 43 minutes respectively for $d_c=2$ (8 inequalities and 2 nuisance parameters), $d_c=3$ (24 inequalities and 3 nuisance parameters), and $d_c=4$ (96 inequalities and 4 nuisance parameters). The subvector sRCC test uses a similar amount of time as the sCC test when $d_c=2$ and $3$. It uses noticeably more time (2.2 hours) when $d_c=4$, which recall is the time to perform the test 5000 times when there are 96 conditional inequalities to begin with and 4 nuisance parameters to eliminate. Thus, for the scale of the job, the computational cost is quite modest.
The advantage of avoiding the computation of $H(C_Z)$ is very large at larger $d_c$. If we do not use the procedure described in Section (ref), but instead compute $H(C_Z)$ every time we perform the sRCC test, then the computation time for $d_c=3$ becomes 14 hours, and that for $d_c=4$ becomes a whopping 300 hours.\footnote{The matrix $H(C_Z)$ is found to have around 19, 1600, and 127550 rows, respectively for $d_c=2$, $3$, and $4$.} Therefore, we recommend our procedure over performing subvector tests via brute force elimination of the nuisance parameters.
For comparison, we also compute the rejection rates of two projection-based tests: the projection RCC test with unconditional variance (Proj-U) and the projection RCC test with conditional variance (Proj-C). Both tests use the full-vector RCC test. That is, the RCC test for the null hypothesis $H_0:(\theta',\delta')' = (\theta'_0,\delta'_0)'$) for the model $\mathbb{E}_{F_z}[\overline{m}_n(\theta) -C_z\delta]\leq 0$.\footnote{In this example, $B_z = I$ and $d_z=\mathbf{0}$.} The difference is that the Proj-C test uses the average conditional variance estimator $\widehat{\Sigma}_n(\theta)$ in ((ref)), while the Proj-U test uses the unconditional variance estimator, $\widehat{\Sigma}_n^{\textup{U}}(\theta,\delta) =$
where $C_{z_i}$ denotes the $C_z$ matrix evaluated with only the $i$th observation for $z$. Suppose that the test statistic and the critical value for the full-vector RCC test using $\widehat{\Sigma}_n(\theta)$ are denoted $T_n^{\textup{C}}(\theta,\delta)$ and $\textup{cv}_n^{\textup{C}}(\theta,\delta,1-\alpha)$, and those using $\widehat{\Sigma}_n^{\textup{U}}(\theta,\delta)$ are denoted $T_n^{\textup{U}}(\theta,\delta)$ and ${\textup{cv}}_n^{\textup{U}}(\theta,\delta,1-\alpha)$. Then the Proj-C and the Proj-U tests are, respectively,
Computing the projection-based tests is tricky even though we already use the computationally simple RCC test for $\textup{cv}$. This is because the infimum over $\delta$ is taken over a non-convex function, and finding global minimum over a non-convex function is challenging. In section (ref) in the appendix, we detail the numerical algorithm used to calculate the global minimum. The algorithm works well, but is not guaranteed to find the global minimum. Not finding the global minimum biases the power upward, and thus what we find is an {\em upper} bound for the power of the projection-based tests.
The results are plotted in Figure (ref) for $d_c=2$ and sample sizes $n=500$ and $1000$. As we can see, the Proj-C test and the Proj-U test perform similarly and both are uniformly dominated by the sRCC test. Computationally, the projection tests (for one $\theta$ value and 5000 Monte Carlo simulations) each took more than 1 hour, or more than 4 times the 13 minutes needed for the sRCC test. Therefore, on the basis of both power and computational cost, we recommend the sRCC test over the projection-based tests.
This paper proposes the refined conditional chi-squared (RCC) test for moment inequality models. This test compares a quasi-likelihood ratio statistic to a chi-squared critical value, where the number of degrees of freedom is the rank of the active inequalities. This test has many desirable properties, including being simple, adaptive, and tuning parameter and simulation free. We show that, with an easy refinement, it has exact size in normal models and has uniformly asymptotically exact size more generally. We also propose a version of the test for subvector inference with conditional moment inequalities and when the nuisance parameters enter linearly that has computational and power advantages.