EconBase
← Back to paper

Constraint Qualifications in Partial Identification

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

77,964 characters · 11 sections · 43 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Constraint Qualifications in Partial Identification

abstractThe literature on stochastic programming typically restricts attention to problems that fulfill constraint qualifications. The literature on estimation and inference under partial identification frequently restricts the geometry of identified sets with diverse high-level assumptions. These superficially appear to be different approaches to closely related problems. We extensively analyze their relation. Among other things, we show that for partial identification through pure moment inequalities, numerous assumptions from the literature essentially coincide with the Mangasarian-Fromowitz constraint qualification. This clarifies the relation between well-known contributions, including within econometrics, and elucidates stringency, as well as ease of verification, of some high-level assumptions in seminal papers.

\onehalfspacing

Introduction

This paper connects two related but largely separate literatures, namely statistical analysis of stochastic programs and estimation and inference for (functions of) partially identified parameters. The bounds that are pervasive in the latter literature are often expressed as values of constrained optimization problems. As such, some similarity to stochastic programming is rather apparent and has been observed before. However, we uncover much deeper connections between these literatures.

Our discussion starts from the econometrics literature. In a seminal paper, Chernozhukov_Hong_etal2007aE provide a comprehensive analysis of consistency of criterion function-based set estimators and their convergence rates in partially identified models. Their work highlights the challenges a researcher faces in this context and puts forward possible solutions in the form of assumptions under which specific rates of convergence attain. While these assumptions can be dispensed with when the researcher's goal is to obtain a confidence set for the partially identified parameter vector that is --pointwise or uniformly-- consistent in level (e.g., AndrewsSoares2010E), related assumptions reappear when the aim is to obtain a confidence interval for a smooth function of the partially identified parameter vector that is --pointwise or uniformly-- consistent in level (e.g., PakesPorterHo2011, BCS14_subv).\footnote{We cite PakesPorterHo2011 because the published version PPHI_ECMA does not contain the inference procedure. However, this procedure has been used in influential papers Eizenberg,HoPakes14,Holmes11.} Some more recent contributions ChoRussell,Gafarov observe a connection to stochastic programming and show that inference becomes much more tractable under the so-called Linear Independence constraint qualification. Some obvious questions arise: How do all these assumptions relate? What is the trade-off between them and possibly other assumptions from the statistics literature?

For a sense of why consistent estimation of identified sets or their projections can be hard even in otherwise well-behaved moment inequality settings, consider Figure (ref). Both panels illustrate a detail of an identified set $\Theta_I$ defined as the collection of parameter values satisfying a finite number of moment inequalities (as in (ref) below, but without equality constraints). Each restriction is represented by a curve and excludes everything above that curve. The resulting $\Theta_I$ is shaded.\footnote{An exact algebraic description of the example in Figure (ref) is as follows. Let $\Theta=\mathbb{R}^2$ with typical element $\theta=(\theta_1,\theta_2)$ and let $\Theta_I=\{\theta:(\theta_2-1)^3+\theta_1E(X_1)\leq 0,(\theta_2-1)^3-\theta_1E(X_1)\leq 0\}$ for the left panel; for the right panel, additionally require $\theta_2\leq E(X_3)$, where $E(X_1)=E(X_2)=E(X_3)=1$. Looking ahead, the right panel of Figure (ref) and both panels of Figure (ref) are qualitative representations.} Suppose now that a researcher wants to find either (i) a set estimator of $\Theta_I$ that is consistent in Hausdorff distance (defined later) or (ii) an estimator of a linear projection of $\Theta_I$ (e.g., the identified set for a component of $\theta$), represented through the support function

equation[equation omitted — 93 chars of source]

in pre-determined direction $p$. (In the figure, $p=(0,1)$, i.e. we maximize $\theta_2$.) Even if all constraints' graphs can be estimated at a specific --e.g., parametric-- rate, it does not follow that their intersection estimates either object of interest at the same or indeed at any rate. For example, if estimators approximate the true constraints from below, the support function may be underestimated including in the limit.

figure[figure omitted — 3,119 chars of source]

The examples may appear “knife-edge." However, note that: (i) While we will, in this paper, take a pointwise perspective to simplify the analysis, the literature on partial identification is usually concerned with inference that is uniformly valid near such irregular cases because asymptotic approximations may otherwise be misleading. Indeed, this is emphasized in the abstract of CS17. (ii) Inference methods that are uniformly valid typically use “Generalized Moment Selection" methods (see again CS17 for details) that account for statistical uncertainty not only of moment conditions that are violated in sample, but also of ones that are local-to-binding. In these methods, the “overidentified" feature of the right-hand panel, i.e. the intersection of more than $d$ constraints at one point in $\mathbb{R}^d$, becomes typical of bootstrap d.g.p.'s. In addition, this feature characterizes the boundary case of overidentifying moment conditions, though some assumptions discussed below will exclude that case anyway.

A reader familiar with constraint qualifications in optimization problems will recognize that both panels of Figure (ref) violate some of these qualifications. Conversely, a reader who is very familiar with the partial identification literature might recognize that they violate assumptions in CHT, PPHI, and elsewhere. We ask if this reflects deeper relations between these bodies of literature. The answer will be affirmative: Under a background assumption of continuous differentiability of moment conditions and abstracting from details of “how uniformly" the assumptions are stated, the literature on partial identification already invokes constraint qualifications; for examples, we show that both papers just cited rely on the Mangasarian-Fromowitz constraint qualification. This implies some previously unrecognized (to our knowledge) logical relations between assumptions made in econometrics.

Some references to the literatures that we connect are as follows. MolinariHOE gives a current overview over the field of partial identification. CS17 provide a definitive treatment of the literature on moment inequalities. We will define constraint qualifications below but refer to baz:she:she06 for a textbook treatment and to BonnansShapiroBook for a textbook on perturbation analysis of stochastic programs.

Assumptions

This section clarifies the general setup and introduces a broad array of assumptions from the literature. For the reader's convenience, these assumptions, and some results that we will explain later, are summarized in Table (ref).

We assume throughout that the model is correctly specified and that individual moment conditions are well-behaved. Specifically, the identified set $\Theta_I$ is characterized by $J$ constraints, namely $J_1\le J$ moment inequalities and $J_2\equiv J-J_1$ moment equalities:

equation[equation omitted — 149 chars of source]

where the functions $(m_1(\cdot),\dots,m_J(\cdot))$ are known up to $\theta \in \Theta$. Then we impose:

assumption(a) $\Theta\subset\mathbb{\ R}^d$ is compact convex with nonempty interior. (b) $\Theta_I \neq \emptyset$ and $\Theta_I \subset \operatorname{int}(\Theta)$. (c) $\sigma^2_j(\theta)\equiv \operatorname{Var}(m_j(X,\theta)) \in(0,\infty)$ for $j=1,\dots, J$. (d) The gradients $D_j(\cdot)\equiv\nabla_\theta\{E[m_j(X,\cdot)]/\sigma_j(\cdot)\},j=1,\dots,J$ exist and are continuous.

These assumptions are standard in the literature. The requirement that $\Theta_I \subset \operatorname{int}(\Theta)$ may appear stronger than the comparable one in CHT, i.e. their condition M2. However, the latter imposes continuous differentiability of moment conditions on a small enlargement of $\Theta$, so it does constrain $\Theta_I$ to be interior to the set on which local linear approximation of moment conditions is valid. This is how we use the assumption, and we could analogously weaken it.

We next state numerous assumptions that are inspired by the aforementioned literature in econometrics. We first state them in a way that maximizes resemblance to the original formulation, subject to the unification that assumptions are stated pointwise (not uniformly) over d.g.p.'s and that their local implications near extreme points of $\Theta_I$ are extracted. Universal constants invoked in assumptions need not take the same value across appearances.

Define the support set of $\Theta_I$ as $$ S(p,\Theta_I) \equiv \{\theta \in \Theta_I: p'\theta=s(p,\Theta_I)\},$$ where the support function $s(\cdot)$ is defined in (ref), and the supporting hyperplane as $$ H(p,\Theta_I) \equiv \{\theta \in \Theta: p'\theta=s(p,\Theta_I)\}.$$ Any element of $S(p,\Theta_I)$ is also called a support point. We use $\theta^*$ to denote a generic support point; to economize on subscripts and because we consider $p$ fixed, we suppress dependence of $\theta^*$ on $p$. Recall also that a constraint is active at $\theta$ if $E(m_j(X,\theta))=0$. Let

equation[equation omitted — 118 chars of source]

denote the set of active inequality constraints and

equation[equation omitted — 142 chars of source]

the full active set of (equality or inequality) constraints at $\theta$.

We first adapt two assumptions, “Degeneracy" and “Polynomial Minorant," from CHT. These are essential for getting rate results for consistency of analog estimators of $\Theta_I$; in particular, Polynomial Minorant ensures rate results for a relaxed sample analog of $\Theta_I$, and Degeneracy allows one to drop the relaxation. We weaken the assumptions insofar as they are only imposed at support points. Also, we do not adapt the high-level Conditions C.2 and C.3 from CHT because, being about sample objects, they restrict the sampling process and not just population moments. To keep these issues separate, we focus on the sufficient conditions that stand in for the assumptions in a moment inequalities setting (i.e., CHT's displays 4.5 and 4.6).

assumptionFor each support point $\theta^*\in S(p,\Theta_I)$, there exist constants $\delta>0$, $\eta>0$, $M>0$, and $C>0$ s.t. \begin{gather} \max_{j=1,\dots,J} E(m_j(X,\theta))/\sigma_j(\theta) \leq -C \varepsilon for all \theta \in \Theta_I^{-\varepsilon} \cap B(\theta^*,\eta) \\ \max_{\theta \in \Theta_I \cap B(\theta^*,\eta)} d(\theta,\Theta_I^{-\varepsilon}) \leq M\varepsilon for all \varepsilon \in [0,\delta], \end{gather} where $\Theta_I^{-\varepsilon}=\{\theta \in \Theta_I:d(\theta,\Theta \setminus \Theta_I) \geq \varepsilon\}$ and $B(\theta^*,\eta)\equiv \{\theta \in \Theta: \Vert \theta-\theta^* \Vert \leq \eta \}$.

This “degeneracy" assumption ensures that for any support point $\theta^*$, there exists a nearby point where moment inequalities hold with slack and whose membership in $\Theta_I$ is therefore easy to determine. In particular, we show in the proof of Theorem (ref) that Assumption (ref) precludes the existence of equality constraints.

In the original version, the equivalent of (ref) was stated using Hausdorff distance, i.e. $d_H(\Theta_I,\Theta_I^{-\varepsilon}) \leq M\varepsilon$, where $d_H(A,B)=\max\{\max_{a \in A} d(a,B),\max_{b \in B} d(b,A)\}$ for generic sets $A,B$ with typical elements $a,b$. However, $\max_{\theta \in \Theta_I^{-\varepsilon}} d(\theta,\Theta_I)=0$ because $\Theta_I^{-\varepsilon} \subset \Theta_I$, so here and in the original, only the implied restriction on $\max_{\theta \in \Theta_I} d(\theta,\Theta_I^{-\varepsilon})$ is nonvacuous. We localize it by restricting $\theta$ to a neighborhood of the support point.

assumptionFor each support point $\theta^*\in S(p,\Theta_I)$, there exist constants $\eta>0$, $c>0$, and $C>0$ such that, for all $\theta \in B(\theta^*,\eta)$, $$ \max\bigl\{0,\max_{j=1,\dots,J_1}E(m_j(X,\theta))/\sigma_j(\theta),\max_{j=J_1+1,\dots,J}\vert E(m_j(X,\theta))/\sigma_j(\theta)\vert\bigr\} \geq C \times \min\{d(\theta,\Theta_I),c\}. $$

This “minorant" assumption ensures that the population criterion increases not too slowly as one moves away from $\Theta_I$. Loosely speaking, it prevents “weak identification" problems. Next, BCS impose a polynomial minorant condition as well.\footnote{In the respective originals, both assumptions raise the r.h.s. to a power $\gamma>0$, hence “polynomial minorant." However, in both cases, $\gamma$ is restricted in ways that imply $\gamma=1$ in the present setting.}

assumptionFor each support point $\theta^*\in S(p,\Theta_I)$, there exist constants $\eta>0$, $c>0$, and $C>0$ such that, for all $\theta \in B(\theta^*,\eta) \cap H(p,\Theta_I)$, $$ \max\bigl\{0,\max_{j=1,\dots,J_1}E(m_j(X,\theta))/\sigma_j(\theta),\max_{j=J_1+1,\dots,J}\vert E(m_j(X,\theta))/\sigma_j(\theta)\vert\bigr\}\geq C \times \min\{d(\theta,S(p,\Theta_I)),c\} . $$

It bears emphasis that the last two assumptions are logically independent. Assumption (ref) forces the population criterion to increase (to first order) as we move away from $\Theta_I$. Assumption (ref) enforces an analogous increase as we move away from the support set $S(p,\Theta_I)$ along the supporting hyperplane. This does not imply Assumption (ref) because it only applies as one leaves $\Theta_I$ in selective directions. But it also is not implied because the directions considered in Assumption (ref) may be tangential to $\Theta_I$, in which case the increase in $d(\theta,\Theta_I)$ may be of low order. Indeed, Assumption (ref) may be considered restrictive: By not allowing directions to be tangential to $\Theta_I$ without also being tangential to $S(p,\Theta_I)$, it excludes smooth maxima, e.g. any identified set whose entire boundary is a smooth manifold.\footnote{Consider the single moment inequality $\theta_1^2+\theta_2^2-E(X) \leq 0$, with $E(X)=1$ and $p=(0,1)$. Then $S(p,\Theta_I)$ is a singleton at $\theta^*=(0,1)$. Let $\theta=(\zeta,1)$, then $\theta_1^2+\theta_2^2-E(X)=\zeta^2$, whereas $d(\theta,S(p,\Theta_I))=\zeta$, violating Assumption (ref) as $\zeta \to 0$; at the same limit, we have $d(\theta,\Theta_I)=\sqrt{1+\zeta^2}-1>\zeta^2/4$, so that Assumption (ref) holds. We note that this and other easy counterexamples to Assumption (ref) do not seem to be counterexamples to the BCS profiling method. We leave to future research the question whether this illustrates the possibility of relaxing their assumptions.}

We next adapt (in this order) assumptions 3, 4(a), and 4(b) from PPHI. These are modified in a few ways: PPHI assume that the support point $\theta^*$ is unique; this assumption is removed. Assumptions are also localized by only looking at $\theta$ near $\theta^*$; this makes the first of them meaningfully weaker. At the same time, PPHI impose assumptions uniformly over d.g.p.'s; to keep this paper focused, such uniform statements are removed throughout.\footnote{To compare the statements of assumptions, also keep in mind notational conventions: PPHI write moment conditions as $E(m_j(\cdot))\geq 0$, set $p=(-1,0,\dots,0)$, and use $1$-subscripts to denote first components of vectors, so that $t_1$ there would be $-p't$ here.}

assumptionFor each support point $\theta^*\in S(p,\Theta_I)$, there exist $\delta,\eta>0$ s.t. $\sup_{t \in T(\theta^*,\eta)}p't \leq -\delta$, where $$T(\theta^*,\eta) \equiv \left\{t=\frac{\theta-\theta^*}{\Vert \theta-\theta^*\Vert}:\theta \in \Theta_I \cap B(\theta^*,\eta),\theta \neq \theta^* \right\}.$$
assumptionThere are no equality constraints, i.e., $J_1=J$. Furthermore, for each support point $\theta^*\in S(p,\Theta_I)$, there exist constants $\delta,\varepsilon>0$ as well as direction (i.e. unit vector) $t \in \mathbb{R}^d$ s.t. $$\max_{j:E(m_j(X,\theta^*))/\sigma_j(\theta^*) \geq -\delta}D_j(\theta^*)t \leq -\varepsilon.$$
assumptionFor each support point $\theta^*\in S(p,\Theta_I)$, there exist constants $\delta<0$, $\varepsilon>0$ s.t. $$ \inf_{t \in \mathbb{R}^d,\Vert t \Vert=1,p't \geq \delta} \max\bigl\{ \max_{j \in \mathcal{J}_1^*(\theta^*)} D_j(\theta^*)t, \max_{j \in \{J_1+1,...J\} } \vert D_j(\theta^*)t\vert \bigr\} > \varepsilon.$$

Brief intuitions for these are as follows. Assumption (ref) ensures that $\Theta_I$ is contained in a cone that has $\theta^*$ as apex and does not otherwise intersect $H(p,\Theta_I)$. In particular, this implies uniqueness of $\theta^*$ (although we will not use this feature) and pointiness of the tangent cone (defined later) at $\theta^*$. Assumption (ref) ensures that locally to $\theta^*$, there exists a direction in which all moment expectations, hence their maximum, strictly decrease. Note in particular that this excludes equality constraints -- while this implication is not explicit in PPHI, they treat equalities as conjunctions of two inequalities, and the assumption cannot possibly hold for both. PPHI point out that their assumption is inspired by CHT's Degeneracy (i.e., our Assumption (ref)), which also excludes equalities. Indeed, both assumptions force $\Theta_I$ to have an interior, and we will have more to say about their relation later. Assumption (ref) enforces that the criterion is strictly increasing in many directions from $\theta^*$, including all that have positive inner product, and some that have negative inner product, with $p$.\footnote{The restriction $\delta<0$ in Assumption (ref) correctly reflects our source. However, we considered the possibility that (in our notation) $\delta>0$ was intended. The assumption then becomes weaker. Specifically, along the lines of our main result below, one can show that it is then equivalent to $\max_{j \in \mathcal{J}^*(\theta^*)} D_j(\theta^*)p >0$ and is implied by Assumption (ref).}

We finally recall some classic constraint qualifications. This requires some standard notation that we will also use later. For any $\theta \in \Theta_I$, define the tangent cone $$ \mathcal{T}(\theta) \equiv \biggl\{t \in \mathbb{R}^d: \exists \{\theta_m\}_{m=1}^{\infty} \subset \Theta_I, \theta_m \to \theta, \lim_{m \to \infty} \frac{\theta_m-\theta}{\Vert \theta_m-\theta \Vert} = \frac{t}{\Vert t \Vert} \biggr\} \cup \{0\} $$ as well as the linearized cone $$ \mathcal{L}(\theta) \equiv \bigl\{t \in \mathbb{R}^d: D_j(\theta)t \leq 0,j \in \mathcal{J}_1^*(\theta);D_j(\theta)t = 0,j \in \{J_1+1,\dots,J\} \bigr\}. $$ We will only invoke these objects for support points $\theta^* \in S(p,\Theta_I)$. Recall that $\mathcal{T}(\theta) \subseteq \mathcal{L}(\theta)$ baz:she:she06.

Both cones are illustrated in Figure (ref). They coincide in “nice" examples like the right panel of Figure (ref) (indeed, this “niceness" is the Abadie constraint qualification defined below) but they may disagree, as in the left panel which reprises the left panel from Figure (ref).

figure[figure omitted — 3,179 chars of source]

We can then state the following constraint qualifications in decreasing order of restrictiveness.

assumption{Linear Independence Constraint Qualification (LICQ)} For each support point $\theta^* \in S(p,\Theta_I)$, the active (equality or inequality) constraints have linearly independent gradients $D_j(\theta^*)$.

Of course, the LICQ requires at most $d$ active constraints, making it quite restrictive.

assumption{Mangasarian-Fromowitz Constraint Qualification (MFCQ)} At each support point $\theta^* \in S(p,\Theta_I)$, the gradients of the equality constraints are linearly independent and there exists $t \in \mathbb{R}^d$ s.t. $D_j(\theta^*)t<0$ for $j \in \mathcal{J}_1^*(\theta^*)$ and $D_j(\theta^*)t=0$ for $j \in \{J_1+1,\dots,J\}$.
assumption{Abadie Constraint Qualification (ACQ)} For each support point $\theta^* \in S(p,\Theta_I)$, we have $\mathcal{L}(\theta^*)=\mathcal{T}(\theta^*)$.

These conditions are frequently invoked in the statistical literature. For example, Shapiro90,Shapiro1991aAOR,Shapiro93 uses either LICQ or uniqueness of Lagrange multipliers, and these two assumptions are essentially the same Wachsmuth. In econometrics, ChoRussell and Kaido:Santos use LICQ; andrews_roth_pakes and Gafarov restrict attention to linear constraints and thereby impose ACQ.

Results

Restating Some Assumptions, and a First Set of Equivalences

We next restate some of these assumptions, exploiting their localization or using the language of constraint qualifications. Between our “localization" of assumptions and the language of tangent cones, several of the assumptions we just introduced can be stated more succinctly. In what follows, recall that $\mathcal{J}^*(\theta)$, defined in (ref), is the active set of (equality or inequality) constraints at $\theta$. Specifically, define:

customassumption{3'} For each support point $\theta^*\in S(p,\Theta_I)$, there exist constants $\eta>0$ and $C>0$ such that, for all $\theta \in B(\theta^*,\eta)$, $$ \max\bigl\{0,\max_{j=1,\dots,J_1}E(m_j(X,\theta))/\sigma_j(\theta),\max_{j=J_1+1,\dots,J}\vert E(m_j(X,\theta))/\sigma_j(\theta)\vert\bigr\} \geq C \times d(\theta,\Theta_I). $$
customassumption{4'} For each support point $\theta^*\in S(p,\Theta_I)$, there exist constants $\eta>0$ and $C>0$ such that, for all $\theta \in B(\theta^*,\eta) \cap H(p,\Theta_I)$, $$ \max\bigl\{0,\max_{j=1,\dots,J_1}E(m_j(X,\theta))/\sigma_j(\theta),\max_{j=J_1+1,\dots,J}\vert E(m_j(X,\theta))/\sigma_j(\theta)\vert\bigr\}\geq C \times d(\theta,S(p,\Theta_I)) . $$
customassumption{5'} For each support point $\theta^*\in S(p,\Theta_I)$, $\max\{p't/\Vert t \Vert: t\in \mathcal{T}(\theta^*) \setminus \{0\}\}<0$.
customassumption{6'} There are no equality constraints. Furthermore, for each support point $\theta^*\in S(p,\Theta_I)$, $\min_{t \in \mathbb{R}^d} \max_{j \in \mathcal{J}^*(\theta^*)}D_j(\theta^*)t/\Vert t \Vert <0$.
customassumption{7'} For each support point $\theta^*\in S(p,\Theta_I)$, $\max\{p't/\Vert t \Vert: t\in \mathcal{L}(\theta^*) \setminus \{0\}\}<0$.

Then we have:

lemmaSuppose that Assumption (ref) holds. Then the following assumptions are equivalent: (ref)$\Leftrightarrow$(ref), (ref)$\Leftrightarrow$(ref), (ref)$\Leftrightarrow$(ref), (ref)$\Leftrightarrow$(ref), and (ref)$\Leftrightarrow$(ref).
proofRegarding Asumptions (ref) and (ref), $\Leftarrow$ is obvious and $\Rightarrow$ holds because in Assumption (ref), one can choose $\eta=c$, ensuring $\min\{d(\theta,\Theta_I),c\}=d(\theta,\Theta_I)$. The argument for (ref)$\Leftrightarrow$(ref) is the same. Next, (ref)$\Rightarrow$(ref) holds because $T(\theta^*,\eta)$ shrinks toward $\mathcal{T}(\theta^*)\setminus \{0\}$ as $\eta \to 0$. Also, suppose $\mathcal{T}(\theta^*)\setminus \{0\}=\emptyset$, then Assumption (ref) holds vacuously; but in this case, $T(\theta^*,\eta)=\emptyset$ for small enough $\eta$ and so Assumption (ref) holds as well. It remains to show that, if $\mathcal{T}(\theta^*)\setminus \{0\}\neq \emptyset$ and therefore $T(\theta^*,\eta)\neq \emptyset$ for any $\eta$, then failure of Assumption (ref) implies failure of Assumption (ref). Now, failure of Assumption (ref) and nonemptiness of $T(\theta^*,\eta)$ jointly imply existence of sequences $(\delta_n,\eta_n)\downarrow (0,0)$ and $\theta_n \in B(\theta^*,\eta_n) \setminus \{\theta^*\}$ s.t. $p'(\theta_n-\theta^*)/\Vert \theta_n-\theta^* \Vert>-\delta_n$. But then any accumulation point $t$ of $(\theta_n-\theta^*)/\Vert \theta_n-\theta^* \Vert$ is in $\mathcal{T}(\theta^*)$ and has $p't \geq 0$, contradicting Assumption (ref). Assumption (ref) obviously implies (ref). To see the reverse implication, suppose (ref) holds, then one can verify Assumption (ref) by choosing $\delta$ to be half the slack of the tightest inactive inequality (or arbitrarily if all inequalities are active). Next, note first that \begin{eqnarray} && \max\bigl\{p't/\Vert t \Vert : t\in \mathcal{L}(\theta^*) \setminus \{0\}\bigr\}<0 \\ &\Longleftrightarrow &\min_{t \in \mathbb{R}^d:\Vert t \Vert=1,p't \geq 0}\max\bigl\{ \max_{j \in \mathcal{J}_1^*(\theta^*)} D_j(\theta^*)t, \max_{j \in \{J_1+1,\dots, J\} } \vert D_j(\theta^*)t\vert \bigr\} > 0. \notag \end{eqnarray} For example, it is easy to see that the above minimum is attained. If it equalled $0$, the vector $t$ attaining it would be in $\mathcal{L}(\theta^*)$, so the first maximum would be at least $0$. The converse argument is similar. The left-hand side of (ref) is Assumption (ref). We next show that its right-hand side is equivalent to Assumption (ref). It is implied by Assumption (ref) because the minimization is over a smaller set. To see the converse, suppose Assumption (ref) fails, then there exist sequences $\delta_n \uparrow 0$, $\varepsilon_n \downarrow 0$, and $t_n$ with $p't_n \geq \delta_n$ and $\max\bigl\{ \max_{j \in \mathcal{J}_1^*(\theta^*)} D_j(\theta^*)t_n, \max_{j \in \{J_1+1,\dots, J\} } \vert D_j(\theta^*)t_n\vert \bigr\} \leq \varepsilon_n$. Any accumulation point of $t_n$ then is a counterexample to the right-hand side of (ref).
remarkSome of these equivalences are due to localization of assumptions. For example, the original Assumption 3 in PPHI, but also both polynomial minorant conditions, are otherwise stronger than their simplifications. Indeed, regarding the equivalences (ref)$\Leftrightarrow$(ref) and (ref)$\Leftrightarrow$(ref), the real insight is that the original assumptions combine local and global conditions, e.g. (for Assumption (ref)) \begin{eqnarray} && d(\theta,\Theta_I) \leq \delta \\ &\implies & \max\bigl\{0,\max_{j=1,\dots,J_1}E(m_j(X,\theta))/\sigma_j(\theta),\max_{j=J_1+1,\dots,J}\vert E(m_j(X,\theta))/\sigma_j(\theta)\vert\bigr\} \geq C \times d(\theta,\Theta_I) \notag \end{eqnarray} and \begin{eqnarray} && d(\theta,\Theta_I) \geq \delta \\ &\implies & \max\bigl\{0,\max_{j=1,\dots,J_1}E(m_j(X,\theta))/\sigma_j(\theta),\max_{j=J_1+1,\dots,J}\vert E(m_j(X,\theta))/\sigma_j(\theta)\vert\bigr\} \geq \varepsilon \notag \end{eqnarray} for some $\varepsilon>0$ (namely, setting $\varepsilon=C\delta$). Only the local condition (ref) is related to constraint qualifications. The other part is really a global identification condition. We will revisit this distinction later.

Econometric Assumptions as Constraint Qualifications

We now present our main insight: Several of the above assumptions are equivalent, or very close to, constraint qualification assumptions, and there are numerous logical relations between them. Our main result, which is also visualized in Figure (ref), follows.

theoremSuppose Assumption (ref) holds. Then the following relations between assumptions hold true. \begin{enumerate} • Assumptions (ref) and (ref) are equivalent. Furthermore, both are equivalent to jointly (i) excluding equality restrictions and (ii) imposing (ref). • Any of Assumptions (ref), (ref), and (ref) (the latter combined with excluding equality constraints) imply (ref). • Assumption (ref) implies (ref). • Assumptions (ref) and (ref) jointly imply (ref). • Assumption (ref) implies (ref). • Assumption (ref) implies that the gradients of active constraints span $\mathbb{R}^d$. In particular, if $\mathcal{J}^*(\theta^*)$ has exactly $d$ elements, (ref) is implied. • Assumption (ref) implies (ref). \end{enumerate}
table[table omitted — 1,140 chars of source]
figure[figure omitted — 2,131 chars of source]
proofThroughout this proof, consider a fixed $\theta^*$. We invoke Lemma (ref) and use the simplified versions of the assumptions. \paragraph{1.} If equalities are excluded, MFCQ reduces to \begin{equation} \min_{t \in \mathbb{R}^d: \Vert t \Vert =1} \max_{j \in \mathcal{J}^*(\theta^*)} D_j(\theta^*)t < 0. \end{equation} This immediately clarifies that excluding inequalities and imposing MFCQ is just Assumption (ref). It remains to show equivalence with Assumption (ref). Assumption (ref) excludes equalities because the presence of even one equality constraint implies that $\max_{j=1,\dots,J} E(m_j(X,\theta))/\sigma_j(\theta)=0$ for all $\theta \in \Theta_I$. This leaves two possibilities: Either $\Theta_I^{-\varepsilon} \neq \emptyset$, in which case (ref) fails, or $\Theta_I^{-\varepsilon} = \emptyset$, in which case $\eqref{eq:CHT-deg-2}$ fails because $d(\theta,\Theta_I^{-\varepsilon})=\infty$. To see that Assumption (ref) also implies (ref), consider a sequence $\varepsilon_n \to 0$. For $n$ large enough we have $\varepsilon_n /M \leq \delta$, where $M$ and $\delta$ are from Assumption (ref). Then by (ref) there exists $\theta_n \in \Theta_I^{-\varepsilon_n/M}$ with $\Vert \theta_n-\theta^* \Vert \leq \varepsilon_n$, and by (ref), we have $\max_{j \in \mathcal{J}^*(\theta^*)} \{E(m_j(X,\theta_n))/\sigma_j(\theta_n)\} \leq -C \varepsilon_n/M$. Next, let $t$ be any accumulation point of $(\theta_n-\theta^*) / \Vert \theta_n-\theta^* \Vert$, then by continuous differentiability one has $\max_{j \in \mathcal{J}^*(\theta^*)} D_j(\theta^*)t \leq -C/M<0$. To see the converse, let \begin{equation} t^* = \arg\min_{t \in \mathbb{R}^d: \Vert t \Vert =1} \max_{j \in \mathcal{J}^*(\theta^*)} D_j(\theta^*)t \end{equation} and let $\mu=\tfrac{1}{2}\vert \max_{j \in \mathcal{J}^*(\theta^*)} D_j(\theta^*)t^* \vert$. By (ref), $\max_{j \in \mathcal{J}^*(\theta^*)} D_j(\theta^*)t^*=-2\mu<0$. We next argue why inactive constraints, i.e. $j \notin \mathcal{J}^*(\theta^*)$, can be ignored in what follows. Note first that by the Mean Value Theorem and Assumption (ref), there exists $\tilde{\theta}$ componentwise between $\theta$ and $\theta^*$ s.t. \begin{multline*} \frac{E(m_j(X,\theta))}{\sigma_j(\theta)}-\frac{E(m_j(X,\theta^*))}{\sigma_j(\theta^*)} = D_j(\tilde{\theta}) ( \theta-\theta^* ) \\ = (D_j(\theta^*)+O(\Vert \theta-\theta^* \Vert)) ( \theta-\theta^* ) = O(\Vert \theta-\theta^* \Vert) + O(\Vert \theta-\theta^* \Vert^2). \end{multline*} Let $\rho \equiv \max_{j \notin \mathcal{J}^*(\theta^*)}E(m_j(X,\theta^*))/\sigma_j(\theta^*)<0$, then it follows that $$ \max_{\theta \in B(\theta^*,\eta)} \max_{j \notin \mathcal{J}^*(\theta)}E(m_j(X,\theta))/\sigma_j(\theta) \leq \rho+\max_{\theta \in B(\theta^*,\eta)} \left\{ \frac{E(m_j(X,\theta))}{\sigma_j(\theta)}-\frac{E(m_j(X,\theta^*))}{\sigma_j(\theta^*)}\right\}<0$$ for a small enough choice of $\eta$. Thus, inactive constraints do not affect the value taken by $\max_{j=1,\dots,J} E(m_j(X,\theta))/\sigma_j(\theta)$ anywhere on $B(\theta^*,\eta)$. (In the remainder of this proof, $\eta$ is understood to be either the $\eta$ just chosen or the $\eta$ from Assumption (ref), whichever is smaller.) Fix $\delta>0$ and consider any $\theta \in B(\theta^*,\eta)$ with $\max_{j \in \mathcal{J}^*(\theta^*)} E(m_j(X,\theta))/\sigma_j(\theta) > -\delta$. Then \begin{eqnarray*} && \max_{j \in \mathcal{J}^*(\theta^*)} E(m_j(X,\theta-\delta t^*/\mu))/\sigma_j(\theta-\delta t^*/\mu) \\ & = & \max_{j \in \mathcal{J}^*(\theta^*)}\left\{ E(m_j(X,\theta))/\sigma_j(\theta) -\delta D_j(\bar{\theta}_j) t^*/\mu\right\} \\ & = & \max_{j \in \mathcal{J}^*(\theta^*)}\left\{ E(m_j(X,\theta))/\sigma_j(\theta) -\delta D_j(\theta^*) t^*/\mu - o(\delta)\right\} \\ &\geq & -\delta + 2\delta - o(\delta) > 0 \end{eqnarray*} for $\delta$ small enough, implying that $\theta-2\delta t^*/\mu \notin \Theta_I$ and therefore that $d(\theta,\Theta \setminus \Theta_I) < 2\delta/\mu$. (Here, $\bar{\theta}_j$ is componentwise between $\theta$ and $\theta-\delta t^*/\mu$ and may change with $j$; $\Theta_I \subset \operatorname{int}\Theta$ ensures $\theta-2\delta t^*/\mu \in \Theta$ for $\delta$ small enough.) Conversely, by setting $\varepsilon=2\delta/\mu$, we find that (for $\varepsilon$ small enough) \begin{equation*} d(\theta,\Theta \setminus \Theta_I) \geq \varepsilon \implies \max_{j \in \mathcal{J}^*(\theta^*)} E(m_j(X,\theta))/\sigma_j(\theta) \leq -\mu\varepsilon/2, \end{equation*} verifying (ref) with $C=\mu/2$. The requirement that $\varepsilon$ be small enough can be enforced by choosing $\eta$ low enough. Next, for any $\theta \in \Theta_I \cap B(\theta^*,\eta)$, we similarly have \begin{eqnarray*} && \max_{j \in \mathcal{J}^*(\theta^*)} E(m_j(X,\theta+\delta t^*/\mu))/\sigma_j(\theta+\delta t^*/\mu) \\ &=& \max_{j \in \mathcal{J}^*(\theta^*)} \left\{ E(m_j(X,\theta))/\sigma_j(\theta) + \delta D_j(\bar{\theta}_j)t^*/\mu \right\} \\ &= & \max_{j \in \mathcal{J}^*(\theta^*)} \left\{ E(m_j(X,\theta))/\sigma_j(\theta) + \delta D_j(\theta^*)t^*/\mu + o(\delta) \right\} \\ & \leq & 0-2\delta + o(\delta) \end{eqnarray*} for $\delta$ small enough. (Again, $\theta+\delta t^*/\mu \in \Theta$ because $\Theta_I \subset \operatorname{int}\Theta$. The interpretation, though not the value taken, of $\bar{\theta}_j$ is as before.) Now, set $\overline{M}= \max_{j \in \mathcal{J}^*(\theta^*)} \Vert D_j(\theta^*) \Vert$ and let $t$ be any unit vector, then \begin{eqnarray*} && \max_{j \in \mathcal{J}^*(\theta^*)} E(m_j(X,\theta+\delta t^*/\mu+\delta t/\overline{M}))/\sigma_j(\theta+\delta t^*/\mu+\delta t/\overline{M}) \\ &=& \max_{j \in \mathcal{J}^*(\theta^*)}\left\{ E(m_j(X,\theta+\delta t^*/\mu))/\sigma_j(\theta+\delta t^*/\mu) + \delta D_j(\bar{\theta}_j)t/\overline{M} \right\} \\ &\leq & -2\delta +\delta + o(\delta) <0 \end{eqnarray*} for $\delta$ small enough, thus \begin{equation*} B(\theta+\delta t^*/\mu,\delta/\overline{M}) \subseteq \Theta_I \implies \theta+\delta t^*/\mu \in \Theta_I^{-\delta/\overline{M}} \implies d(\theta,\Theta_I^{-\delta/\overline{M}}) \leq \delta /\mu. \end{equation*} Setting $\varepsilon=\delta/\overline{M}$, we find $d(\theta,\Theta_I^{-\varepsilon}) \leq \varepsilon \overline{M}/\mu$, verifying (ref) with $M=\overline{M}/\mu$. \paragraph{2.} Suppose that (ref) applies and let $\mu$ and $t^*$ be as in (ref). Fix any scalar \begin{equation*} \gamma \in \left(0,\frac{1}{\mu}\max_{\theta \in B(\theta^*,\eta)} \max_{j \in \mathcal{J}^*(\theta^*)} E(m_j(X,\theta))/\sigma_j(\theta) \right], \end{equation*} noting that by Assumption (ref)(d), the upper bound on $\gamma$ vanishes as $\eta \to 0$. Consider any $\theta \in B(\theta^*,\eta)$ s.t. $\max_{j \in \mathcal{J}^*(\theta^*)} E(m_j(X,\theta))/\sigma_j(\theta)<\mu\gamma$. Then by a use of the mean value theorem very similar to preceding displays, $$\max_{j \in \mathcal{J}^*(\theta^*)} E(m_j(X,\theta+\gamma t^*))/\sigma_j(\theta+\gamma t^*)<\mu\gamma-2\mu\gamma+o(\gamma)<0$$ for $\gamma$ small enough (which can be ensured by choosing $\eta$ small enough). It follows that $\theta+\gamma t^* \in \operatorname{int}\Theta_I$, hence $d(\theta,\Theta_I)<\gamma$. Conversely, $d(\theta,\Theta_I)=\gamma$ then implies $\max_{j \in \mathcal{J}^*(\theta^*)} E(m_j(X,\theta))/\sigma_j(\theta)\geq \mu\gamma$ and therefore $\max_j E(m_j(X,\theta))/\sigma_j(\theta)\geq \mu\gamma$. As $\gamma$ was arbitrary, this verifies Assumption (ref) with $C=\mu$. \paragraph{3.} Fix any $t \in \mathcal{L}(\theta^*)$. By differentiability of $E(m_j(X,\cdot))/\sigma_j(\cdot)$ and the definition of $\mathcal{L}(\cdot)$, we then have \begin{equation*} \lim \sup \bigl\{ n \times \max\bigl\{ \max_{j \in \mathcal{J}_1^*(\theta^*)} D_j(\theta^*)t/n, \max_{j \in \{J_1+1,...J\} } \vert D_j(\theta^*)t/n\vert \bigr\}\bigr\} = 0 as n \to \infty. \end{equation*} But then Assumption (ref) implies $n \times d(\theta^*+t/n,\Theta_I) \to 0$. Next, let $\theta_n=\arg \min_{\theta \in \Theta_I}\Vert \theta - (\theta^*+t/n) \Vert$ and therefore $\Vert \theta_n-(\theta^*+t/n) \Vert=d(\theta^*+t/n,\Theta_I)$, then \begin{multline*} \lim \frac{\theta_n-\theta^*}{\Vert\theta_n-\theta^* \Vert}= \lim \frac{\theta_n-(\theta^*+t/n)+(\theta^*+t/n)-\theta^*}{\Vert \theta_n-(\theta^*+t/n)+(\theta^*+t/n)-\theta^* \Vert} \\ = \lim \frac{n \times (\theta_n-(\theta^*+t/n))+t}{\Vert n \times (\theta_n-(\theta^*+t/n))+t \Vert}=\frac{t}{\Vert t \Vert}, \end{multline*} hence $t \in \mathcal{T}(\theta^*)$. \paragraph{4.} Under ACQ, $\mathcal{L}(\theta^*)=\mathcal{T}(\theta^*)$, hence Assumptions (ref) and (ref) are then equivalent. \paragraph{5.} This holds because $\mathcal{T}(\theta^*)\subseteq \mathcal{L}(\theta^*)$. \paragraph{6.} Suppose the conclusion fails, thus no $d$ gradients of active constraints span $\mathbb{R}^d$. Then there exists a unit vector $t$ s.t. $D_j(\theta^*)t=0$ for all $j \in \mathcal{J}^*(\theta^*)$. This implies $$\min_{t \in \mathbb{R}^d:\Vert t \Vert=1,p't \geq 0} \max_{j \in \mathcal{J}^*(\theta^*)} D_j(\theta^*)t = 0$$ because at least one of $(t,-t)$ is feasible in this minimization problem, contradicting Assumption (ref); compare in particular the equivalent representation of this assumption on the right-hand side of (ref). \paragraph{7.} Because $p't=0 \Leftrightarrow \theta^*+t \in H(p,\Theta_I)$, the right-hand side of (ref), and thereby Assumption (ref), can be restated as $$\min_{\theta \in H(p,\Theta_I)} \max\bigl\{ \max_{j \in \mathcal{J}_1^*(\theta^*)} D_j(\theta^*)(\theta-\theta^*)/\Vert\theta-\theta^*\Vert, \max_{j \in \{J_1+1,...J\} } \vert D_j(\theta^*)(\theta-\theta^*)\vert/\Vert\theta-\theta^*\Vert \bigr\} > 0.$$ By continuous differentiability (using arguments very similar to above), this then implies that, for small enough $\eta>0$, \begin{multline*} \max\bigl\{0,\max_{j=1,\dots,J_1}E(m_j(X,\theta))/\sigma_j(\theta),\max_{j=J_1+1,\dots,J}\vert E(m_j(X,\theta))/\sigma_j(\theta)\vert\bigr\} \geq C \Vert \theta-\theta^* \Vert \\ for all \theta \in B(\theta^*,\eta) \cap H(p,\Theta_I), \end{multline*} where $C = \tfrac{1}{2} \min_{t \in \mathbb{R}^d:\Vert t \Vert=1,p't = 0} \max\bigl\{ \max_{j \in \mathcal{J}_1^*(\theta^*)} D_j(\theta^*)t, \max_{j \in \{J_1+1,...J\} } D_j(\theta^*)t \bigr\}>0.$

Tightness of Theorem (ref)

We next clarify that Theorem (ref) is tight, i.e., no implication that is not indicated in Figure (ref) holds true. This subsection can be skipped without loss of continuity.

Let $m_j(X,\theta) = \mu_j(\theta) +X_j$ for functions $\mu_j:\Theta \mapsto \mathbb{R}$ defined below and suppose that all $X_j$ are standard normal; thus, $E(m_j(X,\theta))/\sigma_j(\theta)=\mu_j(\theta)$. Also, $\Theta=[-1,1]^2$. All examples are constructed s.t. the support point of interest is $\theta^*=(0,0)$. Direction of projection is $p=(0,1)$ unless explicitly indicated otherwise. There are no equality constraints, so for the purpose of these examples, “MFCQ and no equalities" is just MFCQ.

\paragraph{Neither Assumption (ref) nor (ref) imply either (ref) or (ref) (ACQ).}

eqnarray*[eqnarray* omitted — 100 chars of source]

We start with this example because several others build on it. Qualitatively, it is the left panel of Figure (ref). The linear cone $\mathcal{L}(\theta^*)$ is spanned by $\{(0,1),(0,-1)\}$, the tangent cone $\mathcal{T}(\theta^*)$ is spanned only by $\{(0,-1)\}$. The one direction, $t=(0,-1)$, that is in $\mathcal{T}(\theta^*)$ is tangential to all constraints. The example violates ACQ (and all stronger assumptions) as well as (ref) but fulfills (ref) as well as (ref).

\paragraph{Assumption (ref) does not imply (ref) (MFCQ).}

eqnarray*[eqnarray* omitted — 128 chars of source]

This example adds a third constraint to the first example; compare the right panel of Figure (ref). Both $\mathcal{L}(\theta^*)$ and $\mathcal{T}(\theta^*)$ are spanned by $(0,-1)$. The example therefore fulfills (ref) (and by implication ACQ) but not MFCQ.

\paragraph{Assumption (ref) (ACQ) does not imply (ref).}

eqnarray*[eqnarray* omitted — 97 chars of source]

Here, $\mathcal{L}(\theta^*)=\mathcal{T}(\theta^*)=\{t:p't\leq 0\}$, so that ACQ is fulfilled. However, (ref) is violated in direction $t=(1,0)$:

eqnarray*[eqnarray* omitted — 233 chars of source]

\paragraph{Assumption (ref) does not imply (ref) (ACQ).}

eqnarray*[eqnarray* omitted — 136 chars of source]

In this example (which is inspired by a well-known counterexample to ACQ), we have that $\mathcal{L}(\theta^*)$ is spanned by $\{(-1,-1),(1,-1)\}$ but $\mathcal{T}(\theta^*)$ is spanned by $(0,-1)$ only.

\paragraph{Assumption (ref) does not imply (ref) or (ref).}

eqnarray*[eqnarray* omitted — 83 chars of source]

In this example, $\mathcal{L}(\theta^*)$ and $\mathcal{T}(\theta^*)$ coincide and are spanned by $\{(-1,0),(1,-1)\}$, contradicting both (ref) and (ref). Assumption (ref), which here only applies if we move in direction $(1,0)$ from $\theta^*$, is fulfilled.

Discussion

Our findings inform a number of clarifying remarks on the existing literature.

itemize• As mentioned above, CHT's polynomial minorant can be disentangled into a local and a global identification condition. The local condition (ref) is a mild strengthening of ACQ and is implied by degeneracy. The global condition (ref) is essentially the weakest additional statement needed to ensure that $\Theta_I$ is a well-separated (if set-valued) minimum of $\max\bigl\{\max_{j=1,\dots,J_1} E(m_j(X,\theta))/\sigma_j(\theta), \max_{j=J_1+1,\dots,J} |E(m_j(X,\theta))/\sigma_j(\theta)|,0\bigr\}$.\footnote{A well-separated minimum requires that for each $\varepsilon>0$, there exists $\delta>0$ s.t. $d(\theta,\Theta_I) \geq \varepsilon$ implies $\max\{\dots\} \geq \delta$. Its role in ensuring “inner consistency" of sample analogs of identified sets is well understood Newey_McFadden1994a.} While the polynomial minorant condition is, therefore, not redundant, an instructive restatement of the assumptions is available. • Regarding assumptions in PPHI, claim 5 of Theorem (ref) owes to our simplification, but in view of the smoothness imposed in their Assumption 7, claim 4 also applies to the original versions. Also, if $J=d$, then the PPHI assumptions imply LICQ. This clarifies relation to recent work by ChoRussell and Gafarov: Both effectively impose LICQ and benefit from this by being able to propose relatively simple inference. However, while stronger than assumptions in CHT, BCS, and certainly KMS, the assumptions exceed those in PPHI only in the sense of excluding “overidentified" support points, i.e. more than $d$ active constraints, and are actually weaker in other respects. • ChoRussell furthermore impose Assumption (ref), i.e. (by Theorem (ref)) ACQ. This is not redundant in their paper because (ref) is imposed for all $\theta \in \Theta_I$ and with universal $C$, i.e. “more uniformly" than LICQ. • Yildiz2012 presents conditions for Hausdorff consistency of $\hat{\Theta}_I$. Some of these are in essence constraint qualifications and can be related to our analysis as follows. For the case of pure inequality constraints, her high-level Assumption 3.1 states that $\Theta_I$ is the closure of the strict level $0$ lower contour set of $\max_j E(m_j(X,\theta))/\sigma_j(\theta)$.\footnote{Molchanov1998 also uses this condition to ensure Hausdorff consistency of set estimators.} For the case of at most $d$ active constraints, this is derived as an implication of the LICQ (Assumption 3.2). For the more general case, it is derived from an assumption (in Lemma 3.1) enforcing that, at any boundary point $\theta^*$ of $\Theta_I$, the criterion function $\max_j E(m_j(X,\theta))/\sigma_j(\theta)$ is strictly increasing in some component of $\theta$. To make it comparable to our assumptions, one would impose this to hold at any support point $\theta^*$. By a minimal extension of step 2 of Theorem (ref) (see expression (ref)), it is then equivalent to Assumption (ref). Therefore, Yildiz2012 essentially imposes MFCQ for the pure inequality case.\footnote{Yildiz2012 writes that her assumptions imply a degeneracy condition in CHT without claiming the reverse implication. This refers to their high-level degeneracy assumption C.3, which is implied whenever the simple sample analog of $\Theta_I$ is consistent.} For the case of mixed equalities and inequalities, she invokes a LICQ (Assumption 4.1(b)). • We close with some remarks on why, for inference on projections $p'\theta$, certain approaches do not require constraint qualifications. Specifically, the profiling approach in BCS gets by with the relatively weak Assumption (ref); KMS or projection of confidence regions for $\theta$ (such as those in AndrewsSoares2010E, Bugni2009E, or Canay2010JE) use no shape restrictions for $\Theta_I$ at all. The reason is that all of these approaches localize inference at a conjectured true value of the support function $s(p,\Theta_I)$ (in BCS) or parameter vector $\theta$ (in all others). Consistent estimation of identified sets for these objects is then not a concern. In particular, BCS need to ensure some form of consistency of a sample analog of $\Theta_I$ that is restricted to the true supporting hyperplane, and this is precisely what the minorant on the support plane achieves. The other approaches need no constraint qualifications at all. Of course, there is no free lunch. All the methods just alluded to are computationally expensive because they effectively invert tests whose critical value depends on the exact value of the parameter under the null. Thus, while some shortcuts may be available in practice (see in particular KMS and KMST_code), critical values must in principle be recomputed at every conceivable value of $\theta$ or $s(p,\Theta)$.

Verifying Assumptions in Some Examples

In this section, we discuss how these restrictions apply to two well-understood examples of partial identification. It will become clear that one cannot take for granted that they “typically" hold and that verifying them could be intricate in more involved examples. That is part of our message: We point out that many of them are basically constraint qualifications, and it is well known that constraint qualifications can be subtle to verify.

The examples are visualized in Figure (ref), which is designed to resemble Figure (ref). Note that in the second example, $\Theta_I$ has zero measure.

figure[figure omitted — 3,034 chars of source]

Linear regression with interval outcome data and discrete regressors

Consider a linear regression model:

align*[align* omitted — 36 chars of source]

where $Z=(Z_1,\dots,Z_d)$ is a $d\times 1$ random vector with $Z_1=1$. We assume that $Z$ has $k$ points of support denoted $\mathcal Z=\{z^1,\dots,z^k \in \mathbb{R}^d\}$ with $\max_{r=1,\dots,k}\Vert z^r \Vert<M<\infty$. The researcher observes $\{W_0,W_1,Z\}$ with $P(W_0\le W \le W_1|Z=z^r)=1,r=1,\dots,k$, where $\pi^r=P(Z=z^r)>0,r=1,\dots,k$ are assumed known. Suppose that $W_0$ and $W_1$ take values in $\mathcal W\subset\mathbb R$. An important special case is missing data, where $W_0$ and $W_1$ both equal the true $W$ if the latter is observed and correspond to some bound on it otherwise.

The identified set is characterized by the following moment inequalities.

align*[align* omitted — 122 chars of source]

Equivalently,

align[align omitted — 178 chars of source]

Define the following objects.

align*[align* omitted — 500 chars of source]

Note that, since $D_j$ does not depend on $\theta$, we drop its argument.

By (ref)-(ref), the identified set is a polytope characterized by $k$ pairs of parallel constraints

align[align omitted — 128 chars of source]

See the left panel of Figure (ref) for an illustration with $k=3$. Each support point is either a vertex of the polytope or a point in the relative interior of a non-singleton support set. We therefore analyze subcases below. Note that, since the gradients $z^1,\dots,z^k$ of the constraints in (ref) are known, one knows without data which case applies.

We first establish conditions implying MFCQ. For $j=1,\dots,2k$, define:

align*[align* omitted — 101 chars of source]

We call $H_j$ a hyperplane and $H_j^-$ a half space. Let $\operatorname{ri}(A)$ denote the relative interior of $A$. A $(d-1)$-dimensional flat face in $\mathbb R^d$ is called a facet. Similarly, a $\ell$-dimensional element of a $d$-dimensional polytope is called a $\ell$-face, where $1\le\ell\le d-2$. For example, a $1$-face is an edge in a 3 dimensional polytope.

lemmaSuppose (i) $\mathcal W$ is compact; (ii) $\Theta=\{\theta\in\mathbb R^d:\|\theta\|^2\le B_0\}$ with $B_0<\infty$ satisfying $C_0B_0>k\sup_{w\in\mathcal W}w^2$, where $C_0\equiv \inf_{p\in\mathbb S^{d-1}}\sum_r (p'z^r)^2$; (iii) $k\ge d$, and any subset $\mathcal C\subseteq \mathcal Z$ with $\#\mathcal C \le d$ is linearly independent; (iv) $E[W_1-W_0|Z=z^r]>0$ for all $r=1,\dots,k.$ Let $\theta^*\in S(p,\Theta_I)$. \begin{enumerate} • (Facet) If $S(p,\Theta_I)$ is not a singleton and is the intersection of a hyperplane and a finite number of half spaces, MFCQ holds at any $\theta^*\in \text{ri}(S(p,\Theta_I))$; • ($\ell$-face) If $S(p,\Theta_I)$ is not a singleton and is the intersection of more than one hyperplanes, MFCQ holds at any $\theta^*\in \text{ri}(S(p,\Theta_I))$; • (Vertex) If $S(p,\Theta_I)$ is a singleton and $\#\mathcal J^*(\theta^*)\le d$, MFCQ holds at $\theta^*$. \end{enumerate}
remarkThe lemma shows that, even if one is not sure about whether $\theta^*$ is on a facet or any other lower dimensional elements of the polytope, MFCQ holds as long as $\#\mathcal J^*(\theta^*)\le d$ (and other conditions of the lemma hold). Also, things simplify if $k=d$, in which case conditions (iii) and (iv) imply $\#{\mathcal J}^*(\theta^*)\le d$ at any support point. Condition (iii) then ensures LICQ, and therefore MFCQ, at $\theta^*$.
proofUnder our assumptions, $\Theta_I$ is nonempty and is in the interior of $\Theta$ by Proposition F.1 in Kaido:Santos. Case 1: Suppose $S(p,\Theta_I)$ is a flat face and $\theta^*\in \operatorname{ri}(S(p,\Theta_I))$. This occurs if and only if $p=cz^r$ (or $p=-cz^r$) for some $r\in\{1,\dots, k\}$ and $c>0$, and $p'\theta^*=E(W_11\{Z=z^r\})/\pi^r$ (or $E(W_01\{Z=z^r\})/\pi^r$). By (iii), such $r$ is unique. Without loss of generality, suppose $p'\theta^*=E(W_11\{Z=z^r\})/\pi^r$. The assumption $E[W_1-W_0|Z=z^r]>0$ ensures \begin{multline} E(W_11\{Z=z^r\})/\pi^r-E(W_01\{Z=z^r\})/\pi^r \\ =E[E[W_1-W_0|Z=z^r]1\{Z=z^r\}]/\pi^r>0. \end{multline} Hence, $p'\theta^*>E(W_0 1\{Z=z^r\})/\pi^r$, implying the lower bound is slack. This and $\theta^*\in \operatorname{ri}(S(p,\Theta_I))$ imply $\mathcal J^*(\theta^*)=\{j^*\}$ with $j^*=r+k$. The gradient of the normalized moment is \begin{align*} D_{j^*}=\frac{z^{r\prime}}{\sigma_{j^*}}=\frac{z^{r\prime}}{\operatorname{Var}(W_11\{Z=z^r\})^{1/2}/\pi^r}. \end{align*} Let $t^*=-z^r$, then \begin{align*} D_{j^*}t^*=-\frac{z^{r\prime}z^r}{\operatorname{Var}(W_11\{Z=z^r\})^{1/2}/\pi^r}<0, \end{align*} which establishes MFCQ. Case 2: We first claim that $\#\mathcal J^*(\theta^*)<d$. Suppose otherwise. Then, there are at least $d$ active inequalities at $\theta^*$. Select a subset $\tilde{\mathcal J}\subset \mathcal J(\theta^*)$ such that $\#\tilde{\mathcal J}=d$. Then, by condition (iii), $\{D_j,j\in\tilde J\}$ are linearly independent, implying $\theta^*$ is a unique solution to the system of linear equations \begin{align*} D_j\theta=b_j, \forall j\in\tilde{\mathcal J}. \end{align*} Furthermore, by the necessary condition of the maximization problem, there is $\lambda\in\mathbb R^{2k}_+$ such that \begin{align} p+\sum_{j\in\tilde{\mathcal J}}\lambda_j D_j'=0, \end{align} where $\lambda_j>0$ for all $j\in\tilde{\mathcal J}$. Let $\tilde\theta\in S(p,\Theta_I)$ and $\tilde\theta\ne\theta^*$. By construction, $p'(\tilde\theta-\theta^*)=0$. This and (ref) imply \begin{align} \sum_{j\in\tilde{\mathcal J}}\lambda_j D_j(\tilde\theta-\theta^*)=0. \end{align} Suppose that $D_j(\tilde\theta-\theta^*)>0$ for some $j\in\tilde{\mathcal J}$. This implies \begin{align*} D_j\tilde\theta>b_j, \end{align*} hence the $j$-th constraint is violated, hence $\tilde\theta$ cannot be in $S(p,\Theta_I)$, a contradiction. Similarly, suppose that $D_j(\tilde\theta-\theta^*)<0$ for some $j\in\tilde{\mathcal J}$, then by (ref) and the positivity of the Lagrange multipliers, there must exist $j'\in\tilde{\mathcal J}$ such that $D_{j'}(\tilde\theta-\theta^*)>0$, and the same argument applies. The only remaining possibility is $D_j(\tilde\theta-\theta^*)=0$ for all $j\in\tilde{\mathcal J}$, in which case \begin{align*} D_j\tilde\theta=b_j, \forall j\in\tilde{\mathcal J}, \end{align*} which contradicts $\theta^*$ being the unique solution to (ref). Therefore, $\#\mathcal J^*(\theta^*)<d$ must hold. Now, by condition (iii) and $\#\mathcal J^*(\theta^*)<d$, LICQ holds at $\theta^*$, which implies the claim. Case 3: By $\#\mathcal J^*(\theta^*)\le d$ and condition (iii), LICQ holds at $\theta^*$, which implies the claim.

We conclude by mentioning other conditions. The identified set $\Theta_I$ is a finite polytope, hence any projection is the value of a linear program. This immediately clarifies that ACQ obtains, which furthermore means that Assumptions (ref) and (ref) coincide. Both restrict $\mathcal{L}(\theta^*)$, which locally just coincides with $\Theta_I$, to not intersect the supporting hyperplane other than at $\theta^*$. This will hold iff the support set is a singleton, i.e. a vertex. Finally, the solution to a linear program is necessarily well-separated, so that Assumption (ref) holds.

A Simple Entry Game

Two-player entry games are a workhorse example of partial identification since Tamer03. We analyze the game specified in BCS but without covariates. The equilibrium concept is pure strategy Nash equilibrium (PSNE). Firm $1$ respectively $2$ enters if

eqnarray*[eqnarray* omitted — 100 chars of source]

where $(A_1,A_2)\in\{0,1\}^2$ are the firms' actions, $(\theta_1,\theta_2)\in [0,1]^2$ are the interaction parameters, and $(\varepsilon_1,\varepsilon_2)$ are observed by the players but not by the researcher. This system is incomplete: For certain realization of $(\varepsilon_1,\varepsilon_2)$, both $(1,0)$ and $(0,1)$ are PSNE and the model is silent on which is played. Hence, different (stochastic) selection mechanisms picking an equilibrium in the region of multiplicity can be coupled with different values of $\theta$ to yield the same distribution of observables $(A_1,A_2)$. Nonetheless, inference can be carried out by bounding the likelihood of different outcomes from below by the respective probabilities of them being unique PSNE. In this simple example, this exhausts all the information in the model and data, leading to a sharp identification region. See BMM_ECMA for characterizations of sharp identification regions in more complex models.

In this model,

itemize$(1,1)$ is the unique PSNE if $\varepsilon_1-\theta_1 A_2 \geq 0$ and $\varepsilon_2-\theta_2 A_1 \geq 0$, • $(1,0)$ is the unique PSNE if $\varepsilon_1-\theta_1 A_2 \geq 0$ and $\varepsilon_2-\theta_2 A_1 \leq 0$, • $(0,1)$ is the unique PSNE if $\varepsilon_1-\theta_1 A_2 \leq 0$ and $\varepsilon_2-\theta_2 A_1 \geq 0$.

Letting $\pi_{jk}=\Pr(A_1=j,A_2=k)$ and assuming (as in BCS) that $(\varepsilon_1,\varepsilon_2)$ are distributed independently uniformly on $[0,1]$, we have

eqnarray[eqnarray omitted — 206 chars of source]

Geometrically, the identified set $\Theta_I$ is an arc segment. The right panel of Figure (ref) depicts its true shape if $\pi_{01}=\pi_{10}=\pi_{11}=1/3$.

Consider now the problem of bounding $\theta_2$ from above; all other bounds on individual parameters are similar. Guessing that (ref)-(ref) will bind at the solution, one can easily solve for

equation*[equation* omitted — 177 chars of source]

Furthermore, the linear and tangent cone are tangent to (ref) and are spanned by $(\pi_{10}+\pi_{11},-\pi_{11}/(\pi_{10}+\pi_{11}))$. This vector has strictly negative inner product with $p=(0,1)$. We conclude:

itemize• Assumption (ref) and any equivalent assumption cannot hold because $\Theta_I$ has no interior. • All of Assumptions (ref), (ref), (ref), (ref), LICQ, and ACQ hold.

However, fulfillment of some of these assumptions is delicate in that it depends on $\theta^*$ being an exposed point of $\Theta_I$. Other directions of optimization (e.g., $p=(1,-1)$, corresponding to testing the null that $\theta_1=\theta_2$) have $\theta^*$ in the relative interior of $\Theta_I$, i.e. at a point where only (ref) is active. Assumptions (ref), (ref), and (ref) will then fail.

Conclusion

The literature on partial identification uses constraint qualifications in many ways: To ensure Hausdorff consistency or rates of convergence for simple estimators of identified sets Yildiz2012, to justify inference for the full parameter vector $\theta$ (CHT) or subvectors ChoRussell,Gafarov, or to justify efficiency bounds Kaido:Santos. However, some of these uses are implicit, making it difficult even for expert readers to compare assumptions. We provide a guide to how different high-level assumptions relate to each other and to well-known constraint qualifications. A simple, important message is that several high-level assumptions are tightly related to the Mangasarian-Fromowitz constraint qualification and are essentially mutually equivalent. We believe that this provides useful guidance to readers trying to make sense of the large menu of inference methods for partially identified vectors and subvectors CS17,MolinariHOE. For example, it clarifies costs and benefits relative to work that has weak-to-no geometric regularization, mostly at the expense of computational effort.