Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
141,229 characters · 25 sections · 113 citation commands
Linear programming approach to partially identified econometric models
\doparttoc \faketableofcontents \parttoc
In many partially identified models, the sharp bounds on parameters correspond to the values of linear programs (LPs) that depend on identified functionals of the underlying probability measure. Examples include conditional moment inequalities andrews2023, generalized IV models santos, revealed preference restrictions klinetartari, intersection bounds honoreadriana2006, dynamic discrete choice panels honoretamer2006 and shape restrictions MP2000.
In these settings, the bounds take the form $B(\mathbb{P}) = B(\theta_0(\mathbb{P}))$, where
and \( \theta_0(\mathbb{P}) \) is the true value of parameter $\theta = (p', \text{vec}(M)', c')'$, estimated via a \( \sqrt{n} \)-consistent estimator \( \hat{\theta}_n \). However, optimization problem (ref) exhibits non-regular behavior, particularly when the underlying model is rich enough that some linear functionals of \( x \) are nearly or exactly point-identified over \( \Theta_I = \{x \in \mathbb{R}^d:Mx \geq c\}\). In such cases, existing estimators of the LP value $B(\mathbb{P})$ are either inconsistent or rate-conservative, creating an undesirable tradeoff in empirical work: richer models provide tighter bounds but complicate their estimation.
To address this issue, we develop a novel debiased penalty function estimator of $B(\mathbb{P})$. Only assuming that the true polytope $\Theta_I$ is non-empty and contained in a known compact set, we show that our estimator is $\sqrt{n}-$consistent, pointwise in the probability measure. In contrast, the plug-in estimator is not generally consistent and may fail to exist with non-vanishing probability, while the alternative set-expansion estimator based on CHT is rate-conservative and may fail to exist in finite samples. Figure (ref) gives a preview of the comparative performance of these estimators.
We obtain an asymptotically normal version of our estimator via sample-splitting and construct confidence regions with exact asymptotic coverage at any fixed probability measure. By comparison, existing procedures either rely on further conditions gafarov2024simple, or result in asymptotically conservative inference chorussel2023. Notably, the approach most commonly used in applied research—combining plug-in estimation with bootstrap (de2017effect, mentalhealth, siddique, kreider2012identifying, pepper, blundell2007changes)—may not provide valid confidence intervals even when the underlying model is far from point-identification. In addition, our procedure is the most computationally efficient in the existing literature\footnote{Both gafarov2024simple (BG) and chorussel2023 (CR) rely on resampling methods, which require to compute one or multiple LPs at each bootstrap iteration. Computing a confidence interval for a LP with 32 variables takes 16.81 seconds with the approach of BG and 40.65 seconds with the approach of CR, according to the latter work. Our approach requires solving a LP once, which takes around 0.0022 seconds on average. The LPs in our application have 160 variables.}.
Turning to uniform asymptotic theory, we first establish a general impossibility result: using Le Cam's binary testing method, we show that no uniformly consistent estimator exists when the estimated functional is discontinuous in the total variation norm. This result implies that the LP value cannot be uniformly consistently estimated over the unrestricted set of distributions $\mathcal{P}$. To make progress, we introduce the `$\delta-$condition' that parametrizes $\mathcal{P}$ by restricting it to the measures at which the smallest singular value of some full-rank submatrix of constraints binding at an optimal vertex is lower-bounded by a $\delta > 0$. This condition is minimal in the sense that any measure from $\mathcal{P}$ satisfies it for some $\delta$, ensuring the family of restricted measures' sets covers $\mathcal{P}$ as $\delta$ grows small. Unlike the conditions in gafarov2024simple, it does not exclude economically relevant problematic cases, such as point-identification and over-identification, nor does it preclude solution multiplicity. Under the $\delta-$condition, our estimator is shown to be uniformly consistent.
To complement our estimation procedure, we derive sharp (and novel) LP bounds for a broad class of causal parameters under affine inequalities over conditional moments\footnote{Including ATE and CATE, among other typically studied parameters, see Section (ref).} (AICM), potentially augmented with affine almost sure restrictions and missing data conditions. In the simplest case, AICM identifying restrictions have the form
where $(Y(d))_{d}$ are continuous potential outcomes corresponding to the legs of treatment $T$ and $Z \in \mathbb{R}^{d_Z}$ are other covariates. Identified matrices $M^*, \tilde{M}$ and vectors $b^*, \tilde{b}$ are chosen by the researcher. In AICM models, $\theta$ from (ref) is usually a function of identified conditional moments $(\mathbb{E}[Y|T = t,Z =z])_{t,z}$ and the identified joint distribution of $T,Z$, while $x$ collects relevant unobserved conditional moments. Our approach accommodates arbitrary combinations of existing `nonparametric bounds' restrictions, allows to conduct sensitivity analysis, and extends to more complex conditions where sharp bounds were previously unavailable\footnote{For example, cMIV and the mixture of all classical MP2000 conditions, see below.}.
Finally, we apply our approach to estimating returns to education in Colombia. To that end, we first introduce a family of conditionally monotone instrumental variables assumptions (cMIV), which are nested within (ref). A variable is a cMIV if the potential outcomes are mean-monotone in it both unconditionally and within selected treatment subgroups\footnote{The collection of treatment subgroups over which monotonicity is assumed is chosen by the researcher, see Section (ref) for details.}. While an explicit form for the sharp bounds under cMIV may not be feasible, their LP representation follows from our general identification result for (ref). The cMIV conditions yield tighter bounds than the classical MIV assumption of MP2000. We argue, however, that they remain unrestrictive in many applications, including ours. While empirical studies (e.g., de2017effect) have assessed the monotonicity of observed conditional moments to justify applying MIV, such monotonicity is instead equivalent to a particular form of cMIV given that MIV holds and under a mild regularity condition. The formal test of cMIV is obtained as an extension of chetverikov_2019. Using Saber test scores as a cMIV, we find that earning a university degree increases average wages by at least $5.5\%$ in Colombia. In contrast, the classical conditions fail to produce an informative bound.
This paper also contributes two auxiliary results. The first one is concerned with an important special case of (ref) - the combination of all classical MP2000 conditions. Since this combination has the strongest identifying power among classical restrictions, it has been used in empirical work even without a formal justification, sometimes leading to incorrect bounds (see laffers). We provide sharp bounds under continuous outcomes in this setting. Another auxiliary contribution is a novel lower bound on the $\ell_1$-deviation from a non-empty bounded polytope in terms of Euclidean distance from the polytope. It may offer insights into the behavior of $\ell_1$-penalized solutions of systems of linear inequalities studied in the control theory literature (e.g. pinar19991).
We briefly note the limitations of our approach. On the identification side, the absence of restrictions on treatment selection prevents us from studying more granular parameters, such as marginal treatment responses. Furthermore, our identification results are given for discrete treatment and instrument. An extension to the continuous case is feasible, but is outside the scope of this paper\footnote{Even when continuous identification results are available, in practice estimation is still carried out with discretized covariates. This is true for all empirical work referenced below.}. On the estimation side, while our estimator is pointwise $\sqrt{n}-$consistent in general, we only establish $\sqrt{n}/w_n$-uniform consistency\footnote{We show, however, that one side of the convergence happens at the uniform rate $\sqrt{n}$, see Section (ref).} for a slowly diverging sequence $w_n$\footnote{Theoretically, $w_n$ can diverge arbitrarily slowly. We use $w_n \propto \ln \ln n$, see Section (ref) and Appendix (ref).}. We provide further evidence on the uniform rate of consistency in Appendix (ref). A theoretically $\sqrt{n}-$uniformly consistent estimator follows from our analysis, but it depends on an unobserved parameter $\delta$ that is difficult to estimate, so we do not recommend using it in practice. Finally, while our inference procedure naturally extends to uniform setup under sufficient regularity conditions, exploring this is left for future work.
The strand of literature relevant to the estimation of (ref) is concerned with statistical inference in the LP estimation framework. semenova2023adaptive considers a LP with an estimated constraint vector $\hat{c}_n$ but known $M$ and $p$, while bhattacharya2009inferring considers LPs with an estimated $\hat{p}_n$ and known $M, c$. Methods developed under a known $M$ assumption do not easily extend to the setting when $M$ is estimated. santos construct a set-expansion estimator and prove its consistency. gafarov2024simple develops uniform inference for a LP described by affine inequalities over unconditional moments, provided uniform Linear Independence Constraint Qualification (LICQ) and Slater's condition (SC) hold. Gafarov's conditions may be restrictive in some applications - for example, under AICM, see Section (ref). chorussel2023 obtain uniformly valid, yet conservative inference for the case when $\theta$ is affine in unconditional moments. Their practical procedure implicitly assumes that the SC holds\footnote{We discuss this in more detail in Section (ref).}. andrews2023 develop an inference procedure for the LP value in a special case in which SC holds. syrgkanis2017inference develop a testing procedure for the failure of LP feasibility. We conduct Monte Carlo simulations to compare our approach with relevant existing methods in Section (ref).
Despite the similar name, AICM approach is unrelated to the model in andrews2023 beyond producing LP bounds. Instead, it generalizes the nonparametric bounds analysis of MP2000, MP2009. The LP sharp bounds in Theorem (ref) coincide with or tighten the bounds in blundell2007changes, boes2009, siddique, kreider2012identifying, de2017effect and mentalhealth. These studies combine the plug-in estimator $B(\hat{\theta}_n)$ with bootstrap for inference, an approach that relies on strong assumptions (see Section (ref)). The exact inference procedure in Algorithm (ref) could be used instead. AICM also complements the findings of santos, who develop identification theory for generalized IV estimators and obtain bounds in the form (ref). Their method imposes a heckman1999local,heckman2005structural selection mechanism in the binary treatment case and thus restricts them to valid IVs\footnote{Additive separability in treatment selection is equivalent to the imbensangrist IV conditions under instrument exogeneity vytlacil2002independence.}, but allows to accommodate arbitrary a.s. restrictions on the marginal treatment response functions and derive bounds for potentially more granular causal parameters. Even though (ref) nests mean-independence conditions, AICM approach is most useful when a valid IV is not available\footnote{Thus, if one is faced with i) a binary treatment setup, ii) has a valid IV and iii) no outcomes' data is missing, the method of santos may be used. If any of these conditions fail, AICM is an alternative.}.
All vectors are column vectors, and $M'$ denotes the transpose of $M \in \mathbb{R}^{n\times m}$. If $A$ is a set, $A'$ stands for its complement. A collection $(x_j)_{j \in J}$ is a column vector. $2^A$ denotes the powerset of set $A$, and $[n]$ is the collection of integers from $1$ to $n \in \mathbb{N}$. $\times$ is a Cartesian product of sets, while $\otimes$ is the Kronecker product. The sign $\sqcup$ denotes a disjoint union. Signs $\land$ and $\lor$ stand for logical `and' and `or' operators respectively. For $M \in \mathbb{R}^{m\times n}$ and $A \subseteq [m]$, $M_A \in \mathbb{R}^{|A| \times n}$ is the submatrix of the rows of $M$ with indices in $A$. If $j \in [m]$, write $M_j \equiv M'_{\{j\}}$. $\mathcal{R}(M)$ stands for the range of $M$, while $\text{rk}{(M)}$ denotes its rank. $\sigma_d(M)$ is the $d-$th largest singular value of $M$, and $M^\dagger$ denotes the Moore-Penrose pseudoinverse of $M$. In a normed space $S$, the distance between $x \in S$ and $A \subseteq S$ is written as $d(x,A) \equiv \inf_{a \in A} ||x - a||$, and $d_H(A,B) \equiv \max\{\sup_{b \in B} d(b, A), \sup_{a \in A} d(a, B)\}$ is the Hausdorff distance between $A, B \subseteq S$. For $A \subseteq S$ the open expansion is $A^\varepsilon \equiv \{s \in S: d(s, A) < \varepsilon\}$. $Int(A)$ and $Cl(A)$ are the interior and closure of $A \subseteq \mathbb{R}^d$, while $Cone(A)$ is its conical hull. If $A$ is a matrix, $Cone(A)$ is the conical hull of its columns. $s(x, A) \equiv {\max}_{{a \in A}} ~ x'a$ for a compact $A \subseteq \mathbb{R}^d$ and $x \in \mathbb{R}^d$ is a support function. For $v = (v_j)_{j \in [d]}$, define $v^+ \equiv (\max \{v_j, 0\})_{j \in [d]}$. For $v, u \in \mathbb{R}^d$ vector inequalities $v > u$ and $v \geq u$ mean $v_i > u_i ~ \forall i \in [d]$ and $v_i \geq u_i ~ \forall i \in [d]$ respectively. $\iota_d \in \mathbb{R}^d$ is a vector of ones, $I_d \in \mathbb{R}^{d\times d}$ is the identity matrix, and the subscript is dropped occasionally. Operator $\mathbb{E}_\mathbb{P}$ is the expectation under a measure $\mathbb{P}$, and the subscript is dropped whenever it does not cause confusion. The statement $w_n \to \infty$ w.p.a.1 means that $\forall M > 0,$ $ \lim _{n\to\infty} \mathbb{P}[w_n > M] = 1$. We adopt the convention $\inf \emptyset = +\infty$, and $\sup \emptyset = - \infty$.
In many partial identification settings, bounds on the parameters of interest can be characterized as the values of linear programs (see review_lp for a review). Readers who prefer to first see an identification framework leading to such bounds may refer to Section (ref), where LP sharp bounds are derived for a general class of AICM models. This section focuses on the estimation theory for such problems. The LP value function is given by
where $M \in \mathbb{R}^{q\times d}$, $c \in \mathbb{R}^q$ and $p \in \mathbb{R}^d$. The vector $\theta \equiv (p', \text{vec}(M)', c')'$ collects parameters of the LP. The estimable value of these parameters at a fixed true measure is denoted by $\theta_0 \in \mathbb{R}^S$, with $S = qd + q + d$. The value of interest is therefore $B(\theta_0)$. Note that (ref) does not rule out equality constraints, as $Ax = b \iff Ax\geq b \land -Ax \geq -b$.
We denote the constraint set by $\Theta_I(\theta) \equiv \{x \in \mathbb{R}^{d}| Mx \geq c \}$ and omit the argument when $\theta_0$ is concerned. In the context of Section (ref) and other existing applications (e.g. santos), the set $\Theta_I$ is the identified set for an unobserved feature $x$ of the underlying distribution. Under Assumption A0.i-ii, the set $\Theta_I$ is a convex polytope.
Assumption A0 is maintained throughout this section, while other conditions are stated explicitly. A0.i ensures feasibility of the population LP. In partially identified models it means that there exists a distribution consistent with the identifying restrictions, which does not imply correct model specification. A0.ii is a mild restriction\footnote{For a polytope to be bounded, it must be that $q \geq d+1$, see Chapters 2 and 3 in grunbaum1967convex } that usually holds in applications, for example under bounded outcomes in AICM models (see Section (ref)). The estimator in A0.iii is typically warranted by CLT and the Delta-Method. We focus on the case in which $\theta_0$ is $\sqrt{n}-$estimable for expositional simplicity, but our results generalize to any rate $r_n \uparrow \infty$.
The following primal and dual solution sets will prove useful in our discussion:
Assumption A0 implies that a finite $B(\theta_0)$ is attained as a minimum in (ref), $\Theta_I(\theta_0)$ and $\mathcal{A}(\theta_0)$ are non-empty compact sets, and $\Uplambda(\theta_0)$ is non-empty. We now briefly discuss the typically imposed regularity conditions.
SC rules out point-identification of any linear functional of $x$. In particular, it precludes exact point-identification of the target $p'x$ and point-identification of $x$, i.e. the case when $|\Theta_I| = 1$. Most existing methods rely on SC explicitly (gafarov2024simple) or implicitly (chorussel2023, andrews2023). Even an `approximate' failure of SC, when the true identified set $\Theta_I$ becomes `thin', can be problematic for the existing methods in finite samples. Our simulation evidence illustrates this, see Section (ref). This creates an undesirable tradeoff in applications: higher identification power comes with poorer estimation quality.
LICQ precludes the existence of overidentifying constraints at the optimum. It may be hard to justify in `bigger' models, like the one we develop and apply in Sections (ref) and (ref). These feature a larger number of inequality constraints that may have similar identifying power, so it is not ex-ante clear why there must not be overidentification at the optimum. LICQ also rules out parameters-on-the-boundary, as the following remark clarifies.
The notion of flat faces thus corresponds to $|\mathcal{A}(\theta_0)| \ne 1$, i.e. the situation in which the bound on the target parameter $p'x$ is achieved at multiple partially identified features $x$.
Assumption A0 does not impose LICQ or SC, nor does it rule out flat faces. Estimating $B(\theta_0)$ without these conditions is challenging due to the irregular behavior of $B(\cdot)$. If SC fails, $B(\cdot)$ may be discontinuous at $\theta_0$, and the plug-in estimator $B(\hat{\theta}_n)$ may not be pointwise consistent.
The proposition above can be illustrated using two simple examples.
In some special cases (e.g., honoretamer2006), SC may be argued to hold, ensuring the continuity of $B(\cdot)$\footnote{Under SC and A0, $B(\cdot)$ is continuous at $\theta_0$, see Appendix (ref).}. However, an additional challenge arises: $B(\cdot)$ is not necessarily Hadamard differentiable unless further regularity conditions hold. This complicates inference, as Proposition (ref) in Section (ref) illustrates.
This section addresses Propositions (ref) and (ref). Section (ref) introduces the penalty function estimator and its debiased version, which we show to be $\sqrt{n}$-pointwise consistent under A0. Section (ref) develops a computationally efficient inference procedure with exact asymptotic coverage under A0. Turning to uniform properties, Section (ref) presents a general impossibility result for discontinuous functionals. It implies that the LP value cannot be uniformly consistently estimated under a uniform version of A0 alone. We then characterize a broad class of measures over which a uniformly consistent estimator exists. Sections (ref) and (ref) establish the uniform rates of the penalty function estimators over this class. Section (ref) provides simulation evidence.
We now develop a consistent estimator that is inspired by the theory of exact penalty functions. The idea is to restate (ref) as an unconstrained penalized problem. Define the $L_1-$penalized version of the LP objective as
and consider the unconstrained problem
We use $\tilde{B}(\cdot)$ to obtain a preliminary estimator of the LP value, which we term the penalty function estimator. Note that $L(x; \theta, w) = p'x$ at any $x \in \Theta_I(\theta)$, i.e. the penalized objective function is equal to the LP objective function whenever the constraints in (ref) are satisfied.
Note that Assumption A1 does not require $w$ to be component-wise larger than all KKT vectors. In any finite and feasible LP there exists at least one $\lambda^* < \infty$, so at any fixed $\theta_0$ there always exists a large enough $w$ that satisfies A1.
The following Lemma is key to understanding the penalty function approach. It asserts that under Assumption A1 the $L_1-$penalty function is exact for the LP in (ref).
The deterministic result in Lemma (ref), combined with the observation that the objective function converges in probability uniformly in $x$ under A0, establish that the penalty function estimator with a fixed $w$ is consistent under A1.
For ease of notation, from now on we treat $w \in \mathbb{R}_+$ as a scalar penalty that induces the penalty vector $w \iota$. Our results extend immediately to the case when the coordinates of the penalty vector differ. Based on Lemma (ref) and Proposition (ref), it might seem that $w$ should be selected to be as large as possible. This, however, yields a generally inconsistent estimator if SC fails.
\addtocounter{excont}{-1} This observation justifies the need to study $w \to \infty$ asymptotic theory. We show that the penalty parameter can be allowed to diverge at the rate dominated by $\sqrt{n}$.
Observe that the estimator in Theorem (ref) does not rely on Assumption A1, as the latter is always satisfied at a fixed measure for a large enough $n$ when $w_n \to \infty$.
The $w_n n^{-1/2}$ rate of convergence in Theorem (ref) is determined by the slowly vanishing penalty term. This term is a product of the deviation from the true polytope, that vanishes at $n^{-1/2}$, and an exploding sequence $w_n$. It is thus reasonable to ask whether the $\sqrt{n}-$rate could be restored by dropping the penalty term, i.e. debiasing the penalty function. We show that this can be done. Before we proceed, let us make the following simplification. Without loss of generality, suppose that
To see why (ref) can be assumed w.l.g., note that one can set $p = e_1 = (1 ~ 0 \dots 0)'$ and add an auxiliary variable for the value of the problem in the first position of $x$ (see gafarov2024simple).
We define the debiased penalty function estimator as
The following theorem establihes its rate, and is one of the main contributions of this paper.
We now outline the main ideas behind Theorem (ref). For $x \in \mathbb{R}^d$, let $J(x;\tilde{\theta}) \equiv \{j \in [q]: \tilde{M}_{j}x = \tilde{c}_{j}\}$ denote the set of constraints that bind at $x$ when evaluated at $\tilde{\theta}$. For ease of notation, from now on we write $\theta_0 = (\text{vec}(M)',c')'$. Let us introduce the following terminology.
Intuitively, the proof of Theorem (ref) proceeds in two steps. First, by anti-concentration arguments we establish that with high probability asymptotically the penalty function estimator manages to select a vertex-solution $\hat{x}_n$ such that the set of constraints that bind at it, $\hat{A}_n = J(\hat{x}_n;\hat{\theta}_n)$, corresponds to a nice face $F = \{x \in\mathbb{R}^d : M_{\hat{A}_n}x = c_{\hat{A}_n}\}$. Once a nice face has been selected, the $\sqrt{n}-$convergence of $p'\hat{x}_n$ to $B(\theta_0)$ obtains as a consequence of $(\hat{M}_{nA}, \hat{c}_{nA})$ converging to $(M_A, c_A)$ at this rate for a fixed $A \subseteq [q]$.
The discussion of uniform asymptotic theory in Section (ref) sheds light on the role of $w_n$ and the trade-off involved in its selection. The practical guidance on selecting $w_n$ is then developed on the basis of our results and random matrix theory in Appendix (ref).
This section develops an inference procedure for a general LP estimator, in which all parameters are inferred from the data. This procedure nests special cases in which some parameters remain fixed, as in semenova2023adaptive or bhattacharya2009inferring.
We suppose that Assumption B0 holds throughout Section 2.2, whereas the rest of the conditions are imposed explicitly.
Assumption B1 is typically warranted by reference to CLT and the Delta Method when $\hat{\theta}_n = g(n^{-1}\sum^n_{i=1}W_i)$ for some smooth $g(\cdot)$, as in the AICM models (ref), see Section (ref).
We construct a method for statistical inference on \( B(\theta_0) \) that achieves exact asymptotic coverage under minimal regularity conditions. Our approach is based on an asymptotically normal version of the debiased penalty estimator with \( w_n \to \infty \) w.p.a.1 and is outlined in Algorithm (ref). Before stating the main result, we introduce key auxiliary constructions and discuss our assumptions.
For the true $\theta_0 = (\text{vec}(M)',c')'$ and some subset of indices $A \subseteq [q]$, consider conditions
Equation (ref) is satisfied if constraints $A$ may bind simultaneously at some solutions of the original LP, while (ref) holds if the objective function's gradient $p$ is a linear combination of the gradients of inequalities from $A$. For example, if $A$ is a set of all binding constraints at some $x \in \mathcal{A}(\theta_0)$, equation (ref) follows from KKT conditions.
Continuing the discussion in Section (ref), we note that subsets $A$ that satisfy (ref) and (ref) correspond to the nice faces.
Define the set $\mathbb{A} \equiv \{A \in 2^{[q]}:|A| \geq d, A \text{ satisfies \eqref{optimality} and \eqref{range_req}}\}$. It is non-empty in any feasible finite LP. With probability approaching $1$, the penalty function estimator manages to select a vertex-solution $\hat{x}_n \in \tilde{\mathcal{A}}(\hat{\theta}_n; w_n)$, determined by the binding constraints $\hat{A}_n = J(\hat{x}_n;\hat{\theta}_n)\in \mathbb{A}$ that satisfy (ref) and (ref) and thus correspond to a nice face by Lemma (ref).
The debiased estimator may hence be understood as a two-stage procedure: one first finds the set of binding inequalities $\hat{A}_n \in \mathbb{A}$, and then estimates $B(\theta_0)$ as $p'\hat{M}_{n\hat{A}_n}^{\dagger}\hat{c}_{n\hat{A}_n}$. Performing inference on that object directly would require working with the complex joint distribution of $\hat{A}_n$, and $\hat{\theta}_n$, and would likely result in an asymptotically non-normal estimator.
We address this by `disentangling' the variation in \(\hat{A}_n\) and \(\hat{\theta}_n\) via sample splitting. Intuitively, a vertex is estimated on one part of the sample, while the noise in the parameter estimation comes from the other. We now state our assumptions and present the main result.
Assumption B2 requires the researcher to possess a consistent estimator of the asymptotic variance of $\hat{\theta}_n$. If $\hat{\theta}_n = g\left(n^{-1}\sum_{i = 1}^n W_i\right)$ for some smooth and known $g(\cdot)$, such estimator can typically be obtained from the estimated covariance matrix of $W_i$ via Delta-method. In more complicated scenarios, bootstrap on $\hat{\theta}_n$ may be employed.
Define the set $\mathcal{S}_{A} \equiv \{v \in \mathbb{R}^{|A|}: p = M'_{A} v\}$ and note that (ref) is equivalent to $\mathcal{S}_A \ne \emptyset$.
Assumption B3 is a technical condition ensuring that we can find a sequence approaching $\mathcal{S}_A$ asymptotically, i.e. $d(\check{v},\mathcal{S}_A) = o_p(1)$ for $\check{v}$ defined in (ref). Practical guidance on choosing $\overline{v}$ is provided in Appendix (ref). Simulation evidence in Figure (ref) in Appendix (ref) suggests that the specific value of \(\overline{v}\) has no impact on inference, as long as it is sufficiently large.
We randomly split $\mathcal{D}_n$ into two disjoint, collectively exhaustive folds $\mathcal{D}^{(f)}_{n}$ of size $n_f$ for $f = 1, 2$, with $n_1 = \lfloor\gamma n\rfloor$ and $n_2 = n - \lfloor \gamma n\rfloor$ for some fixed $\gamma \in (0;1)$. Our inference procedure uses the data from $\mathcal{D}^{(1)}$ to estimate an optimal triplet $(\hat{A},\hat{x},\hat{v})$. The vertex\footnote{While we assume that $\tilde{\mathcal{A}}(\hat{\theta}^{(1)};w_{n_1})$ is estimated precisely, the results do not change if one is only able to estimate a single optimum. This may occur if numerical errors do not allow the LP-solver to find all of the LP solutions. Such optimum will satisfy $\text{rk}(\hat{M}_{J(x;\hat{\theta}^{(1)})}) = d$ by definition, and so will be a valid vertex-solution.} is estimated as
the set of binding constraints that define it is denoted by $\hat{A} \equiv J(\hat{x};\hat{\theta}^{(1)})$, and
Finally, we define $\hat{v} \in \mathbb{R}^{q}$ so that $\hat{v}_{\hat{A}} = \check{v}$ and $\hat{v}_j = 0$ for $j \notin \hat{A}$.
Our procedure is then based on showing that, for large $n$,
where $\sigma^2(\cdot)$ is derived in Appendix (ref).
An inspection of the proof of Theorem (ref) below reveals that Assumption B4 rules out the scenarios when finding an $A \in \mathbb{A}$ determines the value $B(\theta_0)$, even though $\hat{\theta}$ is noisy. This may occur, for example, if the corresponding $\hat{c}_A, \hat{M}_A$ are deterministic. In this case, if $A$ is also unique, meaning $|\mathbb{A}| = 1$, the debiased estimator has $0$ asymptotic variance, because $A$ and therefore $B(\theta_0)$ are correctly estimated with probability approaching $1$.
To further justify the need for our inferential procedure, we examine the properties of the approach that combines bootstrap on $\hat{\theta}_n$ with the plug-in estimator $B(\hat{\theta}_n)$. This method is widely used in empirical literature applying AICM conditions (blundell2007changes, kreider2012identifying, pepper, siddique, de2017effect, and mentalhealth). In light of Proposition (ref), this approach is inapplicable when SC fails. In practice, researchers attempting to apply it to a LP with a small or empty interior of $\Theta_I$ may encounter frequent LP infeasibility in the bootstrap draws.
In some cases, SC may be established. This is true, for example, if the bound of interest can be expessed as an intersection bound $B(\theta_0) = \max \{c_1,c_2, \dots, c_q\} = \min_{t \in \mathbb{R}} t ~~ \text{s.t. } t \geq c_i, ~ i \in [q]$, where $\theta_0 = (1 ~ \iota' ~ c')'$ (as in CLR). Yet, even if SC holds, we demonstrate that bootstrap inference based on the plug-in estimator is not valid, unless further regularity conditions hold. This observation can be viewed as a generalization of the parameter-on-the-boundary problem of andrews1999estimation,andrews2000.
For simplicity of exposition, we abstract from the case of `true equalities' in $\Theta_I$. The results extend trivially to this case if SC is defined in terms of relative interior.
Hadamard directional differentiability of $B(\cdot)$ is sufficient for convergence in law.
Gaussianity of $B'_{\theta_0}(\mathbb{G}_0)$ is a necessary condition for bootstrap consistency santos2019. Consequently, the empirical literature using bootstrap with the plug-in estimator has implicitly relied on this assumption. However, $B'_{\theta_0}(\mathbb{G}_0)$ is not normal unless full Hadamard differentiability holds, i.e. $B'_{\theta_0}(h)$ is linear in $h$. As (ref) suggests, this is not generally the case. Theorem 3.1 in santos2019 establishes that bootstrap is inconsistent for the distribution when $B'_{\theta_0}(h)$ fails to be linear. The typically applied plug-in and bootstrap combination is then only valid under further restrictive assumptions\footnote{The full-support condition in Proposition (ref) is imposed for expositional purposes. Sufficiency of conditions i, ii holds generally, whereas necessity obtains whenever the derivative $B_{\theta_0}'(h)$ is not linear over $\text{Supp}(\mathbb{G}_0)$ when $\mathcal{A}(\theta_0), \Uplambda(\theta_0)$ are not singletons, see Lemma (ref) in Appendix.}:
The optimization problem (ref) is challenging to study under no further assumptions, as it may exhibit instability under arbitrary perturbations of parameters. We now show that this not only leads the plug-in $B(\hat{\theta}_n)$ to fail pointwise, but also precludes the existence of a uniformly consistent LP estimator over the unrestricted set of measures.
We first establish a new impossibility result that complements the findings of hirano. Specifically, we show that no uniformly consistent estimator exists for any discontinuous functional from the space of probability measures endowed with the total variation norm to an arbitrary metric space $(\mathcal{V}, \rho)$.
In this section, we treat the parameter $\theta_0$ as a functional of the underlying probability measure $\mathbb{P} \in \mathcal{P}$. We then make the following assumption on the pair $\theta_0(\cdot), \mathcal{P}$:
Assumption U0 formalizes the notion of the unrestricted set of measures. U0.i demands that the true parameter be continuous in $P$, which holds, for example, in AICM models (see Example (ref)). U0.ii assumes that $\theta_0$ has full support over $\mathcal{P}$, meaning that any $\theta$ corresponding to a consistent model with $\Theta_I(\theta) \subseteq \mathcal{X}$ is attained at some $P \in \mathcal{P}$.
Given this negative result, it is natural to seek a minimal restriction on $\mathcal{P}$ for which a uniformly consistent estimator may exist. We now show that the condition ensuring uniform consistency of the penalty function approach can be considered minimal in the sense to be made precise in Proposition (ref) .
Our examination of Example (ref) has shown that $w_n$ cannot be allowed to diverge faster than $\sqrt{n}$, as otherwise the penalty approach may fail at measures where SC fails. At the same time, if all KKT vectors $\lambda^* \in \Uplambda(\theta_0)$ grow large, an arbitrarily large $w$ is needed for Assumption A1 to hold. This occurs when optimal vertices become `sharp', i.e. all relevant full-rank submatrices of binding inequality constraints grow closer to being degenerate. The condition that ensures uniform consistency of the penalty function approach should therefore bound such `sharpness'.
We begin our construction with an existence result based on the Caratheodory’s Conical Hull Theorem.
Proposition (ref) asserts that any finite and feasible LP has an optimal vertex $x^*$ at which there is a subset $J^*$ of binding constraints, such that i) the corresponding gradients form a full-rank square matrix, and ii) the objective function gradient belongs to the conical hull formed by the gradients of the constraints from $J^*$.
The $\delta$-condition does not rule out the failure of LICQ, SC or NFF, and is weaker than the conditions under which uniform consistency of LP estimators has been established. To formalize this, let us introduce three families of measures. Firstly, denote the family of measures satisfying U1 for a given $\delta > 0$ by $\mathcal{P}^\delta$. A measure satisfies the Slater's condition if $\mathbb{P} \in \mathcal{P}^{SC} \equiv \{\mathbb{P} \in \mathcal{P}|\text{Int}(\Theta_I(\theta(\mathbb{P}))) \ne \emptyset \}$. Similarly, a measure satisfies a uniform $\varepsilon$-LICQ condition (as in gafarov2024simple) if $\mathbb{P} \in \mathcal{P}^{LICQ; \varepsilon}$, where
where the set $\mathcal{V}(\mathbb{P}) \subseteq 2^{[q]}$ consists of sets of indices of binding inequalities that define vertices of the polytope $\Theta_I(\theta(\mathbb{P}))$.
Intuitively, the $\delta>0$ in Assumption U1 merely parametrizes the degree of irregularity that the researcher is willing to allow for the identified polytope over the considered set of measures. The resulting family of sets `covers' the unconstrained set of measures asymptotically as $\delta$ decreases to $0$. Uniform LICQ and SC, on the contrary, both restrict the set of measures. That is because measures like the one in Figure (ref) do not belong to either $\mathcal{P}^{LICQ;\varepsilon}$ for any $\varepsilon \geq 0$, or $\mathcal{P}^{SC}$.
The penalty function estimator converges to the LP value $B(\mathbb{P})$ a.s. at rate $\sqrt{n}w_n^{-1}$, uniformly over the set of distributions that satisfy the $\delta-$condition for some $\delta >0$.
In applications, condition i) in Theorem (ref) can usually be established with rate $\overline{r}_n = \sqrt{n}$ by reference to the uniform LLN, provided $\theta_0$ is uniformly continuous in population moments.
This subsection provides a variation of the $\delta-$condition under which the debiased penalty estimator is shown to be uniformly consistent. Intuitively, the debiased estimator achieves uniform consistency whenever for all considered measures it is true that if the distance of $x \in \mathcal{X}$ from the polytope $\Theta_I(\mathbb{P})$ is large, then the constraint violation $\iota'(c(\mathbb{P})-M(\mathbb{P})x)^+$ is also large. To develop that notion formally, we introduce the following two terms.
The following proposition simplifies the interpretation of $\kappa(\Theta)$. It asserts that in the definition of $\kappa(\Theta)$ it suffices to only consider the vertices of $\Theta$,
We are now ready to define the polytope $\delta-$condition.
The polytope $\delta-$condition lower bounds the smallest singular values of all full-rank matrices that can be constructed from the vertices of the polytope. Similarly to Assumption U1, it parametrizes the unconstrained set of probability measures, since at any fixed $\mathbb{P} \in \mathcal{P}$ we have $\kappa(\Theta_I(\mathbb{P})) > 0$ by definition. Proposition (ref) continues to hold for U2, if one substitutes $\mathcal{P}^\delta$ with $\mathcal{P}_p^\delta$ - the family of measures satisfying U2 for a given $\delta > 0$.
The following Lemma provides a minorization of the $L_1-$deviation from a polytope, and appears to be mathematically novel. Intuitively, $\kappa(\Theta)$ determines `sharpness' of the vertices of a polytope. The less sharp is a vertex, the greater the polytope constraints' violations need to be in terms of the distance from the polytope. Figure (ref) illustrates that point.
By Theorem (ref), the biased penalty estimator is uniformly consistent at rate $\frac{\sqrt{n}}{w_n}$ under the $\delta$-condition. The following Theorem asserts that the debiased estimator converges at least at the same rate.
Consistency and inference simulations use $10^4$ and $10^3$ repetitions respectively.
We first describe Example (ref) in more detail. Consider a setup with $d = 2$ variables and $q = 4$ constraints. Recall that $\theta = (p', \text{vec}(M)', c')'$. In this case, the parameters are defined as
The sample analogues $\hat{M}_n, \hat{c}_n$ of parameters in (ref) are obtained by substituting $b$ with its estimator $\hat{b}_n$. In our simulation study, that estimator is given by $\hat{b}_n = b + n^{-1}\sum^n_{i=1} U^b_i$, where $U^b_i \sim U[-1; 1]$ for $i \in [n]$ are i.i.d.
The true parameter $b$ is in one-to-one correspondence with the underlying measure, and $\theta$ is completely described by it: $\theta_0(\mathbb{P}) = \theta_0(b(\mathbb{P}))$. Consequently, it is sufficient to index the parameter $\theta$, the identified polytope $\Theta_I$ and the program's value $B$ by $b$ only.
Observe that the example (ref) is so engineered that $\Theta_I(\hat{b}_n)$ would never be empty or unbounded, ensuring the existence of the plug-in estimator $B(\hat{b}_n)$. A slight modification to that setup, that we consider in the next subsection, would lead the identified polytope to be empty with a potentially non-vanishing probability.
\paragraph{Measures} We consider two values of $b$: $b = 0$ and $b = -0.05$. At $b = 0$, SC fails, because $\Theta_I(0)$ has an empty interior. LICQ also fails at $b = 0$, as at the optimum $x^* = (-1,-1)'$, inequalities $1,2,4$ are binding. There are no flat faces at $b = 0$. At $b = -0.05$, SC, LICQ and NFF all hold. The values $b = 0$ and $b = -0.05$ result in the smallest singular values of $M_{J^*(b)} (b)$ matrices that correspond to the $75-$th and the $19-$th percentiles of the Tao-Vu distribution given in Theorem (ref) in the Online Appendix.
If $b < 0$, the norm of the `smallest' KKT vector $\lambda$ in the true LP corresponding to (ref) is proportional to $|b^{-1}|$. So, for a small negative $b$, the $\delta$-condition is only satisfied for a small value of $\delta > 0$. In this case the `optimal' $w$ is large. By that logic, the plug-in estimator, which in this case obtains as the limit $w \to \infty$, should perform well at such $b$, and potentially outperform the debiased penalty estimator. This is partly an artifact of the setup in (ref): additional noise in one of the slated lines' intercepts would render the problem infeasible with positive probability, which would worsen the performance of the plug-in estimator in finite samples (see Figure (ref)).
At $b \geq 0$, the $\delta$-condition is satisfied with a relatively large $\delta = \frac{\sqrt{5} - 1}{2}$, and so the `optimal' $w$ is relatively small. If $b = 0$, in 50% of the cases, $\hat{b}_n < 0$ is estimated. If, moreover, $w_n > \text{const}\times |\hat{b}_n^{-1}|$, the debiased penalty estimator would select the incorrect maximum of $0$ in such cases. A larger $w$ at $b = 0$ thus hampers the performance of the debiased penalty estimator.
\paragraph{Parameters} We set $w_n = \delta^{-1}_{0.2}(d) ||p||\frac{\ln \ln n}{\ln \ln 100}$, where $\delta_{\alpha}(d)$ is the $\alpha-$quantile of the $d-$dimensional Tao-Vu distribution given in Theorem (ref) in Appendix (ref). To ensure the same expansion rate for the set-expansion estimator, we set $\sqrt{\kappa_n} = \kappa_0 \times \ln \ln n$. There is no guidance as to the selection of $\kappa_0$. Our baseline is $\kappa_0 = 0.1$, and we explore other values in Figure (ref).
\paragraph{Discussion} Consider the right panel of Figure (ref), corresponding to $b=0$. The plug-in estimator is inconsistent, while the set-expansion estimator performs well, approaching $-1$ from below for larger $n$. The debiased penalty-function estimator is slightly upward-biased for smaller $n$, but yields the value of almost exactly $-1$ in larger samples. In contrast, the set-expansion estimator has a conservative rate, and remains slightly downward-biased even in larger samples.
The case of $b = -0.05$ is depicted in the left panel of Figure (ref). The plug-in estimator is consistent and appears to be the best estimator out of the three. The set-expansion estimator is severely downward-biased. This is because when the optimal vertex has a `sharp' angle, a small expansion of the inequalities' RHS may lead to a large shift of the vertex. To see that, consider shifting both inequalities outwards in Figure (ref) for a small and a large absolute value of $b$, when $b$ is negative. Once the expansion grows smaller, the set-expansion estimator slowly converges to the true value of $0$ from below. While selecting a smaller $\kappa_0$ parameter would improve the performance of the set-expansion estimator at $b = -0.05$, in the next part of our analysis we demonstrate that in this example $\kappa_0 = 0.1$ is close to being optimal in the uniform sense, because smaller $\kappa_0$ worsens the estimator's performance at $b = 0$. The debiased penalty estimator, in contrast, converges rather quickly. It is slightly conservative at smaller $n$, as it selects the incorrect vertex of $-1$ whenever $\hat{b}_n < 0$ and $|\hat{b}^{-1}_n| > \text{const} \times w_n$, i.e. when the penalty parameter is not large enough.
\paragraph{Robustness of the debiased penalty estimator} For $b < 0$, the debiased estimator may be expected to perform better at measures with a larger $|b|$. Figure (ref) illustrates that point:
Researchers applying any LP estimator should exercise caution when operating in smaller samples due to irregularity inherent in (ref). In example (ref), the debiased penalty estimator exhibits desirable behavior even along highly irregular measures for sample sizes of order $n = 5000$. Such sample sizes are not uncommon in partially identified settings. In our application, estimation is performed on $664633$ observations.
\paragraph{Alternative $\kappa_0$} We now investigate to which extent the behavior of the set-expansion estimator can be improved by selecting an alternative $\kappa_0$ parameter. Clearly, a larger $\kappa_n$ makes the set-expansion estimator more conservative. However, $\kappa_n$ cannot be selected as small as possible. If the expansion is too small, it may be insufficient to counteract the noise involved in estimating $\Theta_I$, leading to poor performance of the estimator at measures where SC fails to hold.
That tradeoff is very clear in Figure (ref). Compare the performance of the estimator with $\kappa_0 = 0.01$ to the baseline of $\kappa_0 = 0.1$. Decreasing $\kappa_0$ down to $0.01$ makes the estimator less conservative at $b = -0.05$, but results in an upward bias at $b = 0$. This occurs, because at $\kappa_0 = 0.01$ the resulting set-expansion sequence $\kappa_n$ is too small to counteract the estimation noise for the considered sample sizes. This logic is `monotone', meaning that selecting an even smaller $\kappa_0$ would worsen the performance at $b = 0$ further. Even at $\kappa_0 = 0.01$, the set-expansion estimator is quite conservative for the measure $b = -0.05$, and the parameter can clearly not be reduced any further without affecting the validity of the estimator at $b = 0$. It appears that the baseline choice of $\kappa_0 = 0.1$ is close to being optimal in our example. Therefore, the conservative behavior of the set-expansion estimator at $b=-0.05$ is not explained by a poor choice of the tuning parameter, but rather is a feature of the estimator itself. \paragraph{Noise in $\hat{c}_n$} We also report the result of consistency simulations for the DGP described in (ref). This DGP features noise on the right-hand-side of the inequalities that describe the polytope, which allows the estimated polytope to be empty.
In Figure (ref), the plug-in and the debiased penalty estimators perform equally well at $b = -0.05$. The rest of the conclusions are qualitatively unchanged.
In this section, we assess the performance of our inferential procedure. We compare it to the performance of two recently developed methods, chorussel2023 (CR) and gafarov2024simple (BG). We consider a slightly modified version of example (ref):
The sample analogues $\hat{M}_n, \hat{c}_n$ of parameters in (ref) are obtained by substituting $b, \zeta, \nu$ with their estimators:
where $U^{k}_i \sim U[-0.5;0.5]$, i.i.d. across $i \in [n]$ and $k \in \{b, \zeta, \nu\}$. We consider measures, for which the true values of $\zeta = \nu = 0$, whereas we still vary $b$ as in the example before. As before, we mainly consider $b = -0.05$ and $b = 0$.
Note that in (ref) the plug-in estimator of the polytope $\hat{\Theta}_I$ may be empty for some realizations of $\{U^k_i\}_{i,k}$. That is the case, for example, if $\hat{b}_n = \hat{\zeta}_n = 0$ and $\hat{\nu}_n > 0$.
\paragraph{Other methods} In case SC fails, the BG estimator is not applicable. It fails to exist in around 25% of cases at $b = 0$. The CR estimator in the main text of the paper also relies on $B(\hat{\theta}_n)$ and is not applicable if SC fails. The procedure described on p.47 of the Supplementary Appendix in chorussel2023 would not be practically applicable without the results contained in the present paper. It can be shown that the CR augmented procedure combines a random, non-vanishing set expansion and objective function perturbation with the penalty function approach\footnote{Because the penalty function estimator can be rewritten in the form of an equivalent problem w.p.a.1.}. The authors, however, treat the analogue of the penalty parameter as `some [fixed] large value'. They proceed to argue that the inequalities, which establish the size of their inference procedure, hold for any value of $w$, which appears to suggest that its selection is unimportant. There is no practical guidance on $w$ selection, and CR do not implement the augmented estimator in their simulations. In this paper, we studied penalized estimation in great detail, and our results suggest that the appropriate choice of $w$ is critical. Both our previous findings and simulations in Figure (ref) demonstrate that for different values of $w$ the performance of the CR procedure ranges from highly conservative to invalid. Selecting `a large value' of $w$ does not yield a valid procedure in finite samples. Our simulations also suggest that CR augmented approach can perform relatively well if combined with our results on the rate and level of $w_n$. Unlike our approach, however, it remains asymptotically conservative due to the use of non-vanishing random expansions.
\paragraph{Implementation details} As mentioned in Section 2.2, one can obtain an estimator $\hat{\Sigma}_n$ for the asymptotic covariance matrix of $\hat{\theta}_n$ via resampling. One then plugs it into the expression for $\sigma(\cdot)$ to obtain the required s.e. An even simpler approach is to obtain an estimator for $\sigma(\hat{A},\hat{x},\hat{v},\Sigma)$ directly by bootstrapping the quantities estimated on the second fold, while keeping the first fold quantities fixed. We have verified that the performance of the two approaches is similar to using the closed-form estimator for $\Sigma$, namely $\hat{\Sigma}_n = G \hat{\Omega}_n G'$, where $G \equiv \frac{\partial \theta}{\partial (b ~ \zeta ~ \nu)'}$ and $\hat{\Omega}_n$ is the sample covariance matrix of $(b_i ~ \zeta_i ~ \nu_i)'$. We employ the latter estimator in our simulations. When implementing the procedure in chorussel2023 (CR), we use uniform noise with the support size of $0.001$, as recommended in the paper. Note that we refer to their parameter $M$ as $w$. The estimator in gafarov2024simple (BG) was implemented using the code kindly provided by the author.
\paragraph{Discussion} Results of our simulations are given in Figure (ref). Overall, it appears that our estimator has the correct nominal level even at smaller sample sizes, whereas the CR with penalty $w_n$ overcovers asymptotically. BG estimator achieves nominal coverage asymptotically, although it over-covers in smaller samples and may yield a very conservative left confidence bound, as is evident from the right panel. It fails at $b = 0$, because the SC fails\footnote{We have still run the corresponding simulations, but are not displaying them. BG fails to exist in around 25% of cases, while the remaining simulations result in highly conservative bounds with incorrect coverage}. For different fixed values of $w$ the Cho-Russell procedure's performance may range from highly conservative to invalid. We illustrate this by adding lines with $w = 2$ and $w = 1000$ to the figures corresponding to $b = -0.05$ and $b = 0$, respectively.
Let $Y \in \mathbb{R}$ denote the outcome of interest\footnote{Univariate case is considered for simplicity of exposition, but the extension to multivariate outcomes is immediate.}, $T \in \mathbb{R}$ stand for the treatment, and $Z \in \mathbb{R}^{d_Z}$ be the candidate instrument. Here treatment is any variable which effect on $Y$ we attempt to infer, whereas the term instrument refers to an auxiliary variable that allows us to partially identify the causal effect of interest. Our approach nests the case when $Z$ is the usual IV. The reader seeking an economic intuition may refer to Section (ref), where $Y$ is the log-wage, $T$ is the level of education, and $Z$ is a proxy for the level of ability.
Denote the supports $Y,T,Z$ as $\mathcal{Y}$, $\mathcal{T}$ and $\mathcal{Z}$ respectively. Throughout this section we consider the case of continuous outcome and discrete treatment and instrument, namely $\mathcal{Y}$ is a (possibly infinite) interval\footnote{We focus on the uncountable case for concreteness. Inclusion (ref) in Theorem (ref) holds generally, while sharpness also holds for the case of arbitrary $\mathcal{Y}$ if there are no almost sure inequalities.}, while $N_T \equiv |\mathcal{T}| < \infty$ and $N_Z \equiv |\mathcal{Z}| < \infty$. In non-parametric bounds literature it is rather conventional to employ a discrete instrument at the estimation stage (see MP2009). While our main identification result may be extended to continuous $Z$, estimation in such settings is left for future research.
Our setup accommodates missing observations of the dependent variable. Namely, we split the set of treatments into two disjoint subsets $\mathcal{T} = \mathcal{O} \sqcup \mathcal{U}$. Whenever $T \in \mathcal{O}$, the researcher observes $Y, T, Z$, whereas if $T \in \mathcal{U}$, only the covariates $T, Z$ are observed. For example, in blundell2007changes the wage is observed only if an individual is employed. Corresponding to the legs of the treatment are the potential outcomes $Y(t), ~ t \in \mathcal{T}$:
Continuing the wages and education example, the value of $Y(t)$ for a fixed $t\in \mathcal{T}$ may then correspond to the potential wage that an individual with the associated random characteristics would get, had she obtained education $t$.
Let us collect the potential outcomes in the vector $\mathbb{Y} \equiv (Y(t))_{t \in \mathcal{T}} \in \mathbb{R}^{N_T}$. Variables $(\mathbb{Y}, T, Z, Y)$ are jointly defined on the true probability space $(\mathbb{P}, \Omega, \mathcal{S})$ and we let $\mathcal{P}$ denote the considered collection of probability measures on $(\Omega, \mathcal{S})$, such that $\mathbb{P} \in \mathcal{P}$. We impose the following conditions on the set of considered measures throughout this section:
Part i) of the Assumption I0 formalizes the assumed identification pattern. It says that the joint distribution of $T, Z$ is always identified and the researcher also observes the joint distribution of $Y, T, Z$ whenever $T \in \mathcal{O}$. Parts ii) and iii) of I0 ensure that all conditional expectations and probabilities are well-defined and finite\footnote{Similar identification results can still be obtained if one relaxes the full-support condition for some known pairs from $\mathcal{Z} \times \mathcal{T}$. Note that it can also be verified in the data.}.
We define the vector $m$ collecting all elementary conditional moments as
and suppose that the researcher is interested in the target parameter $\beta^*$, given by
where $\mu^{*}: \mathcal{P} \to \mathbb{R}^{N^2_T N_Z}$ is identified\footnote{By this we mean that $\mu^*(P) = \mu^*(\mathbb{P})$ for all $P \in \mathcal{P}$.} and chosen by the researcher. It parametrizes the choice of the outcome of interest, as the following remark clarifies.
We now introduce the general class of identifying conditions described by affine inequalities over conditional moments, potentially augmented with affine a.s. restrictions. These restrict the set of admissible measures to $\mathcal{P}^*$:
where $b^*: \mathcal{P} \to \mathbb{R}^R$, $M^*: \mathcal{P} \to \mathbb{R}^{R \times N^2_T N_Z}$, $\tilde{b}: \mathcal{P} \to \mathbb{R}^{\tilde{R}}$ and $\tilde{M}: \mathcal{P} \to \mathbb{R}^{\tilde{R} \times N_T}$ are identified parameters, chosen by the researcher. These parametrize the choice of $R \in \mathbb{N}$ identifying inequalities on conditional moments of potential outcomes as well as $\tilde{R} \in \mathbb{N}$ almost sure inequalities on the potential outcomes. In general, $\psi(P) \equiv (\mu^*(P),b^*(P), M^*(P), \tilde{b}(P), \tilde{M}(P))$ and $m(P)$ are functionals of $P$. We omit this dependence whenever it does not cause confusion.
The family of models that can be written in the form (ref) is very rich, as illustrated by the following examples.
We now provide the general identification result for the models described by $\mathcal{P}^*$. Let us construct $x$ that collects unobserved pointwise-conditional moments\footnote{Formally, $x$ is a partially identified functional $x = x(P)$, whereas $\overline{x} = \overline{x}(P)$ is point-identified.} and the vector $\overline{x}$ of those pointwise-conditional moments that are identified:
For known selector matrices $P_m, \overline{P}_m$, one can then decompose $m$ as
It is also straightforward to observe that $\tilde{M} \mathbb{Y} + \tilde{b} \geq 0$ a.s. implies
Before we state the main identification result, let us construct the matrix $M^{**}$ and the vector $b^{**}$ that combine the conditional restrictions with the implications of the almost sure restrictions in (ref), as
The matrix $\tilde{M}_{MTR}$ is defined in Appendix (ref).
Theorem (ref) postulates that bounds on the target parameter $\beta^*$ under $\mathcal{P}^*$ can be obtained by solving two linear programs. The LP bounds are sharp if there are no a.s. inequalities in the model, or if the a.s. inequalities parametrize three special cases that are typically used in the literature. Otherwise, the LP bounds are valid, but not necessarily sharp, as we show in Appendix (ref). This is because, in general, the entire distribution of $\mathbb{Y}, T, Z$ is relevant for $\beta^*$ under a.s. restrictions. The naive approach of searching over such joint distributions, however, would involve infinite-dimensional optimization, because $|\mathcal{Y}| = \infty$.
A particular family of identifying conditions that can be written in the form (ref) is the conditional monotonicity class of assumptions. These impose that potential outcomes are mean-monotone in the instrument even within some treatment subgroups. While more restrictive than the conventional MIV, conditionally monotone instrumental variables (cMIV) allow to sharpen the bounds on the outcomes of interest. Throughout this section, we assume that outcomes are bounded $Y(t) \in [K_0; K_1]~ a.s.$ for known $K_0, K_1 \in \mathbb{R}$, $K_0 < K_1$. We also suppose that there are no missing data\footnote{Although it is hopefully clear from our general approach how cMIV conditions extend to the missing data case.}, i.e. $\mathcal{T} = \mathcal{O}$.
We argue that cMIV assumptions are reasonable in classical applications, discuss the difference between MIV and cMIV and develop a formal testing strategy for a particular version of cMIV. This testing procedure relies on the observed outcomes' monotonicity, which has been typically used in applied work to justify applying MIV. Our results imply that if such monotonicity is observed and the researcher is comfortable with MIV, the cMIV assumption is inexpensive, and can be applied to sharpen the bounds on the outcomes of interest. In some applications, as is the case in Section 5, cMIV yields informative bounds even when the classical conditions fail to do so.
While we only discuss three variations of cMIV, the class of such assumptions is potentially richer\footnote{One can consider the class of conditional restrictions $\mathbb{E}[Y(t)|T \in A, Z = z'] \geq \mathbb{E}[Y(t)|T \in A, Z = z], ~ ~ \forall A \in \mathcal{F}_t$ for all $t \in T$ where subcollections $\mathcal{F}_t \subseteq \mathcal{T}$ are chosen by the researcher. The triplet $M, p,c$ from Theorem (ref) under any such restrictions can be constructed by following the procedure in Appendix (ref).}, and Theorem (ref) applies in any such framework.
The strong conditional monotonicity assumption possesses the greatest identifying power across all considered cMIV conditions. To see that cMIV-s implies MIV, set $A = \mathcal{T}$ in the definition above.
The weak conditional monotonicity assumption allows for closed-form expressions for sharp bounds that are easy to compute and perform inference on.
The pointwise conditional monotonicity assumption is directly testable under a mild homogeneity condition, see Section (ref).
Conditional monotonicity restrictions differ in the collection of treatment subsets over which monotonicity in the instrument is assumed. The strong conditionally monotone instruments are such that, among individuals from any given counterfactual treatment subgroup, higher values of $Z$ are, on average, associated with higher potential outcomes. The weak conditional monotonicity restriction only imposes the same mean-monotonicity on the whole population and on the untreated, whereas the pointwise form assumes it over the entire population as well as conditional on each counterfactual level of treatment.
While it is possible for the general apporach of form (ref), cMIV conditions avoid assuming monotonicity over the observed treatment subset $\{T = t\}$. This is because such monotonicity is identified. If it holds, it should not add any identifying power to our conditions in theory. However, large violations of the observed outcomes' monotonicity will lead the test developed in Section (ref) to reject cMIV-p and cMIV-s.
The following observation motivates the use of cMIV assumptions.
Sharp bounds for all versions for cMIV follow from Theorem (ref). We also show that under cMIV-w the bounds can be characterized explicitly, which is especially convenient if the treatment is binary, so that all cMIV assumptions coincide (see Appendix (ref)). For didactic purposes, we also provide detailed guidance on how to construct the triplet $M, c, p$ from Theorem (ref) under cMIV-s, cMIV-p and MIV in Appendix (ref).
This section illustrates the difference between MIV and cMIV by considering two parametric examples with classical applications.
Consider the following empirical setup. Suppose $T$ is an indicator of whether or not an individual has a university degree, $Y(t)$ are potential log wages and $Z$ is an observed indicator of ability.
MIV assumption on $Z$ implies that more able individuals can do better both with and without a college degree on average: $\mathbb{E}[Y(t)|Z = z]$ - monotone in $z$. cMIV additionally imposes that: i) among those who have a college degree, a smarter individual could have done relatively better on average than their counterpart if both did not have it: $\mathbb{E}[Y(0)|Z = z, T = 1]$ - monotone in $z$; and ii) among those who do not have a college degree, a smarter individual could have done relatively better on average than their counterpart if both had it: $\mathbb{E}[Y(1)|Z = z, T = 0]$ - monotone in $z$.
We now consider a parametric example. Suppose that $\eta$ measures how diligent one is from birth and is ex-ante mean-independent of $Z$. While $Z$ is observed by both the employers and the econometrician (e.g. an IQ score), the employer additionally observes the employee effort level $\eta + \varepsilon$ with $\varepsilon \rotatebox[origin=c]{90}{$\models$} (Z, T, \eta)$. Suppose $Var(Z) = Var(\eta) = 1$ and $\mathbb{E}[Z]=\mathbb{E}[\eta] = \mathbb{E}[\varepsilon] = 0$. Suppose that, on average, employees choose $T$ to maximize their expected earnings. This motivates a stylized Roy selection model, with
where $\nu \rotatebox[origin=c]{90}{$\models$} (Z, \eta, \varepsilon)$ is remaining heterogeneity, and $\varepsilon(t) \equiv \beta_2(t) \varepsilon$. MIV demands that
MIV postulates that the direct effect of ability on potential earnings is positive. It seems reasonable to suppose that $\beta_i(t) \geq 0, ~ i = 1,2, ~ t = 0, 1$, i.e. both effort and ability increase potential wages. Letting $\delta_z \equiv \beta_1(1) - \beta_1(0)$ and $\delta_\eta \equiv \beta_2(1) - \beta_2(0)$ denote the differentials in the effects of ability and effort respectively, the additional requirement of cMIV is that:
where $\tilde{\nu} \equiv \beta_0(1) - \beta_0(0) + \nu$.
Notice that if $\delta_z$ and $\delta_\eta$ are of different signs, for example because the jobs that one may apply for with a college degree are more ability-intensive ($\delta_z > 0$), whereas those which are available otherwise are more skill-intensive ($\delta_\eta < 0$), the additional conditional monotonicity requirements (ref)-(ref) are less strict than MIV. This is because, conditional on both having a degree and not having it, ability and effort are positively associated.
Intuitively, among those who do not have a degree ($T=0$), people of higher ability must have had stronger incentives to forgo college. This should have been because a higher level of diligence gives them a comparative advantage in effort-intensive jobs. Among those with a degree, higher ability implies a comparative advantage in ability-intensive occupations, which explains their willingness to select into this option ($T=1$). It does not, therefore, signal as low an effort level as it would for a less capable individual.
Now consider the same setup with $\delta_z = \delta_\eta > 0$ and $\tilde{\nu} = 0$,
This selection mechanism can be explained by the fact that to get a degree one needs to be either hard-working or of high ability. The requirement of MIV is unchanged, and cMIV necessitates that:
In this case, conditional on each level of education, effort level $\eta$ and ability $Z$ are negatively associated, so the conditional selection terms in (ref)-(ref) make cMIV a stricter assumption than MIV. Intuitively, a more able individual with a college degree did not need to work as hard to get it as her counterpart with a lower ability. Similarly, if an individual is capable, but does not have a degree, she has to be of lower effort as otherwise she would have selected into education.
Even if MIV holds, cMIV can thus fail if employer prefers effort over ability to the extent that the conditional negative association between the two outweighs the direct impact of ability on wages as well as any ex-ante positive correlation between the employer-observed signal of diligence and the ability.
An examination of equations (ref) and (ref) suggests that cMIV is more likely to hold whenever $\delta_z$ is small relative to $\delta_\eta$, while $\beta_1(\cdot)$ is large relative to $\beta_2(\cdot)$. This means that $Z$ should be relatively weak in the parlance of the classical IV models, and strongly monotone.
Overall, it seems reasonable to use a proxy for the level of ability as a conditionally monotone instrument in the estimation of returns to schooling. One would be inclined to think that while $Z$ does enter selection, it affects the potential outcomes directly and strongly enough, so that there are no subgroups by schooling for which a higher value of ability would correspond to lower potential wages on average.
As some aspects of mathematical intuition may be muted in discrete models, we also consider a simple continuous setup to confirm the insights derived from the previous analysis. For illustrative purposes, we drop the boundedness and discreteness assumptions and consider the supply and demand simultaneous equations,
The observed log-price $P$ clears the market in expectation,
where $\eta, Z$ are continuous unobserved and observed random variables respectively, with $\mathbb{E}[\eta|Z] = 0$ a.s.\footnote{Mean independence is not restrictive, as it can always be enforced by redefining the d.g.p. in an observationally equivalent way.} and $\mathbb{E}[\varepsilon^k] = 0$ with $\varepsilon^k \rotatebox[origin=c]{90}{$\models$} (\eta, Z, \varepsilon^{-k})$ for $k \in \{s,d\}$. Further assume that all functions of $p$ are continuous.
Potential price $p$ indexes the potential outcomes, giving rise to the demand and supply schedules. Suppose we aim to identify the elasticity of supply, $\mathbb{E}[(q^s(p_1) - q^s(p_0))/(p_1 - p_0)]$ for some $p_1 > p_0$, and $Z$ is a monotone instrument for $q^s(p)$, while $P$ can be interpreted as treatment. $\eta$ is unobserved heterogeneity and $\varepsilon^k$ are random violations from the market clearing condition or measurement errors independent of the rest of the model. For an individual realization of market clearing an econometrician observes $\{P, \{q^k(P)\}_k, Z\}$, but does not observe the schedules at other prices $\{q^k(p)\}_k \text{ for } p \ne P$, nor disturbances $\{\eta, \{\varepsilon^k\}_k\}$.
Define $\delta_z(p) \equiv \beta^s(p) - \beta^d(p)$ and similarly for $\eta$, with $\delta_p(p) \equiv \alpha^s(p) - \alpha^d(p)$. As stated, the model is potentially incomplete or incoherent, as for a given vector $(Z, \eta)$ equation (ref) may have multiple or no solutions. To avoid that, so long as that the support of $Z, \eta, \varepsilon^k$ is full, it is necessary that $\delta_z(p)$, $\delta_\eta(p)$ be constant. We shall assume that for simplicity. Provided that $\delta_p(p)$, which determines the excess supply at fixed $(Z, \eta)$, is strictly increasing and has full image, the model is complete and coherent, and
Equation (ref) introduces a deterministic linear relationship between $Z$ and $\eta$ conditional on each given value of $P$. As we saw in the previous example, this constitutes the worst-case scenario for cMIV, if $\delta_z$ and $\delta_\eta$ have the same sign. A noisier selection mechanism would relax the conditional link between $Z$ and $\eta$, and would thus weaken the conditional selection channel.
Note that the reduced-form error is $u \equiv \gamma^s(P) \eta + \kappa^s(P) \varepsilon^s$ and there is a simultaneity bias, as
In this setup, MIV requires that
Whereas cMIV-p additionally imposes that
Suppose that $\delta_z, \delta_\eta \ne 0$ to rule out uninteresting cases. (ref) then implies
For concreteness, consider two positive supply shocks, i.e. $\beta^s(p), \gamma^s(p) > 0$. Equation (ref) means that either $\eta$ and $Z$ affect the reduced-form equilibrium price in different directions (recall the comparative advantage example), or the effect of $Z$ on the equilibrium price relative to its effect on the supply schedule is smaller than that of $\eta$,
Under $\text{sgn}(\delta_\eta) = \text{sgn}(\delta_z)$, equation (ref) once again requires that $Z$ be strongly monotone and relatively weak for cMIV-p to hold. The logic we described may help the researcher navigate the potential economic forces in a given application to decide whether cMIV-p is a suitable assumption.
For example, consider estimating the supply elasticity in the market for plane tickets in the early days of Covid-19 pandemic. Suppose $Z$ is an inverse Covid-stringency index for the economy, while $\eta$ may be interpreted as residual cost shocks, defined to be mean-independent of $Z$. It is likely that $\delta_\eta \approx \gamma^s$, i.e. residual cost shocks affect mainly the supply in that sector, and not the demand. It is also likely that either supply is less responsive to $Z$ than demand (so that cMIV is implied by MIV), or the effects are of the same order of magnitude. $Z$ is therefore likely to be a conditionally monotone instrument.
One could argue against cMIV conditions whenever $\mathbb{E}[Y(t)|T = t, Z= z]$ fail to be monotone in the data. In general, the power and size of that test are unclear. There is, however, a special case when cMIV can be tested directly, provided that the researcher is willing to assume MIV. In some applications one may conjecture that the potential outcomes' functions $Y(t)$, either in the reduced or in the structural form, are such that the relative effects of $Z$ and the unobserved variable(s) $\eta$, potentially correlated with $Z$, are unchanged across outcome indices $t$.
Researchers often impose even stricter versions of this homogeneity assumption. For example, MP2009 discuss MIV identification under HLR condition: $Y(t) = \beta t + \eta$. Conditions in Proposition (ref) relax HLR to an arbitrary shape of response of a potential outcome to treatment and allow for a generally heteroscedastic/treatment-specific response to unobserved variables and instrument, so long as the relative effects are unchanged across potential outcomes. All functions in the proposition below are assumed to be measurable.
Note that whether or not $h(t) \ne 0$ is observable in the data for case $(a)$ and whether or not $h(t)/h(d)>0$ is also identified for $(b)$.
A test of cMIV-p is thus the test of all $f_t(z) \equiv \mathbb{E}[Y(t)|T = t, Z = z]$ being monotone,
To obtain such a test, we may extend the procedure in chetverikov_2019\footnote{This test is developed for continuous $Z$, which is used in our application. Although the instrument is discretized at the estimation stage, the monotonicity of $\mathbb{E}[Y(t)|T = t, Z = z]$ for continuous $Z$ clearly implies the monotonicity of the discretized moments. The procedure we describe straightforwardly accommodates testing discrete instruments. As noted in chetverikov_2019, for discrete conditioning variable the test is a simple parametric problem, since the conditional moment function can be $\sqrt{n}-$consistently estimated at each point from the support.}. Denote the set of all observations with treatment level $t$ as $\mathcal{I}_t \equiv \{i \in [n]: T_i = t\}$ with $n_t \equiv |\mathcal{I}_t|$. Suppose $\phi^{t}_{n_t}$ is the corresponding Chetverikov's regression monotonicity test (or a corresponding parametric test for discrete $Z$) with the confidence level $\alpha_t \in (0; 0.5)$. We define the joint test as
Denote $\mathcal{P}^C$ to be the set of probability measures, such that for all $P \in \mathcal{P}^C$ and all $t \in \mathcal{T}$ the conditional probability measure given $T=t$ that $P$ generates satisfies the regularity conditions in Theorem 3.1 in chetverikov_2019. Similarly, let $\mathcal{P}^C_t$ be the set of all the conditional probability measures given $T = t$ that measures from $\mathcal{P}^C$ generate.
Our data is comprised of 861492 observations from Colombian labor force. The sample represents a snapshot of those individuals who could be matched across the educational, formal employment and census datasets in 2021\footnote{Educational dataset was assembled by the testing authority Instituto Colombiano para la Evaluación de la Educación (ICFES), formal employment dataset comes from social security data based on Planilla Integrada de Liquidación de Aportes (PILA), whereas census data is handled by Departamento Administrativo Nacional de Estadística (DANE). The data was merged and anonymized by ICFES.}. For 664633 individuals from this dataset we observe their average lifetime wages, education level and Saber 5 or Saber 11 scores for Mathematics and Spanish language tests\footnote{Saber 5 and 11 tests are taken at different ages, but designed to be comparable between each other, which justifies merging them.}.
The outcome variable we consider ($Y_i$) is a log-wage, and $T_i$ is the education level. We distinguish four education levels: primary, secondary and high school, as well as `university'\footnote{$T_i$ is based on the number of years of schooling, $S_i$. If $S_i < 9$, set $T_i \equiv 0$ meaning the individual only graduated from primary school. $S_i \in [9;11)$ and $T_i \equiv 1$ correspond to completing compulsory education (secondary school), $S_i = 11$ and $T_i \equiv 2$ means that the individual is a high-school graduate, whereas $S_i > 11$ with $T_i \equiv 3$ means university education. Unfortunately, $S_i$ is capped at $17$ years in our sample, making it impossible to distinguish between those who continued to graduate education and those who just finished the $6-$years degree.}. Our measure of ability $Z_i$ is constructed as a CES aggregator of Mathematics and Spanish language test scores. To apply Theorem (ref), we then split $Z_i$ into deciles. Formally, we have
We first test whether cMIV is a reasonable assumption in our setup by implementing the test developed in Section (ref). To that end, we use the parameters and kernel functions recommended by chetverikov_2019 and focus on the theoretically most powerful procedure, the step-down approach. The estimated $p$-value of the test is $0.29$, see Table (ref). We thus conclude that cMIV-p is a credible assumption provided that MIV holds.
The data we study is rather noisy. One would expect a considerable measurement error in the construction of both treatment levels and the outcome variable\footnote{In particular, age is self-reported when filling an online questionnaire and appears to be of low quality, so we are forced to merge multiple cohorts.}. In line with that, the strongest form of cMIV is not sufficient to provide identification in the absence of further assumptions. While the resulting bounds are tighter than that under MIV, they remain uninformative.
To achieve identification, we supplement our assumptions with the MTR condition, which appears reasonable in our setting.\footnote{See MP2000 for a discussion of MTR in the context of returns to education. From the results in Section (ref) it also follows that MTR need only hold conditionally on all pairs in \(\mathcal{T} \times \mathcal{Z}\)—it does not need to hold almost surely.} The estimated bounds and confidence intervals are presented in Figure (ref). While MIV and cMIV-w remain uninformative, both cMIV-p and cMIV-s yield positive lower bounds on the ATEs. Under cMIV-p, the effect of obtaining a university education is estimated to be at least \(3.54\%\), and at least \(5.5\%\) under cMIV-s. These findings align with previous evidence: causal estimates for the U.S. (card1993using, propensity_score, angrist2011schooling) suggest returns of at least \(10\%\) for a four-year college degree. Recent evidence indicates that this figure may be substantially lower for Colombia gomez2022returns.
We also find significantly positive effects at other education stages, see Figure (ref). Further results and details on the estimation procedure are available in Appendix (ref).
\selectlanguage{english}
\nocite{*}
\addcontentsline{toc}{section}{6 References.}