EconBase
← Back to paper

Linear programming approach to partially identified econometric models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

141,229 characters · 25 sections · 113 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Linear programming approach to partially identified econometric models

\doparttoc \faketableofcontents \parttoc

abstractSharp bounds on partially identified parameters are often given by the values of linear programs (LPs). This paper introduces a novel estimator of the LP value. Unlike existing procedures, our estimator is $\sqrt{n}$-consistent, pointwise in the probability measure, whenever the population LP is feasible and finite. Our estimator is valid under point-identification, over-identifying constraints, and solution multiplicity. Turning to uniformity properties, we prove that the LP value cannot be uniformly consistently estimated without restricting the set of possible distributions. We then show that our estimator achieves uniform consistency under a condition that is minimal for the existence of any such estimator. We obtain computationally efficient, asymptotically normal inference procedure with exact asymptotic coverage at any fixed probability measure. To complement our estimation results, we derive LP sharp bounds in a general identification setting. We apply our findings to estimating returns to education. To that end, we propose the conditionally monotone IV assumption (cMIV) that tightens the classical monotone IV (MIV) bounds and is testable under a mild regularity condition. Under cMIV, university education in Colombia is shown to increase the average wage by at least $5.5\%$, whereas classical conditions fail to yield an informative bound. \begin{keyword} partial identification, linear programming, bounds estimation, stochastic programming, uniform estimation, returns to education. \end{keyword}

Introduction

In many partially identified models, the sharp bounds on parameters correspond to the values of linear programs (LPs) that depend on identified functionals of the underlying probability measure. Examples include conditional moment inequalities andrews2023, generalized IV models santos, revealed preference restrictions klinetartari, intersection bounds honoreadriana2006, dynamic discrete choice panels honoretamer2006 and shape restrictions MP2000.

In these settings, the bounds take the form $B(\mathbb{P}) = B(\theta_0(\mathbb{P}))$, where

align[align omitted — 87 chars of source]

and \( \theta_0(\mathbb{P}) \) is the true value of parameter $\theta = (p', \text{vec}(M)', c')'$, estimated via a \( \sqrt{n} \)-consistent estimator \( \hat{\theta}_n \). However, optimization problem (ref) exhibits non-regular behavior, particularly when the underlying model is rich enough that some linear functionals of \( x \) are nearly or exactly point-identified over \( \Theta_I = \{x \in \mathbb{R}^d:Mx \geq c\}\). In such cases, existing estimators of the LP value $B(\mathbb{P})$ are either inconsistent or rate-conservative, creating an undesirable tradeoff in empirical work: richer models provide tighter bounds but complicate their estimation.

To address this issue, we develop a novel debiased penalty function estimator of $B(\mathbb{P})$. Only assuming that the true polytope $\Theta_I$ is non-empty and contained in a known compact set, we show that our estimator is $\sqrt{n}-$consistent, pointwise in the probability measure. In contrast, the plug-in estimator is not generally consistent and may fail to exist with non-vanishing probability, while the alternative set-expansion estimator based on CHT is rate-conservative and may fail to exist in finite samples. Figure (ref) gives a preview of the comparative performance of these estimators.

figure[figure omitted — 506 chars of source]

We obtain an asymptotically normal version of our estimator via sample-splitting and construct confidence regions with exact asymptotic coverage at any fixed probability measure. By comparison, existing procedures either rely on further conditions gafarov2024simple, or result in asymptotically conservative inference chorussel2023. Notably, the approach most commonly used in applied research—combining plug-in estimation with bootstrap (de2017effect, mentalhealth, siddique, kreider2012identifying, pepper, blundell2007changes)—may not provide valid confidence intervals even when the underlying model is far from point-identification. In addition, our procedure is the most computationally efficient in the existing literature\footnote{Both gafarov2024simple (BG) and chorussel2023 (CR) rely on resampling methods, which require to compute one or multiple LPs at each bootstrap iteration. Computing a confidence interval for a LP with 32 variables takes 16.81 seconds with the approach of BG and 40.65 seconds with the approach of CR, according to the latter work. Our approach requires solving a LP once, which takes around 0.0022 seconds on average. The LPs in our application have 160 variables.}.

Turning to uniform asymptotic theory, we first establish a general impossibility result: using Le Cam's binary testing method, we show that no uniformly consistent estimator exists when the estimated functional is discontinuous in the total variation norm. This result implies that the LP value cannot be uniformly consistently estimated over the unrestricted set of distributions $\mathcal{P}$. To make progress, we introduce the `$\delta-$condition' that parametrizes $\mathcal{P}$ by restricting it to the measures at which the smallest singular value of some full-rank submatrix of constraints binding at an optimal vertex is lower-bounded by a $\delta > 0$. This condition is minimal in the sense that any measure from $\mathcal{P}$ satisfies it for some $\delta$, ensuring the family of restricted measures' sets covers $\mathcal{P}$ as $\delta$ grows small. Unlike the conditions in gafarov2024simple, it does not exclude economically relevant problematic cases, such as point-identification and over-identification, nor does it preclude solution multiplicity. Under the $\delta-$condition, our estimator is shown to be uniformly consistent.

To complement our estimation procedure, we derive sharp (and novel) LP bounds for a broad class of causal parameters under affine inequalities over conditional moments\footnote{Including ATE and CATE, among other typically studied parameters, see Section (ref).} (AICM), potentially augmented with affine almost sure restrictions and missing data conditions. In the simplest case, AICM identifying restrictions have the form

align[align omitted — 166 chars of source]

where $(Y(d))_{d}$ are continuous potential outcomes corresponding to the legs of treatment $T$ and $Z \in \mathbb{R}^{d_Z}$ are other covariates. Identified matrices $M^*, \tilde{M}$ and vectors $b^*, \tilde{b}$ are chosen by the researcher. In AICM models, $\theta$ from (ref) is usually a function of identified conditional moments $(\mathbb{E}[Y|T = t,Z =z])_{t,z}$ and the identified joint distribution of $T,Z$, while $x$ collects relevant unobserved conditional moments. Our approach accommodates arbitrary combinations of existing `nonparametric bounds' restrictions, allows to conduct sensitivity analysis, and extends to more complex conditions where sharp bounds were previously unavailable\footnote{For example, cMIV and the mixture of all classical MP2000 conditions, see below.}.

Finally, we apply our approach to estimating returns to education in Colombia. To that end, we first introduce a family of conditionally monotone instrumental variables assumptions (cMIV), which are nested within (ref). A variable is a cMIV if the potential outcomes are mean-monotone in it both unconditionally and within selected treatment subgroups\footnote{The collection of treatment subgroups over which monotonicity is assumed is chosen by the researcher, see Section (ref) for details.}. While an explicit form for the sharp bounds under cMIV may not be feasible, their LP representation follows from our general identification result for (ref). The cMIV conditions yield tighter bounds than the classical MIV assumption of MP2000. We argue, however, that they remain unrestrictive in many applications, including ours. While empirical studies (e.g., de2017effect) have assessed the monotonicity of observed conditional moments to justify applying MIV, such monotonicity is instead equivalent to a particular form of cMIV given that MIV holds and under a mild regularity condition. The formal test of cMIV is obtained as an extension of chetverikov_2019. Using Saber test scores as a cMIV, we find that earning a university degree increases average wages by at least $5.5\%$ in Colombia. In contrast, the classical conditions fail to produce an informative bound.

This paper also contributes two auxiliary results. The first one is concerned with an important special case of (ref) - the combination of all classical MP2000 conditions. Since this combination has the strongest identifying power among classical restrictions, it has been used in empirical work even without a formal justification, sometimes leading to incorrect bounds (see laffers). We provide sharp bounds under continuous outcomes in this setting. Another auxiliary contribution is a novel lower bound on the $\ell_1$-deviation from a non-empty bounded polytope in terms of Euclidean distance from the polytope. It may offer insights into the behavior of $\ell_1$-penalized solutions of systems of linear inequalities studied in the control theory literature (e.g. pinar19991).

We briefly note the limitations of our approach. On the identification side, the absence of restrictions on treatment selection prevents us from studying more granular parameters, such as marginal treatment responses. Furthermore, our identification results are given for discrete treatment and instrument. An extension to the continuous case is feasible, but is outside the scope of this paper\footnote{Even when continuous identification results are available, in practice estimation is still carried out with discretized covariates. This is true for all empirical work referenced below.}. On the estimation side, while our estimator is pointwise $\sqrt{n}-$consistent in general, we only establish $\sqrt{n}/w_n$-uniform consistency\footnote{We show, however, that one side of the convergence happens at the uniform rate $\sqrt{n}$, see Section (ref).} for a slowly diverging sequence $w_n$\footnote{Theoretically, $w_n$ can diverge arbitrarily slowly. We use $w_n \propto \ln \ln n$, see Section (ref) and Appendix (ref).}. We provide further evidence on the uniform rate of consistency in Appendix (ref). A theoretically $\sqrt{n}-$uniformly consistent estimator follows from our analysis, but it depends on an unobserved parameter $\delta$ that is difficult to estimate, so we do not recommend using it in practice. Finally, while our inference procedure naturally extends to uniform setup under sufficient regularity conditions, exploring this is left for future work.

Relationship to literature

The strand of literature relevant to the estimation of (ref) is concerned with statistical inference in the LP estimation framework. semenova2023adaptive considers a LP with an estimated constraint vector $\hat{c}_n$ but known $M$ and $p$, while bhattacharya2009inferring considers LPs with an estimated $\hat{p}_n$ and known $M, c$. Methods developed under a known $M$ assumption do not easily extend to the setting when $M$ is estimated. santos construct a set-expansion estimator and prove its consistency. gafarov2024simple develops uniform inference for a LP described by affine inequalities over unconditional moments, provided uniform Linear Independence Constraint Qualification (LICQ) and Slater's condition (SC) hold. Gafarov's conditions may be restrictive in some applications - for example, under AICM, see Section (ref). chorussel2023 obtain uniformly valid, yet conservative inference for the case when $\theta$ is affine in unconditional moments. Their practical procedure implicitly assumes that the SC holds\footnote{We discuss this in more detail in Section (ref).}. andrews2023 develop an inference procedure for the LP value in a special case in which SC holds. syrgkanis2017inference develop a testing procedure for the failure of LP feasibility. We conduct Monte Carlo simulations to compare our approach with relevant existing methods in Section (ref).

Despite the similar name, AICM approach is unrelated to the model in andrews2023 beyond producing LP bounds. Instead, it generalizes the nonparametric bounds analysis of MP2000, MP2009. The LP sharp bounds in Theorem (ref) coincide with or tighten the bounds in blundell2007changes, boes2009, siddique, kreider2012identifying, de2017effect and mentalhealth. These studies combine the plug-in estimator $B(\hat{\theta}_n)$ with bootstrap for inference, an approach that relies on strong assumptions (see Section (ref)). The exact inference procedure in Algorithm (ref) could be used instead. AICM also complements the findings of santos, who develop identification theory for generalized IV estimators and obtain bounds in the form (ref). Their method imposes a heckman1999local,heckman2005structural selection mechanism in the binary treatment case and thus restricts them to valid IVs\footnote{Additive separability in treatment selection is equivalent to the imbensangrist IV conditions under instrument exogeneity vytlacil2002independence.}, but allows to accommodate arbitrary a.s. restrictions on the marginal treatment response functions and derive bounds for potentially more granular causal parameters. Even though (ref) nests mean-independence conditions, AICM approach is most useful when a valid IV is not available\footnote{Thus, if one is faced with i) a binary treatment setup, ii) has a valid IV and iii) no outcomes' data is missing, the method of santos may be used. If any of these conditions fail, AICM is an alternative.}.

Notation

All vectors are column vectors, and $M'$ denotes the transpose of $M \in \mathbb{R}^{n\times m}$. If $A$ is a set, $A'$ stands for its complement. A collection $(x_j)_{j \in J}$ is a column vector. $2^A$ denotes the powerset of set $A$, and $[n]$ is the collection of integers from $1$ to $n \in \mathbb{N}$. $\times$ is a Cartesian product of sets, while $\otimes$ is the Kronecker product. The sign $\sqcup$ denotes a disjoint union. Signs $\land$ and $\lor$ stand for logical `and' and `or' operators respectively. For $M \in \mathbb{R}^{m\times n}$ and $A \subseteq [m]$, $M_A \in \mathbb{R}^{|A| \times n}$ is the submatrix of the rows of $M$ with indices in $A$. If $j \in [m]$, write $M_j \equiv M'_{\{j\}}$. $\mathcal{R}(M)$ stands for the range of $M$, while $\text{rk}{(M)}$ denotes its rank. $\sigma_d(M)$ is the $d-$th largest singular value of $M$, and $M^\dagger$ denotes the Moore-Penrose pseudoinverse of $M$. In a normed space $S$, the distance between $x \in S$ and $A \subseteq S$ is written as $d(x,A) \equiv \inf_{a \in A} ||x - a||$, and $d_H(A,B) \equiv \max\{\sup_{b \in B} d(b, A), \sup_{a \in A} d(a, B)\}$ is the Hausdorff distance between $A, B \subseteq S$. For $A \subseteq S$ the open expansion is $A^\varepsilon \equiv \{s \in S: d(s, A) < \varepsilon\}$. $Int(A)$ and $Cl(A)$ are the interior and closure of $A \subseteq \mathbb{R}^d$, while $Cone(A)$ is its conical hull. If $A$ is a matrix, $Cone(A)$ is the conical hull of its columns. $s(x, A) \equiv {\max}_{{a \in A}} ~ x'a$ for a compact $A \subseteq \mathbb{R}^d$ and $x \in \mathbb{R}^d$ is a support function. For $v = (v_j)_{j \in [d]}$, define $v^+ \equiv (\max \{v_j, 0\})_{j \in [d]}$. For $v, u \in \mathbb{R}^d$ vector inequalities $v > u$ and $v \geq u$ mean $v_i > u_i ~ \forall i \in [d]$ and $v_i \geq u_i ~ \forall i \in [d]$ respectively. $\iota_d \in \mathbb{R}^d$ is a vector of ones, $I_d \in \mathbb{R}^{d\times d}$ is the identity matrix, and the subscript is dropped occasionally. Operator $\mathbb{E}_\mathbb{P}$ is the expectation under a measure $\mathbb{P}$, and the subscript is dropped whenever it does not cause confusion. The statement $w_n \to \infty$ w.p.a.1 means that $\forall M > 0,$ $ \lim _{n\to\infty} \mathbb{P}[w_n > M] = 1$. We adopt the convention $\inf \emptyset = +\infty$, and $\sup \emptyset = - \infty$.

LP estimation framework

In many partial identification settings, bounds on the parameters of interest can be characterized as the values of linear programs (see review_lp for a review). Readers who prefer to first see an identification framework leading to such bounds may refer to Section (ref), where LP sharp bounds are derived for a general class of AICM models. This section focuses on the estimation theory for such problems. The LP value function is given by

align[align omitted — 81 chars of source]

where $M \in \mathbb{R}^{q\times d}$, $c \in \mathbb{R}^q$ and $p \in \mathbb{R}^d$. The vector $\theta \equiv (p', \text{vec}(M)', c')'$ collects parameters of the LP. The estimable value of these parameters at a fixed true measure is denoted by $\theta_0 \in \mathbb{R}^S$, with $S = qd + q + d$. The value of interest is therefore $B(\theta_0)$. Note that (ref) does not rule out equality constraints, as $Ax = b \iff Ax\geq b \land -Ax \geq -b$.

We denote the constraint set by $\Theta_I(\theta) \equiv \{x \in \mathbb{R}^{d}| Mx \geq c \}$ and omit the argument when $\theta_0$ is concerned. In the context of Section (ref) and other existing applications (e.g. santos), the set $\Theta_I$ is the identified set for an unobserved feature $x$ of the underlying distribution. Under Assumption A0.i-ii, the set $\Theta_I$ is a convex polytope.

namedass[A0 (Pointwise setup)] Suppose that at the fixed true parameter $\theta_0$: i) The identified set is non-empty, $\Theta_I(\theta_0) \ne \emptyset$; ii) $\Theta_I(\theta_0) \subseteq \mathcal{X}$ for a known compact $\mathcal{X} \subseteq \mathbb{R}^d$ and iii) There is an estimator $\hat{\theta}_n \equiv (\hat{p}'_n, \text{vec}(\hat{M}_n)',\hat{c}'_n)'$: $||\hat{\theta}_n - \theta_0|| = O_p(1/\sqrt{n})$

Assumption A0 is maintained throughout this section, while other conditions are stated explicitly. A0.i ensures feasibility of the population LP. In partially identified models it means that there exists a distribution consistent with the identifying restrictions, which does not imply correct model specification. A0.ii is a mild restriction\footnote{For a polytope to be bounded, it must be that $q \geq d+1$, see Chapters 2 and 3 in grunbaum1967convex } that usually holds in applications, for example under bounded outcomes in AICM models (see Section (ref)). The estimator in A0.iii is typically warranted by CLT and the Delta-Method. We focus on the case in which $\theta_0$ is $\sqrt{n}-$estimable for expositional simplicity, but our results generalize to any rate $r_n \uparrow \infty$.

The following primal and dual solution sets will prove useful in our discussion:

align*[align* omitted — 178 chars of source]

Assumption A0 implies that a finite $B(\theta_0)$ is attained as a minimum in (ref), $\Theta_I(\theta_0)$ and $\mathcal{A}(\theta_0)$ are non-empty compact sets, and $\Uplambda(\theta_0)$ is non-empty. We now briefly discuss the typically imposed regularity conditions.

definitionSlater's condition (SC) is the assertion that $\text{Int}(\Theta_I) \ne \emptyset$.\footnote{We give simplified versions of assumptions here for simplicity of exposition. In the presence of `true' equalities $Ax = b$, SC should be stated in terms of relative interior, allowing for point-identification along `true' equalities. LICQ should similarly be restated to account for equalities, as in gafarov2024simple.}

SC rules out point-identification of any linear functional of $x$. In particular, it precludes exact point-identification of the target $p'x$ and point-identification of $x$, i.e. the case when $|\Theta_I| = 1$. Most existing methods rely on SC explicitly (gafarov2024simple) or implicitly (chorussel2023, andrews2023). Even an `approximate' failure of SC, when the true identified set $\Theta_I$ becomes `thin', can be problematic for the existing methods in finite samples. Our simulation evidence illustrates this, see Section (ref). This creates an undesirable tradeoff in applications: higher identification power comes with poorer estimation quality.

definitionLinear independence constraint qualification (LICQ) is the assertion that the submatrix of binding inequality constraints at any $x \in \mathcal{A}(\theta_0)$ is full-rank.

LICQ precludes the existence of overidentifying constraints at the optimum. It may be hard to justify in `bigger' models, like the one we develop and apply in Sections (ref) and (ref). These feature a larger number of inequality constraints that may have similar identifying power, so it is not ex-ante clear why there must not be overidentification at the optimum. LICQ also rules out parameters-on-the-boundary, as the following remark clarifies.

remarkExample 2.1 in santos2019 can be restated as a LP: $B(\theta_0) = \max\{0, \mathbb{E}[X]\} = \min_{t \in \mathbb{R}} t ~ \text{s.t.} ~ t \geq 0, t \geq \mathbb{E}[X]$. LICQ fails in that program if $\mathbb{E}[X] = 0$, which corresponds to the parameter-on-the-boundary case from andrews1999estimation, andrews2000.
definitionNo flat faces condition (NFF) is the assertion that $|\mathcal{A}(\theta_0)| = 1$.

The notion of flat faces thus corresponds to $|\mathcal{A}(\theta_0)| \ne 1$, i.e. the situation in which the bound on the target parameter $p'x$ is achieved at multiple partially identified features $x$.

Assumption A0 does not impose LICQ or SC, nor does it rule out flat faces. Estimating $B(\theta_0)$ without these conditions is challenging due to the irregular behavior of $B(\cdot)$. If SC fails, $B(\cdot)$ may be discontinuous at $\theta_0$, and the plug-in estimator $B(\hat{\theta}_n)$ may not be pointwise consistent.

figure[figure omitted — 2,311 chars of source]
propFix any $d \in \mathbb{N}$. Then, i) for any $q \geq d+2$, there exist $\theta_0$ and $\hat{\theta}_n$ satisfying Assumption A0, with $B(\hat{\theta}_n) - B(\theta_0) \ne o_p(1)$ and $|B(\hat{\theta}_n)| < +\infty$ for all $n \in \mathbb{N}$ surely; and ii) for any $q \geq d+1$, there exist $\theta_0$ and $\hat{\theta}_n$ satisfying Assumption A0, with $B(\hat{\theta}_n) - B(\theta_0) \ne o_p(1)$ and $B(\hat{\theta}_n) = +\infty$ with non-vanishing probability.

The proposition above can be illustrated using two simple examples.

exampleFor the first part of Proposition (ref), consider \begin{align} B(b) = \underset{x \in \mathbb{R}^2}{\min} x_1 \quad s.t.: x_2 \geq (1+b)x_1, x_2 \leq x_1, x_1 \in [-1;1], \end{align} where $b$ is estimated via $\hat{b}_n = b + \frac{1}{n}\sum_{i = 1}^n U_i$ with $U_i \sim U[-1;1]$ i.i.d. Suppose in population $b = 0$, as in Figure (ref). The true value is then $B(0) = -1$, attained at $x^* = - \iota$. The plug-in estimator collapses to $B(\hat{b}_n) = - \mathds{1}\{ \hat{b}_n \geq 0\}$, and does not converge to $-1$ in probability.
exampleTo illustrate the second part of Proposition (ref), consider \begin{align*} B(a) = \underset{x \in \mathbb{R}^2}{\min} x_1 \quad s.t.: x_2 \geq x_1 + a, x_2 \leq x_1, x_1 \in [-1;1], \end{align*} where $a$ is estimated via $\hat{a}_n = a + \frac{1}{n}\sum^n_{i =1}U_i$ with $U_i \sim U[-1;1]$ i.i.d. Suppose $a = 0$. If $\hat{a}_n > 0$, the sample LP is infeasible, so $B(\hat{b}_n) = +\infty$. This occurs w.p. $1/2$ for any $n \in \mathbb{N}$.

In some special cases (e.g., honoretamer2006), SC may be argued to hold, ensuring the continuity of $B(\cdot)$\footnote{Under SC and A0, $B(\cdot)$ is continuous at $\theta_0$, see Appendix (ref).}. However, an additional challenge arises: $B(\cdot)$ is not necessarily Hadamard differentiable unless further regularity conditions hold. This complicates inference, as Proposition (ref) in Section (ref) illustrates.

This section addresses Propositions (ref) and (ref). Section (ref) introduces the penalty function estimator and its debiased version, which we show to be $\sqrt{n}$-pointwise consistent under A0. Section (ref) develops a computationally efficient inference procedure with exact asymptotic coverage under A0. Turning to uniform properties, Section (ref) presents a general impossibility result for discontinuous functionals. It implies that the LP value cannot be uniformly consistently estimated under a uniform version of A0 alone. We then characterize a broad class of measures over which a uniformly consistent estimator exists. Sections (ref) and (ref) establish the uniform rates of the penalty function estimators over this class. Section (ref) provides simulation evidence.

Consistency

Penalty function estimator

We now develop a consistent estimator that is inspired by the theory of exact penalty functions. The idea is to restate (ref) as an unconstrained penalized problem. Define the $L_1-$penalized version of the LP objective as

align*[align* omitted — 61 chars of source]

and consider the unconstrained problem

align[align omitted — 227 chars of source]

We use $\tilde{B}(\cdot)$ to obtain a preliminary estimator of the LP value, which we term the penalty function estimator. Note that $L(x; \theta, w) = p'x$ at any $x \in \Theta_I(\theta)$, i.e. the penalized objective function is equal to the LP objective function whenever the constraints in (ref) are satisfied.

namedass[A1 (Penalty parameter)] The penalty vector $w \in \mathbb{R}^q$ and the true parameter $\theta_0 \in \mathbb{R}^S$ are such that in the problem (ref) evaluated at $\theta_0$ there exists a KKT vector $\lambda^* \in \Uplambda(\theta_0)$, such that $w > \lambda^*$.

Note that Assumption A1 does not require $w$ to be component-wise larger than all KKT vectors. In any finite and feasible LP there exists at least one $\lambda^* < \infty$, so at any fixed $\theta_0$ there always exists a large enough $w$ that satisfies A1.

remarkIf it is known that i) $B(\theta_0) < K$ for some $K > 0$, and ii) $c > \underline{c} > 0$ for some known $\underline{c} > 0$, Assumption A1 is satisfied by $\theta_0$ and known $w = \iota K/\underline{c}$ by duality.

The following Lemma is key to understanding the penalty function approach. It asserts that under Assumption A1 the $L_1-$penalty function is exact for the LP in (ref).

lemmaFor any $\theta \in \mathbb{R}^S$, and $w \in \mathbb{R}^q_{+}$, if $\Theta_I(\theta) \subseteq \mathcal{X}$, then \begin{align} \tilde{B}(\theta;w) \leq B(\theta). \end{align} Moreover, if $\theta_0$ and $w$ satisfy Assumption A1, then: i) (ref) holds with an equality at $\theta = \theta_0$, and ii) solutions coincide, $\tilde{\mathcal{A}}(\theta_0;w) = \mathcal{A}(\theta_0)$.

The deterministic result in Lemma (ref), combined with the observation that the objective function converges in probability uniformly in $x$ under A0, establish that the penalty function estimator with a fixed $w$ is consistent under A1.

propUnder Assumption A1, \begin{align*} \tilde{B}(\hat{\theta}_n;w) \xrightarrow{p} B(\theta_0). \end{align*}

For ease of notation, from now on we treat $w \in \mathbb{R}_+$ as a scalar penalty that induces the penalty vector $w \iota$. Our results extend immediately to the case when the coordinates of the penalty vector differ. Based on Lemma (ref) and Proposition (ref), it might seem that $w$ should be selected to be as large as possible. This, however, yields a generally inconsistent estimator if SC fails.

excont[cont'd] Consider (ref) with $b = 0$ and suppose $w > 2$. If $\hat{b}_n < 0$, there exists a sample KKT vector $\hat{\lambda}_n \in \Uplambda(\hat{b}_n)$, whose largest coordinate is $||\hat{\lambda}_n||_{\infty} = |\hat{b}^{-1}_n|$. So, if also $w > |\hat{b}^{-1}_n|$, the penalty estimator selects an incorrect optimum $(0,0)$ in light of Lemma (ref). Since $\hat{b}_n \xrightarrow{p} 0$, at a large enough sample size $|\hat{b}^{-1}_n|$ will exceed any fixed $w$ with high probability and the correct minimum of $-1$ will be estimated. However, that logic fails in finite samples if $w$ is `large'.

\addtocounter{excont}{-1} This observation justifies the need to study $w \to \infty$ asymptotic theory. We show that the penalty parameter can be allowed to diverge at the rate dominated by $\sqrt{n}$.

theorFor any $w_n \to \infty$ w.p.a.1 with $\frac{w_n}{\sqrt{n}} \xrightarrow{p} 0$, we have \begin{align*} |\tilde{B}_n(\hat{\theta}_n, w_n) - B(\theta_0)| = O_p\left(\frac{w_n}{\sqrt{n}}\right) \end{align*}

Observe that the estimator in Theorem (ref) does not rely on Assumption A1, as the latter is always satisfied at a fixed measure for a large enough $n$ when $w_n \to \infty$.

Debiased penalty estimator

The $w_n n^{-1/2}$ rate of convergence in Theorem (ref) is determined by the slowly vanishing penalty term. This term is a product of the deviation from the true polytope, that vanishes at $n^{-1/2}$, and an exploding sequence $w_n$. It is thus reasonable to ask whether the $\sqrt{n}-$rate could be restored by dropping the penalty term, i.e. debiasing the penalty function. We show that this can be done. Before we proceed, let us make the following simplification. Without loss of generality, suppose that

align[align omitted — 66 chars of source]

To see why (ref) can be assumed w.l.g., note that one can set $p = e_1 = (1 ~ 0 \dots 0)'$ and add an auxiliary variable for the value of the problem in the first position of $x$ (see gafarov2024simple).

We define the debiased penalty function estimator as

align*[align* omitted — 110 chars of source]

The following theorem establihes its rate, and is one of the main contributions of this paper.

theorSuppose $\mathcal{A}(\theta_0) \subseteq Int(\mathcal{X})$. For any $w_n \to \infty$ w.p.a.1 with $\frac{w_n}{\sqrt{n}} \xrightarrow{p} 0$, \begin{align*} \underset{x \in \tilde{\mathcal{A}}(\hat{\theta}_n,w_n)}{\max} \left|p'x - B(\theta_0)\right| = O_p\left(\frac{1}{\sqrt{n}}\right). \end{align*}
remarkThe result in Theorem (ref) is uniform over the argmin set, so one may use any measurable selection from $\tilde{\mathcal{A}}(\hat{\theta}_n;w_n)$ to obtain a $\sqrt{n}-$consistent estimator. In the context of lower/upper bound estimation, $\max/\min_{\tilde{\mathcal{A}}(\hat{\theta}_n;w_n)} p'x$ respectively yield the tightest bound.
remarkThe argmin set of the penalty function estimator can be computed by solving a LP with $d + q$ variables and $2 q$ constraints. Specifically, in Appendix (ref) we show that, w.p.a.1, $\tilde{\mathcal{A}}(\hat{\theta}_n;w_n) = \arg \min_{x, a \in \mathbb{R}^d \times \mathbb{R}^q} p'x + w_n\iota'a, \text{ s.t.: } a\geq 0, a \geq \hat{c}_n - \hat{M}_n x$.
remarkAn alternative estimator can be constructed using a set-expansion argument: $\check{B}_n = \min_{x \in \mathbb{R}^d} \hat{p}_n'x ~ \text{s.t.} ~ \hat{M}_n x \geq \hat{c}_n - \sqrt{\kappa_n} n^{-1/2}\iota $. In Appendix (ref), we show that the results from CHT and the geometry of polytopes imply that $\check{B}_n$ with an appropriately chosen, diverging $\kappa_n$, is consistent for $B(\theta_0)$. However, it can be rate-conservative, converging at $\sqrt{{n}}\kappa_n^{-1/2}$. It appears to perform worse than the debiased penalty function estimator in our simulations, see Section (ref).

We now outline the main ideas behind Theorem (ref). For $x \in \mathbb{R}^d$, let $J(x;\tilde{\theta}) \equiv \{j \in [q]: \tilde{M}_{j}x = \tilde{c}_{j}\}$ denote the set of constraints that bind at $x$ when evaluated at $\tilde{\theta}$. For ease of notation, from now on we write $\theta_0 = (\text{vec}(M)',c')'$. Let us introduce the following terminology.

definition[Vertex] We call $x \in \tilde{\mathcal{A}}(\hat{\theta}_n;w_n)$ a vertex-solution if the corresponding matrix of binding constraints, $\hat{M}_{nJ(x;\hat{\theta}_n)}$, has full column rank.
definition[Nice face] We say that a nonempty set $A \subseteq [q]$ corresponds to a nice face\footnote{It should be noted that a nice face is not necessarily a valid $k-$face of the true polytope $\Theta_I$.} $F \equiv \{x \in \mathbb{R}^d:M_A x = c_A\}$ if $p'x = B(\theta_0)$ for any $x \in F$.

Intuitively, the proof of Theorem (ref) proceeds in two steps. First, by anti-concentration arguments we establish that with high probability asymptotically the penalty function estimator manages to select a vertex-solution $\hat{x}_n$ such that the set of constraints that bind at it, $\hat{A}_n = J(\hat{x}_n;\hat{\theta}_n)$, corresponds to a nice face $F = \{x \in\mathbb{R}^d : M_{\hat{A}_n}x = c_{\hat{A}_n}\}$. Once a nice face has been selected, the $\sqrt{n}-$convergence of $p'\hat{x}_n$ to $B(\theta_0)$ obtains as a consequence of $(\hat{M}_{nA}, \hat{c}_{nA})$ converging to $(M_A, c_A)$ at this rate for a fixed $A \subseteq [q]$.

The discussion of uniform asymptotic theory in Section (ref) sheds light on the role of $w_n$ and the trade-off involved in its selection. The practical guidance on selecting $w_n$ is then developed on the basis of our results and random matrix theory in Appendix (ref).

Inference

This section develops an inference procedure for a general LP estimator, in which all parameters are inferred from the data. This procedure nests special cases in which some parameters remain fixed, as in semenova2023adaptive or bhattacharya2009inferring.

namedass[B0 (Random sample)] Suppose $\hat{\theta}_n = \hat{\theta}_n(\mathcal{D}_n)$ is a measurable function of the sample $\mathcal{D}_n \equiv \{W_1, W_2, \dots, W_n\}$, where $W_i \in \mathbb{R}^{d_W}, ~ i \in [n]$ are i.i.d. random vectors.

We suppose that Assumption B0 holds throughout Section 2.2, whereas the rest of the conditions are imposed explicitly.

namedass[B1 (Asymptotic normality)] The estimator $\hat{\theta}_n$ is such that, for $\Sigma < \infty$, \begin{align*} \sqrt{n}(\hat{\theta}_n - \theta_0) \xrightarrow{L} \mathbb{G}_0 \sim \mathcal{N}(0, \Sigma). \end{align*}

Assumption B1 is typically warranted by reference to CLT and the Delta Method when $\hat{\theta}_n = g(n^{-1}\sum^n_{i=1}W_i)$ for some smooth $g(\cdot)$, as in the AICM models (ref), see Section (ref).

Exact inference on a debiased estimator

We construct a method for statistical inference on \( B(\theta_0) \) that achieves exact asymptotic coverage under minimal regularity conditions. Our approach is based on an asymptotically normal version of the debiased penalty estimator with \( w_n \to \infty \) w.p.a.1 and is outlined in Algorithm (ref). Before stating the main result, we introduce key auxiliary constructions and discuss our assumptions.

algorithm[algorithm omitted — 2,549 chars of source]

For the true $\theta_0 = (\text{vec}(M)',c')'$ and some subset of indices $A \subseteq [q]$, consider conditions

align[align omitted — 138 chars of source]

Equation (ref) is satisfied if constraints $A$ may bind simultaneously at some solutions of the original LP, while (ref) holds if the objective function's gradient $p$ is a linear combination of the gradients of inequalities from $A$. For example, if $A$ is a set of all binding constraints at some $x \in \mathcal{A}(\theta_0)$, equation (ref) follows from KKT conditions.

Continuing the discussion in Section (ref), we note that subsets $A$ that satisfy (ref) and (ref) correspond to the nice faces.

lemmaIf $A \subseteq [q]$ satisfies (ref) and (ref), then $F = \{x \in \mathbb{R}^d: M_Ax = c_A\}$ is a nice face, \begin{align*} B(\theta_0) = p'x, \quad \forall x \in F. \end{align*}

Define the set $\mathbb{A} \equiv \{A \in 2^{[q]}:|A| \geq d, A \text{ satisfies \eqref{optimality} and \eqref{range_req}}\}$. It is non-empty in any feasible finite LP. With probability approaching $1$, the penalty function estimator manages to select a vertex-solution $\hat{x}_n \in \tilde{\mathcal{A}}(\hat{\theta}_n; w_n)$, determined by the binding constraints $\hat{A}_n = J(\hat{x}_n;\hat{\theta}_n)\in \mathbb{A}$ that satisfy (ref) and (ref) and thus correspond to a nice face by Lemma (ref).

The debiased estimator may hence be understood as a two-stage procedure: one first finds the set of binding inequalities $\hat{A}_n \in \mathbb{A}$, and then estimates $B(\theta_0)$ as $p'\hat{M}_{n\hat{A}_n}^{\dagger}\hat{c}_{n\hat{A}_n}$. Performing inference on that object directly would require working with the complex joint distribution of $\hat{A}_n$, and $\hat{\theta}_n$, and would likely result in an asymptotically non-normal estimator.

We address this by `disentangling' the variation in \(\hat{A}_n\) and \(\hat{\theta}_n\) via sample splitting. Intuitively, a vertex is estimated on one part of the sample, while the noise in the parameter estimation comes from the other. We now state our assumptions and present the main result.

namedass[B2 (Variance estimator)] B1 holds, and there exists an estimator $\hat{\Sigma}_n \xrightarrow{p} \Sigma$.

Assumption B2 requires the researcher to possess a consistent estimator of the asymptotic variance of $\hat{\theta}_n$. If $\hat{\theta}_n = g\left(n^{-1}\sum_{i = 1}^n W_i\right)$ for some smooth and known $g(\cdot)$, such estimator can typically be obtained from the estimated covariance matrix of $W_i$ via Delta-method. In more complicated scenarios, bootstrap on $\hat{\theta}_n$ may be employed.

Define the set $\mathcal{S}_{A} \equiv \{v \in \mathbb{R}^{|A|}: p = M'_{A} v\}$ and note that (ref) is equivalent to $\mathcal{S}_A \ne \emptyset$.

namedass[B3] For a constant $\overline{v} > 0$, $\underset{A \in \mathbb{A}}{\max} \underset{v \in \mathcal{S}_A}{\min} ||v||\leq \overline{v}$.

Assumption B3 is a technical condition ensuring that we can find a sequence approaching $\mathcal{S}_A$ asymptotically, i.e. $d(\check{v},\mathcal{S}_A) = o_p(1)$ for $\check{v}$ defined in (ref). Practical guidance on choosing $\overline{v}$ is provided in Appendix (ref). Simulation evidence in Figure (ref) in Appendix (ref) suggests that the specific value of \(\overline{v}\) has no impact on inference, as long as it is sufficiently large.

definition[Optimal triplet] We call $(A, x, v) \in 2^{[q]}\times \mathbb{R}^d \times \mathbb{R}^q$ an optimal triplet if i) $|A| \geq d$, ii) $x \in \mathcal{A}(\theta_0)$, iii) $M_Ax = c_A$, iv) $p = M_A'v_A$, and v) $A = \text{Supp}(v)$.

We randomly split $\mathcal{D}_n$ into two disjoint, collectively exhaustive folds $\mathcal{D}^{(f)}_{n}$ of size $n_f$ for $f = 1, 2$, with $n_1 = \lfloor\gamma n\rfloor$ and $n_2 = n - \lfloor \gamma n\rfloor$ for some fixed $\gamma \in (0;1)$. Our inference procedure uses the data from $\mathcal{D}^{(1)}$ to estimate an optimal triplet $(\hat{A},\hat{x},\hat{v})$. The vertex\footnote{While we assume that $\tilde{\mathcal{A}}(\hat{\theta}^{(1)};w_{n_1})$ is estimated precisely, the results do not change if one is only able to estimate a single optimum. This may occur if numerical errors do not allow the LP-solver to find all of the LP solutions. Such optimum will satisfy $\text{rk}(\hat{M}_{J(x;\hat{\theta}^{(1)})}) = d$ by definition, and so will be a valid vertex-solution.} is estimated as

align*[align* omitted — 177 chars of source]

the set of binding constraints that define it is denoted by $\hat{A} \equiv J(\hat{x};\hat{\theta}^{(1)})$, and

align[align omitted — 125 chars of source]

Finally, we define $\hat{v} \in \mathbb{R}^{q}$ so that $\hat{v}_{\hat{A}} = \check{v}$ and $\hat{v}_j = 0$ for $j \notin \hat{A}$.

Our procedure is then based on showing that, for large $n$,

align*[align* omitted — 208 chars of source]

where $\sigma^2(\cdot)$ is derived in Appendix (ref).

namedass[B4 (Non-degeneracy)] B1 holds, and $\sigma(A, x, v, \Sigma) > 0$ for any optimal triplet $A, x, v$.

An inspection of the proof of Theorem (ref) below reveals that Assumption B4 rules out the scenarios when finding an $A \in \mathbb{A}$ determines the value $B(\theta_0)$, even though $\hat{\theta}$ is noisy. This may occur, for example, if the corresponding $\hat{c}_A, \hat{M}_A$ are deterministic. In this case, if $A$ is also unique, meaning $|\mathbb{A}| = 1$, the debiased estimator has $0$ asymptotic variance, because $A$ and therefore $B(\theta_0)$ are correctly estimated with probability approaching $1$.

theorSuppose $\mathcal{A}(\theta_0) \subseteq \text{Int}(\mathcal{X})$ and Assumptions B1, B3, B4 hold. Moreover, \begin{align*} \hat{\sigma}_n(A, x, v) \xrightarrow{p} \sigma(A, x, v, \Sigma) \end{align*} for any optimal triplet $(A, x, v)$ with $||v|| \leq \overline{v}$, which holds for $\hat{\sigma}_n(A, x, v) = \sigma(A, x, v, \hat{\Sigma}_n)$ under Assumption B2. Then, for any $\alpha \in (0;1)$, and any $w_n \to \infty$ w.p.a.1 such that $w_n = o_p(\sqrt{n})$, \begin{align*} \mathbb{P}\left[\frac{\sqrt{n_2}}{\hat{\sigma}_n (\hat{A},\hat{v},\hat{x})} \left(\check{v}'(\hat{c}^{(2)}_{\hat{A}} - \hat{M}^{(2)}_{\hat{A}} \hat{x}) + p'\hat{x} - B(\theta_0)\right) \leq z_{1-\alpha}\right] = 1-\alpha + o(1). \end{align*}
remarkFollowing gafarov2024simple, one can drop Assumption B4 by using $\max \{\hat{\sigma}_n(\cdot), \underline{\sigma}\}$ for some small $\underline{\sigma} > 0$ instead of $\hat{\sigma}_n(\cdot)$ in Theorem (ref). In that case, the test controls level, but may have a conservative size.
remarkIf an estimator $\hat{\Sigma}_n$ is not available and is not obtained via bootstrap on $\hat{\theta}_n$, one may alternatively construct confidence intervals using the quantiles of $\sqrt{n}_2 (\breve{B} - B(\theta_0))$. These can be computed via bootstrap on $\sqrt{n}_2\hat{v}'_{\hat{A}}(\hat{c}_{\hat{A}}^{(2)} - c_{\hat{A}} - (\hat{M}^{(2)}_{\hat{A}} - M_{\hat{A}})\hat{x})$, where $\hat{A}, \hat{x},\hat{v}$ remain fixed, while $\hat{c}^{(2)}$ and $\hat{M}^{(2)}$ are bootstrapped over the second fold data $\mathcal{D}^{(2)}$.

Bootstrapping the plug-in fails even under SC

To further justify the need for our inferential procedure, we examine the properties of the approach that combines bootstrap on $\hat{\theta}_n$ with the plug-in estimator $B(\hat{\theta}_n)$. This method is widely used in empirical literature applying AICM conditions (blundell2007changes, kreider2012identifying, pepper, siddique, de2017effect, and mentalhealth). In light of Proposition (ref), this approach is inapplicable when SC fails. In practice, researchers attempting to apply it to a LP with a small or empty interior of $\Theta_I$ may encounter frequent LP infeasibility in the bootstrap draws.

In some cases, SC may be established. This is true, for example, if the bound of interest can be expessed as an intersection bound $B(\theta_0) = \max \{c_1,c_2, \dots, c_q\} = \min_{t \in \mathbb{R}} t ~~ \text{s.t. } t \geq c_i, ~ i \in [q]$, where $\theta_0 = (1 ~ \iota' ~ c')'$ (as in CLR). Yet, even if SC holds, we demonstrate that bootstrap inference based on the plug-in estimator is not valid, unless further regularity conditions hold. This observation can be viewed as a generalization of the parameter-on-the-boundary problem of andrews1999estimation,andrews2000.

definitionLet $\mathbb{D}$ and $\mathbb{E}$ be Banach spaces, and $f: \mathbb{D}_f \subseteq \mathbb{D} \to \mathbb{E}$. The map $f$ is said to be Hadamard directionally differentiable (H.d.d.) at $\upsilon \in \mathbb{D}_f$ tangentially to $\mathbb{D}_0 \subseteq \mathbb{D}$, if there is a continuous map $f'_\upsilon: \mathbb{D}_0 \to \mathbb{E}$, such that \begin{align*} \underset{n \to \infty}{\lim}\left \lVert \frac{f(\upsilon + t_n h_n) - f(\upsilon)}{t_n} - f'_{\upsilon}(h)\right \rVert_{\mathbb{E}} = 0, \end{align*} for all sequences $\{h_n\} \subset \mathbb{D}$ and $\{t_n\} \subset \mathbb{R}_+$ such that $t_n \to 0^+$, $h_n \to h \in \mathbb{D}_0$ as $n \to \infty$ and $\upsilon + t_n h_n \in \mathbb{D}_f$ for all $n$. If, moreover, $f'_v(h)$ is linear in $h$, the map $f$ is said to be fully Hadamard differentiable at $\upsilon$.

For simplicity of exposition, we abstract from the case of `true equalities' in $\Theta_I$. The results extend trivially to this case if SC is defined in terms of relative interior.

lemma[duan2020hadamard] Under SC, $B(\cdot)$ is Hadamard directionally differentiable at $\theta_0$. The directional derivative is given by \begin{align} B'_{\theta_0}(h) = \underset{x \in \mathcal{A}(\theta_0)}{\inf} \underset{\lambda \in \Uplambda(\theta_0)}{\sup} h'_p x + \sum^Q_{i = 1} \lambda_{i}(h_{c_i} - h'_{M_i} x), \end{align} where $h = (h'_p, h'_{M_{1}},\dots, h'_{M_{q}}, h_{c_1},\dots,h_{c_q})'$ is the direction of the increment in $\theta$.

Hadamard directional differentiability of $B(\cdot)$ is sufficient for convergence in law.

propUnder SC and Assumption B1, it follows that \begin{align*} \sqrt{n}(B(\hat{\theta}_n) - B(\theta_0)) \xrightarrow{L} B'_{\theta_0}(\mathbb{G}_0) \end{align*}
proofsantos2019 Theorem 2.1. combined with Lemma (ref).

Gaussianity of $B'_{\theta_0}(\mathbb{G}_0)$ is a necessary condition for bootstrap consistency santos2019. Consequently, the empirical literature using bootstrap with the plug-in estimator has implicitly relied on this assumption. However, $B'_{\theta_0}(\mathbb{G}_0)$ is not normal unless full Hadamard differentiability holds, i.e. $B'_{\theta_0}(h)$ is linear in $h$. As (ref) suggests, this is not generally the case. Theorem 3.1 in santos2019 establishes that bootstrap is inconsistent for the distribution when $B'_{\theta_0}(h)$ fails to be linear. The typically applied plug-in and bootstrap combination is then only valid under further restrictive assumptions\footnote{The full-support condition in Proposition (ref) is imposed for expositional purposes. Sufficiency of conditions i, ii holds generally, whereas necessity obtains whenever the derivative $B_{\theta_0}'(h)$ is not linear over $\text{Supp}(\mathbb{G}_0)$ when $\mathcal{A}(\theta_0), \Uplambda(\theta_0)$ are not singletons, see Lemma (ref) in Appendix.}:

propIf SC and Assumption B1 hold, $\text{Supp}(\mathbb{G}_0) = \mathbb{R}^S$ and $\theta^*_n$ satisfies Assumption 3 in santos2019, bootstrap is consistent in the sense that \begin{align*} \underset{f \in BL_1}{\sup} \left|\mathbb{E}[f(\sqrt{n}(B(\theta^{*}_n) - B(\hat{\theta}_n)))|\mathcal{D}_n] - \mathbb{E}[f(B'_{\theta_0}(\mathbb{G}_0))]\right| = o_p(1), \ \end{align*} where $BL_1 \equiv \{f:\mathbb{R}\to\mathbb{R} \text{ s.t. }|f(a)|\leq 1, |f(a) - f(b)|~ \leq |a - b| ~ \forall a, b \in \mathbb{R}\}$, if and only if i) NFF and ii) SMFCQ also hold at $\theta_0$.
remarkStrict Mangasarian-Fromovitz Constraint Qualification (SMFCQ) that we define in Appendix (ref) is equivalent to $|\Uplambda(\theta_0)| = 1$, which is necessary for bootstrap consistency. However, it depends on both the vector $p$ and on $\Theta_I$. In Appendix (ref), we show that \( |\Uplambda(\cdot)| = 1 \) uniformly over all objective functions minimized at some \( x \in \Theta_I \) if and only if (pointwise) LICQ holds at \( x \). Both SMFCQ and LICQ are high-level and may be restrictive in models with potentially overidentifying constraints, as in Sections (ref) and (ref).
remarkA consistent estimator for the distribution of $B(\hat{\theta}_n)$ under SC can be obtained by combining the Functional Delta Method (FDM) of santos2019 with the Numerical Delta Method (NDM) given in hong2015numerical, see Appendix (ref). In Appendix (ref), we also show that the penalty function estimator is H.d.d. in $\theta$ , so FDM + NDM combination yields exact inference for it. This approach relies on an arbitrarily selected NDM step size and a fixed $w$, and so does not appear satisfactory. Finally, the set-expansion estimator enforces SC and has a Lipschitz-bounded bias (see Appendix (ref)). A conservative inference procedure based on it can then also be obtained via FDM and NDM.

Impossibility result

The optimization problem (ref) is challenging to study under no further assumptions, as it may exhibit instability under arbitrary perturbations of parameters. We now show that this not only leads the plug-in $B(\hat{\theta}_n)$ to fail pointwise, but also precludes the existence of a uniformly consistent LP estimator over the unrestricted set of measures.

We first establish a new impossibility result that complements the findings of hirano. Specifically, we show that no uniformly consistent estimator exists for any discontinuous functional from the space of probability measures endowed with the total variation norm to an arbitrary metric space $(\mathcal{V}, \rho)$.

lemmaSuppose a functional $V: (\mathcal{P}, ||\cdot||_{TV}) \to (\mathcal{V}, \rho)$ is discontinuous at $\mathbb{P}_0 \in \mathcal{P}$. Then, there exists no uniformly consistent estimator $\hat{V}_n = \hat{V}_n(X)$, which is a sequence of measurable functions of the data $X \sim \mathbb{P}^n$. Moreover, if $\varepsilon > 0$ is a lower bound on the discontinuity, i.e. for any $\delta > 0$ there exists $\mathbb{P}_1 \in \mathcal{P}$ such that $||\mathbb{P}_0 - \mathbb{P}_1||_{TV}< \delta$ and $\rho(V(\mathbb{P}_0), V(\mathbb{P}_1)) > \varepsilon$, then \begin{align*} \underset{\hat{V}_n}{\inf }\underset{\mathbb{P} \in \mathcal{P}}{\sup } \mathbb{E}_\mathbb{P}\left[\rho\left(V(\mathbb{P}), \hat{V}_n(X(\mathbb{P}^n))\right)\right] \geq \frac{\varepsilon}{4}, \quad \forall n \in \mathbb{N}, \end{align*} where infinum is taken over all measurable functions of the data.

In this section, we treat the parameter $\theta_0$ as a functional of the underlying probability measure $\mathbb{P} \in \mathcal{P}$. We then make the following assumption on the pair $\theta_0(\cdot), \mathcal{P}$:

namedass[U0 (Uniform setup)] The functional $\theta_0(\cdot)$ and the set of probability measures $\mathcal{P}$ are such that: i) $\theta_0: (\mathcal{P}, ||\cdot||_{TV}) \to (\mathbb{R}^S, ||\cdot||_2)$ is continuous; ii) $\theta_0(\mathcal{P}) = \{ y \in \mathbb{R}^S \text{ s.t. } \Theta_I(y) \ne \emptyset, \Theta_I(y) \subseteq \mathcal{X}\}$ for a known and fixed compact $\mathcal{X}$

Assumption U0 formalizes the notion of the unrestricted set of measures. U0.i demands that the true parameter be continuous in $P$, which holds, for example, in AICM models (see Example (ref)). U0.ii assumes that $\theta_0$ has full support over $\mathcal{P}$, meaning that any $\theta$ corresponding to a consistent model with $\Theta_I(\theta) \subseteq \mathcal{X}$ is attained at some $P \in \mathcal{P}$.

theorUnder U0, there exists no uniformly consistent estimator of $B(\theta_0)$.
proofCombining U0, Lemma (ref) and Proposition (ref).

Given this negative result, it is natural to seek a minimal restriction on $\mathcal{P}$ for which a uniformly consistent estimator may exist. We now show that the condition ensuring uniform consistency of the penalty function approach can be considered minimal in the sense to be made precise in Proposition (ref) .

Our examination of Example (ref) has shown that $w_n$ cannot be allowed to diverge faster than $\sqrt{n}$, as otherwise the penalty approach may fail at measures where SC fails. At the same time, if all KKT vectors $\lambda^* \in \Uplambda(\theta_0)$ grow large, an arbitrarily large $w$ is needed for Assumption A1 to hold. This occurs when optimal vertices become `sharp', i.e. all relevant full-rank submatrices of binding inequality constraints grow closer to being degenerate. The condition that ensures uniform consistency of the penalty function approach should therefore bound such `sharpness'.

We begin our construction with an existence result based on the Caratheodory’s Conical Hull Theorem.

propThe problem (ref) admits a solution $x^*$ and the associated KKT vector $\lambda^*$ such that for some index subset $J^{*} \subseteq \{1, \dots, q\}$ with $|J^{*}| = d$, $M_{J^{*}}$ is invertible and: \begin{align*} &x^* = M_{J^{*}}^{-1}c_{J^{*}},\\ &\lambda^{*}_{J^{*}} = M_{J^{*}}^{-1 \prime}p,\\ &\lambda^{*}_i = 0 if i \notin J^{*}. \end{align*}

Proposition (ref) asserts that any finite and feasible LP has an optimal vertex $x^*$ at which there is a subset $J^*$ of binding constraints, such that i) the corresponding gradients form a full-rank square matrix, and ii) the objective function gradient belongs to the conical hull formed by the gradients of the constraints from $J^*$.

namedass[U1 ($\delta$-condition)] The class of measures $\overline{\mathcal{P}}$ satisfies the $\delta-$condition for a given $\delta > 0$, if \begin{align} \underset{\mathbb{P} \in \overline{\mathcal{P}}}{\inf} \underset{J^{*} \in \mathcal{J}^*(\theta(\mathbb{P}))}{\max} \sigma_d(M_{J^{*}}(\mathbb{P})) > \delta, \end{align} where $\mathcal{J}^*(\theta(\mathbb{P}))$ collects all $J^{*}$ defined in Proposition (ref) at a given $\theta(\mathbb{P})$.

The $\delta$-condition does not rule out the failure of LICQ, SC or NFF, and is weaker than the conditions under which uniform consistency of LP estimators has been established. To formalize this, let us introduce three families of measures. Firstly, denote the family of measures satisfying U1 for a given $\delta > 0$ by $\mathcal{P}^\delta$. A measure satisfies the Slater's condition if $\mathbb{P} \in \mathcal{P}^{SC} \equiv \{\mathbb{P} \in \mathcal{P}|\text{Int}(\Theta_I(\theta(\mathbb{P}))) \ne \emptyset \}$. Similarly, a measure satisfies a uniform $\varepsilon$-LICQ condition (as in gafarov2024simple) if $\mathbb{P} \in \mathcal{P}^{LICQ; \varepsilon}$, where

align*[align* omitted — 211 chars of source]

where the set $\mathcal{V}(\mathbb{P}) \subseteq 2^{[q]}$ consists of sets of indices of binding inequalities that define vertices of the polytope $\Theta_I(\theta(\mathbb{P}))$.

propThe following hold: \begin{enumerate} • $\mathcal{P}^{Slater} \cup \mathcal{P}^{LICQ;0} \subset \mathcal{P} = \bigcup_{\delta > 0} \mathcal{P}^{\delta}$, where the inclusion is strict • $\mathcal{P}^{LICQ; \varepsilon} \subset \mathcal{P}^{\delta}$ for any $\delta \leq \varepsilon$, where the inclusion is strict \end{enumerate}
proofPart 1 follows by Proposition (ref), the definition of $\mathcal{P}^\delta$, and using part ii) in Appendix (ref). Part 2 follows by definition of $\mathcal{P}^{LICQ;\varepsilon}$.

Intuitively, the $\delta>0$ in Assumption U1 merely parametrizes the degree of irregularity that the researcher is willing to allow for the identified polytope over the considered set of measures. The resulting family of sets `covers' the unconstrained set of measures asymptotically as $\delta$ decreases to $0$. Uniform LICQ and SC, on the contrary, both restrict the set of measures. That is because measures like the one in Figure (ref) do not belong to either $\mathcal{P}^{LICQ;\varepsilon}$ for any $\varepsilon \geq 0$, or $\mathcal{P}^{SC}$.

figure[figure omitted — 1,409 chars of source]
excont[cont'd] Figure (ref) plots the range of $b-$values that correspond to the measures satisfying the $\delta-$condition for a given $\delta > 0$ in (ref). The $\delta-$condition cuts off an interval of $b$ along which the optimal vertex becomes `too sharp'. Recall that the case $b = 0$ leads to the failure of SC. However, it satisfies the $\delta-$condition for a relatively large $\delta$, because at the optimum $x^* = (-1, -1)'$ there is a set $J^* = \{1,3\}$ from Proposition (ref) such that the relevant matrix of binding constraints $M_{J^*}$ has the smallest singular value $\sigma_2(M_{J^*}) \approx 0.62 \gg 0$. \begin{figure}[h] \begin{tikzpicture} \draw[thick,->] (-3, 0) -- (3, 0) node[below right] {\( b \)}; \draw[thick] (0, -0.2) -- (0, 0.2); \node[above] at (0, 0.2) {\( 0 \)}; \node[above] at (-0.65, 0.4) {\( \mathcal{P}\setminus \mathcal{P}^\delta \)}; \fill[pattern= north east lines, pattern color = red, opacity=0.6] (-3, -0.2) -- (-1.3, -0.2) -- (-1.3, 0.2) -- (-3, 0.2) -- cycle; \fill[pattern = north east lines, pattern color = red, opacity=0.6] (0, -0.2) -- (3, -0.2) -- (3, 0.2) -- (0, 0.2) -- cycle; \node[above, red] at (-1.8, 0.5) {\( \mathcal{P}^{\delta} \)}; \node[above, red] at (1.8, 0.5) {\( \mathcal{P}^{\delta} \)}; \end{tikzpicture} \caption{The set of $b$ in problem (ref) satisfying a $\delta-$condition.} \end{figure}
exampleThere are LPs in which SC, LICQ and NFF all fail, but the $\delta-$condition is satisfied for a relatively large $\delta$. One example is the problem \begin{align*} \min -x_1 + x_2, \quad s.t. \quad x_2 \leq x_1, x_2 \geq x_1, x_1 \in [-1;1], \end{align*} where the $\delta-$condition is satisfied for $\delta > \sigma_2(M_{\{1,3\}}) \approx 0.62$. It is large relative to the values of $\delta$, for which the penalty vector suggested in Appendix (ref) is valid for small $n$. There are flat faces, as any pair with $x_2 = x_1$ and $x_1 \in [-1;1]$ is a solution, SC fails, as $\text{Int}(\Theta_I) = \emptyset$, and LICQ fails, as at an optimal $x_2 = x_1 = -1$ there are three binding constraints.

Uniform consistency of penalty function estimator

The penalty function estimator converges to the LP value $B(\mathbb{P})$ a.s. at rate $\sqrt{n}w_n^{-1}$, uniformly over the set of distributions that satisfy the $\delta-$condition for some $\delta >0$.

theorSuppose that i) $\hat{\theta}_n = \hat{\theta}_n(\mathbb{P})$ converges to $\theta_0(\mathbb{P})$ a.s. uniformly over $\mathcal{P}$ at rate $\overline{r}_n \uparrow \infty$, i.e. for all $r_n \uparrow \infty$ with $r_n = o(\overline{r}_n)$ and any $\varepsilon > 0$, \begin{align} \underset{n \to \infty}{\lim} \underset{\mathbb{P} \in \mathcal{P}}{\sup } \mathbb{P}[\underset{m \geq n}{\sup}r_m||\hat{\theta}_m - \theta(\mathbb{P})|| \geq \varepsilon] = 0, \end{align} and ii) $w_n(\mathbb{P}) = w_n \to \infty$ w.p.a.1 a.s. uniformly over $\mathcal{P}$, i.e. for any $M > 0$, \begin{align*} \lim_{n \to \infty} \inf_{\mathbb{P} \in \mathcal{P}} \mathbb{P}[\inf_{m \geq n} w_m > M ] = 1. \end{align*} Then, for all $r_n \uparrow \infty$ with $r_n = o(\frac{\overline{r}_n}{w_n})$ and any $\varepsilon > 0$, \begin{align*} \sup_{\delta > 0} \underset{n \to \infty}{\lim} \underset{\mathbb{P} \in \mathcal{P}^\delta}{\sup } \mathbb{P}[\underset{m \geq n}{\sup}r_m|\tilde{B}(\hat{\theta}_m;w_m) - B(\theta_0(\mathbb{P}))| \geq \varepsilon] = 0. \end{align*}

In applications, condition i) in Theorem (ref) can usually be established with rate $\overline{r}_n = \sqrt{n}$ by reference to the uniform LLN, provided $\theta_0$ is uniformly continuous in population moments.

exampleIn the AICM models (ref), $\theta_0$ is linear in moments of interactions of $Y(t)$ with treatment indicators and linear or hyperbolic in probabilities of $\{T = t, Z = z\}$ (see Section 3). Thus, condition i) in Theorem (ref) is established by, firstly, imposing integrability \begin{align*} \underset{C \to \infty}{\lim} \underset{\mathbb{P} \in \mathcal{P}}{\sup } \mathbb{E}_\mathbb{P}||\mathbb{Y}|| \mathds{1}\{||\mathbb{Y}|| > C\} = 0, \quad \mathbb{Y} \equiv (Y(t))_{t \in \mathcal{T}}, \end{align*} which holds whenever bounded outcomes are assumed. If also a full-support condition holds, \begin{align*} \inf_{\mathbb{P} \in \mathcal{P}, (t, z) \in \mathcal{T}\times\mathcal{Z}}\mathbb{P}[T = t, Z = z] > C, \end{align*} for some $C > 0$, then $\theta(\cdot)$ is uniformly continuous in population moments. Combining this with the LLN uniform in probability measure (see Proposition A.5.1 on p. 456 of van1996weak) yields condition (ii). If one additionally assumes that \begin{align*} \underset{C \to \infty}{\lim} \underset{\mathbb{P} \in \mathcal{P}}{\sup } \mathbb{E}_\mathbb{P}||\mathbb{Y}||^2 \mathds{1}\{||\mathbb{Y}|| > C\} = 0, \end{align*} the rate in (ref) is $r_n = \sqrt{n}$ (see Proposition A.5.2 on p. 457 of van1996weak).
remarkThe existence of a uniformly consistent estimator over $\mathcal{P}$ implies that $B(\cdot)$ is continuous over $\theta_0(\mathcal{P})$. However, as Example (ref) illustrates, it is not necessarilly continuous over the support of $\hat{\theta}_n$. It is straightforward to see that the case $b = 0$ in (ref) satisfies the $\delta$-condition for a relatively large $\delta$, but the plug-in estimator is still pointwise inconsistent at such measure.

Uniform consistency of the debiased estimator

This subsection provides a variation of the $\delta-$condition under which the debiased penalty estimator is shown to be uniformly consistent. Intuitively, the debiased estimator achieves uniform consistency whenever for all considered measures it is true that if the distance of $x \in \mathcal{X}$ from the polytope $\Theta_I(\mathbb{P})$ is large, then the constraint violation $\iota'(c(\mathbb{P})-M(\mathbb{P})x)^+$ is also large. To develop that notion formally, we introduce the following two terms.

definition[Face condition number] For a proper $k$-face\footnote{See Chapters 2 and 3 in grunbaum1967convex for the definition. We call a $k-$face proper if $0 \leq k<d$. A $0-$face is a vertex.} of a polytope $\Theta = \{x \in \mathbb{R}^d:\tilde{M}x \geq \tilde{c}\}$, given by $f$ and described by binding constraints $A \subseteq [q]$ with $|A| \geq d - k$ such that $\text{rk}(\tilde{M}_A) = d - k$, define the face condition number to be \begin{align*} \tilde{\kappa}(f) \equiv \underset{B \subseteq A: rk (\tilde{M}_B) = d-k}{\min} \sigma_{d - k}(\tilde{M}_B). \end{align*}
definition[Polytope condition number] For a polytope $\Theta$, define the polytope condition number as \begin{align*} \kappa(\Theta) = \underset{f - proper face of $\Theta$}{\min} \tilde{\kappa}(f). \end{align*}

The following proposition simplifies the interpretation of $\kappa(\Theta)$. It asserts that in the definition of $\kappa(\Theta)$ it suffices to only consider the vertices of $\Theta$,

propFor any polytope $\Theta$, $\kappa(\Theta)> 0$, and \begin{align*} \kappa(\Theta) = \underset{f - vertex of $\Theta$}{\min} \tilde{\kappa}(f). \end{align*}

We are now ready to define the polytope $\delta-$condition.

namedass[U2 (Polytope $\delta$-condition)] The class of measures $\overline{\mathcal{P}}$ satisfies the polytope $\delta-$condition for a given $\delta > 0$, if \begin{align*} \underset{\mathbb{P} \in \overline{\mathcal{P}}}{\inf} \kappa(\Theta_I(\mathbb{P})) \geq \delta. \end{align*}
figure[figure omitted — 1,991 chars of source]

The polytope $\delta-$condition lower bounds the smallest singular values of all full-rank matrices that can be constructed from the vertices of the polytope. Similarly to Assumption U1, it parametrizes the unconstrained set of probability measures, since at any fixed $\mathbb{P} \in \mathcal{P}$ we have $\kappa(\Theta_I(\mathbb{P})) > 0$ by definition. Proposition (ref) continues to hold for U2, if one substitutes $\mathcal{P}^\delta$ with $\mathcal{P}_p^\delta$ - the family of measures satisfying U2 for a given $\delta > 0$.

The following Lemma provides a minorization of the $L_1-$deviation from a polytope, and appears to be mathematically novel. Intuitively, $\kappa(\Theta)$ determines `sharpness' of the vertices of a polytope. The less sharp is a vertex, the greater the polytope constraints' violations need to be in terms of the distance from the polytope. Figure (ref) illustrates that point.

lemmaFor any non-empty and bounded polytope $\Theta = \{x \in \mathbb{R}^d|\tilde{M}x \geq \tilde{c}\}$: \begin{align*} \iota'(\tilde{c} - \tilde{M}x)^+ \geq d(x, \Theta)\kappa(\Theta) \end{align*}

By Theorem (ref), the biased penalty estimator is uniformly consistent at rate $\frac{\sqrt{n}}{w_n}$ under the $\delta$-condition. The following Theorem asserts that the debiased estimator converges at least at the same rate.

theorSuppose that the hypotheses i) and ii) of Theorem (ref) are satisfied. Then, for all $r_n \uparrow \infty$ with $r_n = o(\overline{r}_nw^{-1}_n)$ and any $\varepsilon > 0$, \begin{align*} \sup_{\delta > 0} \underset{n \to \infty}{\lim} \underset{\mathbb{P} \in \mathcal{P}_p^\delta}{\sup } \mathbb{P}[\underset{m \geq n}{\sup}r_m|\hat{B}(\hat{\theta}_m;w_m) - B(\theta(\mathbb{P}))| \geq \varepsilon] = 0. \end{align*} Moreover, for all $r_n \uparrow \infty$ with $r_n = o(\overline{r}_n)$ and any $\varepsilon > 0$, \begin{align*} \sup_{\delta > 0} \underset{n \to \infty}{\lim} \underset{\mathbb{P} \in \mathcal{P}_p^\delta}{\sup } \mathbb{P}[\underset{m \geq n}{\sup}r_m(B(\theta(\mathbb{P})) - \hat{B}(\hat{\theta}_m;w_m) )) \geq \varepsilon] = 0. \end{align*}
remarkTheorem (ref) establishes that under the polytope $\delta-$condition the debiased estimator converges at least at the rate $\sqrt{n}w^{-1}_n$ for some slowly growing $w_n$. Moreover, it converges from below at the rate $\sqrt{n}$ uniformly.
remarkIt is unclear if the $\sqrt{n}w_n^{-1}$ uniform rate is sharp. The simulation evidence in Appendix (ref) suggests that this may be the case, but only along the sequences of measures $\{\mathbb{P}_k\}_{k \in \mathbb{N}}\subseteq \mathcal{P}$, for which SC, LICQ and NFF all fail in the limit. Once such sequences are ruled out, the rate of $\sqrt{n}$ appears to be restored uniformly.

Monte Carlo

Consistency and inference simulations use $10^4$ and $10^3$ repetitions respectively.

Consistency

We first describe Example (ref) in more detail. Consider a setup with $d = 2$ variables and $q = 4$ constraints. Recall that $\theta = (p', \text{vec}(M)', c')'$. In this case, the parameters are defined as

align[align omitted — 316 chars of source]

The sample analogues $\hat{M}_n, \hat{c}_n$ of parameters in (ref) are obtained by substituting $b$ with its estimator $\hat{b}_n$. In our simulation study, that estimator is given by $\hat{b}_n = b + n^{-1}\sum^n_{i=1} U^b_i$, where $U^b_i \sim U[-1; 1]$ for $i \in [n]$ are i.i.d.

The true parameter $b$ is in one-to-one correspondence with the underlying measure, and $\theta$ is completely described by it: $\theta_0(\mathbb{P}) = \theta_0(b(\mathbb{P}))$. Consequently, it is sufficient to index the parameter $\theta$, the identified polytope $\Theta_I$ and the program's value $B$ by $b$ only.

Observe that the example (ref) is so engineered that $\Theta_I(\hat{b}_n)$ would never be empty or unbounded, ensuring the existence of the plug-in estimator $B(\hat{b}_n)$. A slight modification to that setup, that we consider in the next subsection, would lead the identified polytope to be empty with a potentially non-vanishing probability.

\paragraph{Measures} We consider two values of $b$: $b = 0$ and $b = -0.05$. At $b = 0$, SC fails, because $\Theta_I(0)$ has an empty interior. LICQ also fails at $b = 0$, as at the optimum $x^* = (-1,-1)'$, inequalities $1,2,4$ are binding. There are no flat faces at $b = 0$. At $b = -0.05$, SC, LICQ and NFF all hold. The values $b = 0$ and $b = -0.05$ result in the smallest singular values of $M_{J^*(b)} (b)$ matrices that correspond to the $75-$th and the $19-$th percentiles of the Tao-Vu distribution given in Theorem (ref) in the Online Appendix.

If $b < 0$, the norm of the `smallest' KKT vector $\lambda$ in the true LP corresponding to (ref) is proportional to $|b^{-1}|$. So, for a small negative $b$, the $\delta$-condition is only satisfied for a small value of $\delta > 0$. In this case the `optimal' $w$ is large. By that logic, the plug-in estimator, which in this case obtains as the limit $w \to \infty$, should perform well at such $b$, and potentially outperform the debiased penalty estimator. This is partly an artifact of the setup in (ref): additional noise in one of the slated lines' intercepts would render the problem infeasible with positive probability, which would worsen the performance of the plug-in estimator in finite samples (see Figure (ref)).

At $b \geq 0$, the $\delta$-condition is satisfied with a relatively large $\delta = \frac{\sqrt{5} - 1}{2}$, and so the `optimal' $w$ is relatively small. If $b = 0$, in 50% of the cases, $\hat{b}_n < 0$ is estimated. If, moreover, $w_n > \text{const}\times |\hat{b}_n^{-1}|$, the debiased penalty estimator would select the incorrect maximum of $0$ in such cases. A larger $w$ at $b = 0$ thus hampers the performance of the debiased penalty estimator.

\paragraph{Parameters} We set $w_n = \delta^{-1}_{0.2}(d) ||p||\frac{\ln \ln n}{\ln \ln 100}$, where $\delta_{\alpha}(d)$ is the $\alpha-$quantile of the $d-$dimensional Tao-Vu distribution given in Theorem (ref) in Appendix (ref). To ensure the same expansion rate for the set-expansion estimator, we set $\sqrt{\kappa_n} = \kappa_0 \times \ln \ln n$. There is no guidance as to the selection of $\kappa_0$. Our baseline is $\kappa_0 = 0.1$, and we explore other values in Figure (ref).

figure[figure omitted — 246 chars of source]

\paragraph{Discussion} Consider the right panel of Figure (ref), corresponding to $b=0$. The plug-in estimator is inconsistent, while the set-expansion estimator performs well, approaching $-1$ from below for larger $n$. The debiased penalty-function estimator is slightly upward-biased for smaller $n$, but yields the value of almost exactly $-1$ in larger samples. In contrast, the set-expansion estimator has a conservative rate, and remains slightly downward-biased even in larger samples.

The case of $b = -0.05$ is depicted in the left panel of Figure (ref). The plug-in estimator is consistent and appears to be the best estimator out of the three. The set-expansion estimator is severely downward-biased. This is because when the optimal vertex has a `sharp' angle, a small expansion of the inequalities' RHS may lead to a large shift of the vertex. To see that, consider shifting both inequalities outwards in Figure (ref) for a small and a large absolute value of $b$, when $b$ is negative. Once the expansion grows smaller, the set-expansion estimator slowly converges to the true value of $0$ from below. While selecting a smaller $\kappa_0$ parameter would improve the performance of the set-expansion estimator at $b = -0.05$, in the next part of our analysis we demonstrate that in this example $\kappa_0 = 0.1$ is close to being optimal in the uniform sense, because smaller $\kappa_0$ worsens the estimator's performance at $b = 0$. The debiased penalty estimator, in contrast, converges rather quickly. It is slightly conservative at smaller $n$, as it selects the incorrect vertex of $-1$ whenever $\hat{b}_n < 0$ and $|\hat{b}^{-1}_n| > \text{const} \times w_n$, i.e. when the penalty parameter is not large enough.

\paragraph{Robustness of the debiased penalty estimator} For $b < 0$, the debiased estimator may be expected to perform better at measures with a larger $|b|$. Figure (ref) illustrates that point:

figure[figure omitted — 224 chars of source]

Researchers applying any LP estimator should exercise caution when operating in smaller samples due to irregularity inherent in (ref). In example (ref), the debiased penalty estimator exhibits desirable behavior even along highly irregular measures for sample sizes of order $n = 5000$. Such sample sizes are not uncommon in partially identified settings. In our application, estimation is performed on $664633$ observations.

\paragraph{Alternative $\kappa_0$} We now investigate to which extent the behavior of the set-expansion estimator can be improved by selecting an alternative $\kappa_0$ parameter. Clearly, a larger $\kappa_n$ makes the set-expansion estimator more conservative. However, $\kappa_n$ cannot be selected as small as possible. If the expansion is too small, it may be insufficient to counteract the noise involved in estimating $\Theta_I$, leading to poor performance of the estimator at measures where SC fails to hold.

figure[figure omitted — 260 chars of source]

That tradeoff is very clear in Figure (ref). Compare the performance of the estimator with $\kappa_0 = 0.01$ to the baseline of $\kappa_0 = 0.1$. Decreasing $\kappa_0$ down to $0.01$ makes the estimator less conservative at $b = -0.05$, but results in an upward bias at $b = 0$. This occurs, because at $\kappa_0 = 0.01$ the resulting set-expansion sequence $\kappa_n$ is too small to counteract the estimation noise for the considered sample sizes. This logic is `monotone', meaning that selecting an even smaller $\kappa_0$ would worsen the performance at $b = 0$ further. Even at $\kappa_0 = 0.01$, the set-expansion estimator is quite conservative for the measure $b = -0.05$, and the parameter can clearly not be reduced any further without affecting the validity of the estimator at $b = 0$. It appears that the baseline choice of $\kappa_0 = 0.1$ is close to being optimal in our example. Therefore, the conservative behavior of the set-expansion estimator at $b=-0.05$ is not explained by a poor choice of the tuning parameter, but rather is a feature of the estimator itself. \paragraph{Noise in $\hat{c}_n$} We also report the result of consistency simulations for the DGP described in (ref). This DGP features noise on the right-hand-side of the inequalities that describe the polytope, which allows the estimated polytope to be empty.

figure[figure omitted — 350 chars of source]

In Figure (ref), the plug-in and the debiased penalty estimators perform equally well at $b = -0.05$. The rest of the conclusions are qualitatively unchanged.

Inference

In this section, we assess the performance of our inferential procedure. We compare it to the performance of two recently developed methods, chorussel2023 (CR) and gafarov2024simple (BG). We consider a slightly modified version of example (ref):

align[align omitted — 258 chars of source]

The sample analogues $\hat{M}_n, \hat{c}_n$ of parameters in (ref) are obtained by substituting $b, \zeta, \nu$ with their estimators:

align*[align* omitted — 165 chars of source]

where $U^{k}_i \sim U[-0.5;0.5]$, i.i.d. across $i \in [n]$ and $k \in \{b, \zeta, \nu\}$. We consider measures, for which the true values of $\zeta = \nu = 0$, whereas we still vary $b$ as in the example before. As before, we mainly consider $b = -0.05$ and $b = 0$.

Note that in (ref) the plug-in estimator of the polytope $\hat{\Theta}_I$ may be empty for some realizations of $\{U^k_i\}_{i,k}$. That is the case, for example, if $\hat{b}_n = \hat{\zeta}_n = 0$ and $\hat{\nu}_n > 0$.

\paragraph{Other methods} In case SC fails, the BG estimator is not applicable. It fails to exist in around 25% of cases at $b = 0$. The CR estimator in the main text of the paper also relies on $B(\hat{\theta}_n)$ and is not applicable if SC fails. The procedure described on p.47 of the Supplementary Appendix in chorussel2023 would not be practically applicable without the results contained in the present paper. It can be shown that the CR augmented procedure combines a random, non-vanishing set expansion and objective function perturbation with the penalty function approach\footnote{Because the penalty function estimator can be rewritten in the form of an equivalent problem w.p.a.1.}. The authors, however, treat the analogue of the penalty parameter as `some [fixed] large value'. They proceed to argue that the inequalities, which establish the size of their inference procedure, hold for any value of $w$, which appears to suggest that its selection is unimportant. There is no practical guidance on $w$ selection, and CR do not implement the augmented estimator in their simulations. In this paper, we studied penalized estimation in great detail, and our results suggest that the appropriate choice of $w$ is critical. Both our previous findings and simulations in Figure (ref) demonstrate that for different values of $w$ the performance of the CR procedure ranges from highly conservative to invalid. Selecting `a large value' of $w$ does not yield a valid procedure in finite samples. Our simulations also suggest that CR augmented approach can perform relatively well if combined with our results on the rate and level of $w_n$. Unlike our approach, however, it remains asymptotically conservative due to the use of non-vanishing random expansions.

\paragraph{Implementation details} As mentioned in Section 2.2, one can obtain an estimator $\hat{\Sigma}_n$ for the asymptotic covariance matrix of $\hat{\theta}_n$ via resampling. One then plugs it into the expression for $\sigma(\cdot)$ to obtain the required s.e. An even simpler approach is to obtain an estimator for $\sigma(\hat{A},\hat{x},\hat{v},\Sigma)$ directly by bootstrapping the quantities estimated on the second fold, while keeping the first fold quantities fixed. We have verified that the performance of the two approaches is similar to using the closed-form estimator for $\Sigma$, namely $\hat{\Sigma}_n = G \hat{\Omega}_n G'$, where $G \equiv \frac{\partial \theta}{\partial (b ~ \zeta ~ \nu)'}$ and $\hat{\Omega}_n$ is the sample covariance matrix of $(b_i ~ \zeta_i ~ \nu_i)'$. We employ the latter estimator in our simulations. When implementing the procedure in chorussel2023 (CR), we use uniform noise with the support size of $0.001$, as recommended in the paper. Note that we refer to their parameter $M$ as $w$. The estimator in gafarov2024simple (BG) was implemented using the code kindly provided by the author.

figure[figure omitted — 857 chars of source]

\paragraph{Discussion} Results of our simulations are given in Figure (ref). Overall, it appears that our estimator has the correct nominal level even at smaller sample sizes, whereas the CR with penalty $w_n$ overcovers asymptotically. BG estimator achieves nominal coverage asymptotically, although it over-covers in smaller samples and may yield a very conservative left confidence bound, as is evident from the right panel. It fails at $b = 0$, because the SC fails\footnote{We have still run the corresponding simulations, but are not displaying them. BG fails to exist in around 25% of cases, while the remaining simulations result in highly conservative bounds with incorrect coverage}. For different fixed values of $w$ the Cho-Russell procedure's performance may range from highly conservative to invalid. We illustrate this by adding lines with $w = 2$ and $w = 1000$ to the figures corresponding to $b = -0.05$ and $b = 0$, respectively.

Special case of LP bounds

Let $Y \in \mathbb{R}$ denote the outcome of interest\footnote{Univariate case is considered for simplicity of exposition, but the extension to multivariate outcomes is immediate.}, $T \in \mathbb{R}$ stand for the treatment, and $Z \in \mathbb{R}^{d_Z}$ be the candidate instrument. Here treatment is any variable which effect on $Y$ we attempt to infer, whereas the term instrument refers to an auxiliary variable that allows us to partially identify the causal effect of interest. Our approach nests the case when $Z$ is the usual IV. The reader seeking an economic intuition may refer to Section (ref), where $Y$ is the log-wage, $T$ is the level of education, and $Z$ is a proxy for the level of ability.

Denote the supports $Y,T,Z$ as $\mathcal{Y}$, $\mathcal{T}$ and $\mathcal{Z}$ respectively. Throughout this section we consider the case of continuous outcome and discrete treatment and instrument, namely $\mathcal{Y}$ is a (possibly infinite) interval\footnote{We focus on the uncountable case for concreteness. Inclusion (ref) in Theorem (ref) holds generally, while sharpness also holds for the case of arbitrary $\mathcal{Y}$ if there are no almost sure inequalities.}, while $N_T \equiv |\mathcal{T}| < \infty$ and $N_Z \equiv |\mathcal{Z}| < \infty$. In non-parametric bounds literature it is rather conventional to employ a discrete instrument at the estimation stage (see MP2009). While our main identification result may be extended to continuous $Z$, estimation in such settings is left for future research.

Our setup accommodates missing observations of the dependent variable. Namely, we split the set of treatments into two disjoint subsets $\mathcal{T} = \mathcal{O} \sqcup \mathcal{U}$. Whenever $T \in \mathcal{O}$, the researcher observes $Y, T, Z$, whereas if $T \in \mathcal{U}$, only the covariates $T, Z$ are observed. For example, in blundell2007changes the wage is observed only if an individual is employed. Corresponding to the legs of the treatment are the potential outcomes $Y(t), ~ t \in \mathcal{T}$:

align*[align* omitted — 125 chars of source]

Continuing the wages and education example, the value of $Y(t)$ for a fixed $t\in \mathcal{T}$ may then correspond to the potential wage that an individual with the associated random characteristics would get, had she obtained education $t$.

Let us collect the potential outcomes in the vector $\mathbb{Y} \equiv (Y(t))_{t \in \mathcal{T}} \in \mathbb{R}^{N_T}$. Variables $(\mathbb{Y}, T, Z, Y)$ are jointly defined on the true probability space $(\mathbb{P}, \Omega, \mathcal{S})$ and we let $\mathcal{P}$ denote the considered collection of probability measures on $(\Omega, \mathcal{S})$, such that $\mathbb{P} \in \mathcal{P}$. We impose the following conditions on the set of considered measures throughout this section:

namedass[I0 (Conditions on $\mathcal{P}$)] $\mathcal{P}$ is such that $P \in \mathcal{P}$ if: i) $P$ generates $F_{T,Z}(\cdot)$ and $\{F_{Y|T = t, Z}(\cdot)\}_{t \in \mathcal{O}}$; ii) $P[T = d, Z = z] > 0$ $\forall d, z \in \mathcal{T} \times \mathcal{Z}$ and iii) $|\mathbb{E}_P[Y(t)|T = d, Z = z]| < \infty$ for all $z \in \mathcal{Z}$ and $t, d \in \mathcal{T}$

Part i) of the Assumption I0 formalizes the assumed identification pattern. It says that the joint distribution of $T, Z$ is always identified and the researcher also observes the joint distribution of $Y, T, Z$ whenever $T \in \mathcal{O}$. Parts ii) and iii) of I0 ensure that all conditional expectations and probabilities are well-defined and finite\footnote{Similar identification results can still be obtained if one relaxes the full-support condition for some known pairs from $\mathcal{Z} \times \mathcal{T}$. Note that it can also be verified in the data.}.

remarkUnder no missing data, i.e. $\mathcal{T} = \mathcal{O}$, condition i) is equivalent to $P$ generating the identified joint distribution $F_{Y, T, Z}(\cdot)$.

We define the vector $m$ collecting all elementary conditional moments as

align*[align* omitted — 136 chars of source]

and suppose that the researcher is interested in the target parameter $\beta^*$, given by

align[align omitted — 69 chars of source]

where $\mu^{*}: \mathcal{P} \to \mathbb{R}^{N^2_T N_Z}$ is identified\footnote{By this we mean that $\mu^*(P) = \mu^*(\mathbb{P})$ for all $P \in \mathcal{P}$.} and chosen by the researcher. It parametrizes the choice of the outcome of interest, as the following remark clarifies.

remarkThe form (ref) nests i) $\mathbb{E}[Y(t)]$, ii) $ATE_{td} = \mathbb{E}[Y(t) - Y(d)]$ and iii) $CATE_{td, A, B} = \mathbb{E}[Y(t) - Y(d)|T \in A, Z \in B]$.

Affine inequalities over conditional moments

We now introduce the general class of identifying conditions described by affine inequalities over conditional moments, potentially augmented with affine a.s. restrictions. These restrict the set of admissible measures to $\mathcal{P}^*$:

align[align omitted — 171 chars of source]

where $b^*: \mathcal{P} \to \mathbb{R}^R$, $M^*: \mathcal{P} \to \mathbb{R}^{R \times N^2_T N_Z}$, $\tilde{b}: \mathcal{P} \to \mathbb{R}^{\tilde{R}}$ and $\tilde{M}: \mathcal{P} \to \mathbb{R}^{\tilde{R} \times N_T}$ are identified parameters, chosen by the researcher. These parametrize the choice of $R \in \mathbb{N}$ identifying inequalities on conditional moments of potential outcomes as well as $\tilde{R} \in \mathbb{N}$ almost sure inequalities on the potential outcomes. In general, $\psi(P) \equiv (\mu^*(P),b^*(P), M^*(P), \tilde{b}(P), \tilde{M}(P))$ and $m(P)$ are functionals of $P$. We omit this dependence whenever it does not cause confusion.

The family of models that can be written in the form (ref) is very rich, as illustrated by the following examples.

example[MIV] $Z \in \mathbb{R}$ is a monotone IV MP2000 if, for each $t \in \mathcal{T}$ and $z, z' \in \mathcal{Z}$, if $z' \geq z$, then $\mathbb{E}[Y(t)|Z = z'] \geq \mathbb{E}[Y(t)|Z = z]$. MIV is nested in AICM for an appropriate choice of matrix $M^* = M_{MIV}$ and $b^* = 0$, and no a.s. restrictions, $\tilde{M} = 0, ~\tilde{b} = 0$. Monotone treatment selection assumption in MP2000 obtains under $Z = T$.
example[IV] $Z \in \mathbb{R}$ is mean-independent of potential outcomes if, for each $t \in \mathcal{T}$ and $z, z' \in \mathcal{Z}$, $ \mathbb{E}[Y(t)|Z = z'] = \mathbb{E}[Y(t)|Z = z]$. This assumption is nested in AICM for $M^* = M_{IV} \equiv (M'_{MIV} \quad - M'_{MIV})'~$, $b^* = 0$, and no a.s. restrictions.
example[MTR] Monotone treatment response assumption MP2000 imposes that, for each $t, t' \in \mathcal{T}$, if $t' > t$, then $Y(t') \geq Y(t)$ a.s. It is nested in AICM for an appropriate choice of matrix $\tilde{M} = \tilde{M}_{MTR}$ with $\tilde{b} = 0$, and no inequalities over conditional moments, $M^* = 0, ~ b^* = 0$.
example[Roy model] laffers2019 imposes that for each $t, z \in \mathcal{T} \times \mathcal{Z}$, the individual's choice is, on average, optimal: $\mathbb{E}[Y(t)|T = t, Z = z] =\max_{d\in \mathcal{T}} ~ \mathbb{E}[Y(d)|T = t, Z = z]$. It is nested for an appropriate choice of matrix $M^{*} = M_{ROY}$ and $b^* = 0$, and no a.s. restictions.
example[Missing data] blundell2007changes derives bounds on $F(w|x)$ - the cdf of wages evaluated at some $w \in \mathbb{R}$, conditional on $X = x$. The wage is observed if the individual is employed, $E = 1$, and unobserved otherwise ($E = 0$). Introduce $\mathcal{O} = \{1\}$ and $ \mathcal{U} = \{0\}$. Let $Y(t) \equiv \mathds{1}\{W \leq w\}$, so that $\mathbb{E}[Y(t)|X = x] = F(w|x)$. Our approach allows to accommodate all identifying conditions in the original paper by appropriately choosing $M^*, b^*$ and $\tilde{M}, \tilde{b}$.
remarkCombinations of assumptions are obtained by stacking the respective matrices, as in Example (ref). Sensitivity analysis can be performed via relaxations $b_{\ell}^* = b^* + \ell$, or $\tilde{b}_\ell = \tilde{b} + \ell$ for some $\ell \geq 0$. For example, given some $\ell = \{\ell(t,z,z')\}_{t,z, z'} \geq 0$, inequalities $\mathbb{E}[Y(t)|Z = z'] - \mathbb{E}[Y(t)|Z = z] \geq -\ell(t,z,z')$ for $z, z' \in \mathcal{Z}$ with $z' > z$ and $t \in \mathcal{T}$ yield a relaxation of MIV. In de2017effect the shape of observed moments may suggest a failure of monotonicity near the boundaries of $\text{Supp}(Z)$. Selecting positive $\ell(z, z')$ for values of $z, z'$ close to the boundaries could constitute a meaningful robustness check.

Linear programming bounds

We now provide the general identification result for the models described by $\mathcal{P}^*$. Let us construct $x$ that collects unobserved pointwise-conditional moments\footnote{Formally, $x$ is a partially identified functional $x = x(P)$, whereas $\overline{x} = \overline{x}(P)$ is point-identified.} and the vector $\overline{x}$ of those pointwise-conditional moments that are identified:

align*[align* omitted — 250 chars of source]

For known selector matrices $P_m, \overline{P}_m$, one can then decompose $m$ as

align*[align* omitted — 130 chars of source]

It is also straightforward to observe that $\tilde{M} \mathbb{Y} + \tilde{b} \geq 0$ a.s. implies

align[align omitted — 108 chars of source]

Before we state the main identification result, let us construct the matrix $M^{**}$ and the vector $b^{**}$ that combine the conditional restrictions with the implications of the almost sure restrictions in (ref), as

align[align omitted — 272 chars of source]
theorSuppose Assumption I0 holds. For any $\psi$, the sharp identified set $\mathcal{B}^*$ for $\beta^*$ satisfies \begin{align} \mathcal{B}^* \equiv (\mu^{*\prime} m)(\mathcal{P}^*) \subseteq \{\beta \in \mathbb{R}|\underset{x: Mx \geq c}{\inf} p'x \leq \beta - \overline{p}'\overline{x} \leq \underset{x: Mx \geq c}{\sup} p'x\}, \end{align} where \begin{align*} \overline{p} \equiv \overline{P}'_m\mu^*, \quad p \equiv P'_m\mu^*, \quad M \equiv M^{**}P_m, \quad c \equiv - b^{**} - M^{**} \overline{P}_m \overline{x}. \end{align*} Reverse inclusion in (ref) holds if $\tilde{M} = 0, \tilde{b}= 0$, or if one of the following is true: \begin{enumerate} • MTR holds, i.e. $Y(t_1) \geq Y(t_0)$ a.s. $\forall t_1, t_0 \in \mathcal{T}$ s.t. $t_1 > t_0$: \begin{align*} \tilde{M} = \tilde{M}_{MTR}, \quad \tilde{b} = 0_{N_{T} - 1} \end{align*} • Outcomes are bounded, $Y(t) \in [K_0; K_1]$, $\forall t \in \mathcal{T}$ a.s. for known $K_1 > K_0$: \begin{align} \tilde{M} = \tilde{M}_b \equiv \begin{pmatrix} I_{N_T}\\ -I_{N_T} \end{pmatrix}, \quad \tilde{b} = \tilde{b}_b \equiv \begin{pmatrix} -K_0 \cdot \iota_{N_T}\\ K_1 \cdot \iota_{N_T} \end{pmatrix}. \end{align} • MTR holds, outcomes are bounded and $(\Omega, \mathcal{S})$ can support a $U[0;1]$ r.v.: \begin{align} \tilde{M} = \begin{pmatrix} \tilde{M}_{MTR}\\ \tilde{M}_{b} \end{pmatrix}, \quad \tilde{b} = \begin{pmatrix} \tilde{b}_{MTR}\\ 0_{N_{T} - 1} \end{pmatrix}. \end{align} \end{enumerate}

The matrix $\tilde{M}_{MTR}$ is defined in Appendix (ref).

Theorem (ref) postulates that bounds on the target parameter $\beta^*$ under $\mathcal{P}^*$ can be obtained by solving two linear programs. The LP bounds are sharp if there are no a.s. inequalities in the model, or if the a.s. inequalities parametrize three special cases that are typically used in the literature. Otherwise, the LP bounds are valid, but not necessarily sharp, as we show in Appendix (ref). This is because, in general, the entire distribution of $\mathbb{Y}, T, Z$ is relevant for $\beta^*$ under a.s. restrictions. The naive approach of searching over such joint distributions, however, would involve infinite-dimensional optimization, because $|\mathcal{Y}| = \infty$.

remarkWe are not aware of affine a.s. restrictions $\tilde{M}, \tilde{b}$ used in applied work that are not special cases 1-3 in Theorem (ref). The special cases 1-3 appear in blundell2007changes, kreider2012identifying, pepper, siddique.
remarkThe empirical literature has extensively relied on the MIV + MTR + MTS combination of MP2000 assumptions, as it yields the tightest bounds out of all classical conditions. In the absence of a theoretical justification, this has led to errors laffers. Theorem (ref) provides the first available sharp bounds under this combination when $Y$ is continuous.
remarkIn any of the sharp cases in Theorem (ref), the polytope $\Theta_I \equiv \{x \in \mathbb{R}^q: Mx \geq c\}$ gives the sharp identified set for the vector of unobserved pointwise-conditional moments $x$ under the corresponding AICM model described by $M, c$.

Conditional monotonicity assumptions

A particular family of identifying conditions that can be written in the form (ref) is the conditional monotonicity class of assumptions. These impose that potential outcomes are mean-monotone in the instrument even within some treatment subgroups. While more restrictive than the conventional MIV, conditionally monotone instrumental variables (cMIV) allow to sharpen the bounds on the outcomes of interest. Throughout this section, we assume that outcomes are bounded $Y(t) \in [K_0; K_1]~ a.s.$ for known $K_0, K_1 \in \mathbb{R}$, $K_0 < K_1$. We also suppose that there are no missing data\footnote{Although it is hopefully clear from our general approach how cMIV conditions extend to the missing data case.}, i.e. $\mathcal{T} = \mathcal{O}$.

We argue that cMIV assumptions are reasonable in classical applications, discuss the difference between MIV and cMIV and develop a formal testing strategy for a particular version of cMIV. This testing procedure relies on the observed outcomes' monotonicity, which has been typically used in applied work to justify applying MIV. Our results imply that if such monotonicity is observed and the researcher is comfortable with MIV, the cMIV assumption is inexpensive, and can be applied to sharpen the bounds on the outcomes of interest. In some applications, as is the case in Section 5, cMIV yields informative bounds even when the classical conditions fail to do so.

While we only discuss three variations of cMIV, the class of such assumptions is potentially richer\footnote{One can consider the class of conditional restrictions $\mathbb{E}[Y(t)|T \in A, Z = z'] \geq \mathbb{E}[Y(t)|T \in A, Z = z], ~ ~ \forall A \in \mathcal{F}_t$ for all $t \in T$ where subcollections $\mathcal{F}_t \subseteq \mathcal{T}$ are chosen by the researcher. The triplet $M, p,c$ from Theorem (ref) under any such restrictions can be constructed by following the procedure in Appendix (ref).}, and Theorem (ref) applies in any such framework.

namedass[cMIV-s] Suppose that for any $t \in \mathcal{T}$, $A \subseteq \mathcal{T}: A \ne \{t\}$ and $z, z' \in \mathcal{Z}$ s.t. $z' > z$ we have: \begin{align*} \mathbb{E}[Y(t)|T \in A, Z = z'] \geq \mathbb{E}[Y(t)|T \in A, Z = z], \end{align*} i.e. the potential outcomes are, on average, non-decreasing in $Z$ for any treatment subgroup.

The strong conditional monotonicity assumption possesses the greatest identifying power across all considered cMIV conditions. To see that cMIV-s implies MIV, set $A = \mathcal{T}$ in the definition above.

namedass[cMIV-w] Suppose MIV holds and for any $t \in \mathcal{T}$ and $z, z' \in \mathcal{Z}$ s.t. $z' > z$ we have: \begin{align*} \mathbb{E}[Y(t)|T \ne t, Z = z'] \geq \mathbb{E}[Y(t)|T \ne t, Z = z], \end{align*} i.e. the potential outcomes are, on average, non-decreasing in $Z$ for the non-treated and the whole population.

The weak conditional monotonicity assumption allows for closed-form expressions for sharp bounds that are easy to compute and perform inference on.

namedass[cMIV-p] Suppose MIV holds and for any $t \in \mathcal{T}, d \in \mathcal{T}\setminus\{t\}$ and $z, z' \in \mathcal{Z}$ s.t. $z' > z$ we have: \begin{align*} \mathbb{E}[Y(t)|T = d, Z = z'] \geq \mathbb{E}[Y(t)|T = d, Z = z], \end{align*} i.e. the potential outcomes are, on average, non-decreasing in $Z$ conditional on any counterfactual level of treatment.

The pointwise conditional monotonicity assumption is directly testable under a mild homogeneity condition, see Section (ref).

Conditional monotonicity restrictions differ in the collection of treatment subsets over which monotonicity in the instrument is assumed. The strong conditionally monotone instruments are such that, among individuals from any given counterfactual treatment subgroup, higher values of $Z$ are, on average, associated with higher potential outcomes. The weak conditional monotonicity restriction only imposes the same mean-monotonicity on the whole population and on the untreated, whereas the pointwise form assumes it over the entire population as well as conditional on each counterfactual level of treatment.

remarkAll cMIV assumptions imply MIV. Moreover, cMIV-w, cMIV-p are implied by cMIV-s. If treatment is binary, cMIV-s, cMIV-w and cMIV-p are equivalent.

While it is possible for the general apporach of form (ref), cMIV conditions avoid assuming monotonicity over the observed treatment subset $\{T = t\}$. This is because such monotonicity is identified. If it holds, it should not add any identifying power to our conditions in theory. However, large violations of the observed outcomes' monotonicity will lead the test developed in Section (ref) to reject cMIV-p and cMIV-s.

The following observation motivates the use of cMIV assumptions.

propMP2000 MIV bounds are not sharp under either cMIV-s, cMIV-w or cMIV-p.
proofConsider a binary treatment $T$, three levels of the instrument $Z \in \{z_0, z_1, z_2\}$ with $z_0 < z_1 < z_2$ and $-K_0 = K_1 = 1$. Suppose for a fixed $t \in \{0,1\}$, we have $\mathbb{E}[Y(t)|T = t, Z = z_i] = 0$, with $P[T \ne t| Z = z_0] = 0.125$, $P[T \ne t| Z = z_1] = 0.5$, $P[T \ne t| Z = z_2] = 0.25$. The no-assumptions lower bounds on $\mathbb{E}[Y(t)|Z = z_i]$ are $(-0.125, ~ -0.5, ~ -0.25)$. MIV `irons' the no-assumptions bounds to $(-0.125, ~ -0.125, ~ -0.125)$, which also implies the lower bounds on $\mathbb{E}[Y(t) | T \ne t, Z = z_i]$: $(-1, ~ -0.25, ~ -0.5 )$. Under cMIV, one can further `iron' these to improve the lower bound for $z_2$ up to $-0.25$, so that the lower bound on $\mathbb{E}[Y(t)|Z = z_2]$ becomes $-1/16 > -1/8$.

Sharp bounds for all versions for cMIV follow from Theorem (ref). We also show that under cMIV-w the bounds can be characterized explicitly, which is especially convenient if the treatment is binary, so that all cMIV assumptions coincide (see Appendix (ref)). For didactic purposes, we also provide detailed guidance on how to construct the triplet $M, c, p$ from Theorem (ref) under cMIV-s, cMIV-p and MIV in Appendix (ref).

Discussion of cMIV

This section illustrates the difference between MIV and cMIV by considering two parametric examples with classical applications.

Education selection

Consider the following empirical setup. Suppose $T$ is an indicator of whether or not an individual has a university degree, $Y(t)$ are potential log wages and $Z$ is an observed indicator of ability.

MIV assumption on $Z$ implies that more able individuals can do better both with and without a college degree on average: $\mathbb{E}[Y(t)|Z = z]$ - monotone in $z$. cMIV additionally imposes that: i) among those who have a college degree, a smarter individual could have done relatively better on average than their counterpart if both did not have it: $\mathbb{E}[Y(0)|Z = z, T = 1]$ - monotone in $z$; and ii) among those who do not have a college degree, a smarter individual could have done relatively better on average than their counterpart if both had it: $\mathbb{E}[Y(1)|Z = z, T = 0]$ - monotone in $z$.

We now consider a parametric example. Suppose that $\eta$ measures how diligent one is from birth and is ex-ante mean-independent of $Z$. While $Z$ is observed by both the employers and the econometrician (e.g. an IQ score), the employer additionally observes the employee effort level $\eta + \varepsilon$ with $\varepsilon \rotatebox[origin=c]{90}{$\models$} (Z, T, \eta)$. Suppose $Var(Z) = Var(\eta) = 1$ and $\mathbb{E}[Z]=\mathbb{E}[\eta] = \mathbb{E}[\varepsilon] = 0$. Suppose that, on average, employees choose $T$ to maximize their expected earnings. This motivates a stylized Roy selection model, with

align*[align* omitted — 159 chars of source]

where $\nu \rotatebox[origin=c]{90}{$\models$} (Z, \eta, \varepsilon)$ is remaining heterogeneity, and $\varepsilon(t) \equiv \beta_2(t) \varepsilon$. MIV demands that

align*[align* omitted — 50 chars of source]

MIV postulates that the direct effect of ability on potential earnings is positive. It seems reasonable to suppose that $\beta_i(t) \geq 0, ~ i = 1,2, ~ t = 0, 1$, i.e. both effort and ability increase potential wages. Letting $\delta_z \equiv \beta_1(1) - \beta_1(0)$ and $\delta_\eta \equiv \beta_2(1) - \beta_2(0)$ denote the differentials in the effects of ability and effort respectively, the additional requirement of cMIV is that:

align[align omitted — 435 chars of source]

where $\tilde{\nu} \equiv \beta_0(1) - \beta_0(0) + \nu$.

Notice that if $\delta_z$ and $\delta_\eta$ are of different signs, for example because the jobs that one may apply for with a college degree are more ability-intensive ($\delta_z > 0$), whereas those which are available otherwise are more skill-intensive ($\delta_\eta < 0$), the additional conditional monotonicity requirements (ref)-(ref) are less strict than MIV. This is because, conditional on both having a degree and not having it, ability and effort are positively associated.

Intuitively, among those who do not have a degree ($T=0$), people of higher ability must have had stronger incentives to forgo college. This should have been because a higher level of diligence gives them a comparative advantage in effort-intensive jobs. Among those with a degree, higher ability implies a comparative advantage in ability-intensive occupations, which explains their willingness to select into this option ($T=1$). It does not, therefore, signal as low an effort level as it would for a less capable individual.

Now consider the same setup with $\delta_z = \delta_\eta > 0$ and $\tilde{\nu} = 0$,

align*[align* omitted — 52 chars of source]

This selection mechanism can be explained by the fact that to get a degree one needs to be either hard-working or of high ability. The requirement of MIV is unchanged, and cMIV necessitates that:

align[align omitted — 201 chars of source]

In this case, conditional on each level of education, effort level $\eta$ and ability $Z$ are negatively associated, so the conditional selection terms in (ref)-(ref) make cMIV a stricter assumption than MIV. Intuitively, a more able individual with a college degree did not need to work as hard to get it as her counterpart with a lower ability. Similarly, if an individual is capable, but does not have a degree, she has to be of lower effort as otherwise she would have selected into education.

Even if MIV holds, cMIV can thus fail if employer prefers effort over ability to the extent that the conditional negative association between the two outweighs the direct impact of ability on wages as well as any ex-ante positive correlation between the employer-observed signal of diligence and the ability.

An examination of equations (ref) and (ref) suggests that cMIV is more likely to hold whenever $\delta_z$ is small relative to $\delta_\eta$, while $\beta_1(\cdot)$ is large relative to $\beta_2(\cdot)$. This means that $Z$ should be relatively weak in the parlance of the classical IV models, and strongly monotone.

Overall, it seems reasonable to use a proxy for the level of ability as a conditionally monotone instrument in the estimation of returns to schooling. One would be inclined to think that while $Z$ does enter selection, it affects the potential outcomes directly and strongly enough, so that there are no subgroups by schooling for which a higher value of ability would correspond to lower potential wages on average.

Simultaneous equations

As some aspects of mathematical intuition may be muted in discrete models, we also consider a simple continuous setup to confirm the insights derived from the previous analysis. For illustrative purposes, we drop the boundedness and discreteness assumptions and consider the supply and demand simultaneous equations,

align*[align* omitted — 117 chars of source]

The observed log-price $P$ clears the market in expectation,

align[align omitted — 115 chars of source]

where $\eta, Z$ are continuous unobserved and observed random variables respectively, with $\mathbb{E}[\eta|Z] = 0$ a.s.\footnote{Mean independence is not restrictive, as it can always be enforced by redefining the d.g.p. in an observationally equivalent way.} and $\mathbb{E}[\varepsilon^k] = 0$ with $\varepsilon^k \rotatebox[origin=c]{90}{$\models$} (\eta, Z, \varepsilon^{-k})$ for $k \in \{s,d\}$. Further assume that all functions of $p$ are continuous.

Potential price $p$ indexes the potential outcomes, giving rise to the demand and supply schedules. Suppose we aim to identify the elasticity of supply, $\mathbb{E}[(q^s(p_1) - q^s(p_0))/(p_1 - p_0)]$ for some $p_1 > p_0$, and $Z$ is a monotone instrument for $q^s(p)$, while $P$ can be interpreted as treatment. $\eta$ is unobserved heterogeneity and $\varepsilon^k$ are random violations from the market clearing condition or measurement errors independent of the rest of the model. For an individual realization of market clearing an econometrician observes $\{P, \{q^k(P)\}_k, Z\}$, but does not observe the schedules at other prices $\{q^k(p)\}_k \text{ for } p \ne P$, nor disturbances $\{\eta, \{\varepsilon^k\}_k\}$.

Define $\delta_z(p) \equiv \beta^s(p) - \beta^d(p)$ and similarly for $\eta$, with $\delta_p(p) \equiv \alpha^s(p) - \alpha^d(p)$. As stated, the model is potentially incomplete or incoherent, as for a given vector $(Z, \eta)$ equation (ref) may have multiple or no solutions. To avoid that, so long as that the support of $Z, \eta, \varepsilon^k$ is full, it is necessary that $\delta_z(p)$, $\delta_\eta(p)$ be constant. We shall assume that for simplicity. Provided that $\delta_p(p)$, which determines the excess supply at fixed $(Z, \eta)$, is strictly increasing and has full image, the model is complete and coherent, and

align[align omitted — 82 chars of source]

Equation (ref) introduces a deterministic linear relationship between $Z$ and $\eta$ conditional on each given value of $P$. As we saw in the previous example, this constitutes the worst-case scenario for cMIV, if $\delta_z$ and $\delta_\eta$ have the same sign. A noisier selection mechanism would relax the conditional link between $Z$ and $\eta$, and would thus weaken the conditional selection channel.

Note that the reduced-form error is $u \equiv \gamma^s(P) \eta + \kappa^s(P) \varepsilon^s$ and there is a simultaneity bias, as

align*[align* omitted — 170 chars of source]

In this setup, MIV requires that

align*[align* omitted — 80 chars of source]

Whereas cMIV-p additionally imposes that

align[align omitted — 192 chars of source]

Suppose that $\delta_z, \delta_\eta \ne 0$ to rule out uninteresting cases. (ref) then implies

align[align omitted — 116 chars of source]

For concreteness, consider two positive supply shocks, i.e. $\beta^s(p), \gamma^s(p) > 0$. Equation (ref) means that either $\eta$ and $Z$ affect the reduced-form equilibrium price in different directions (recall the comparative advantage example), or the effect of $Z$ on the equilibrium price relative to its effect on the supply schedule is smaller than that of $\eta$,

align[align omitted — 222 chars of source]

Under $\text{sgn}(\delta_\eta) = \text{sgn}(\delta_z)$, equation (ref) once again requires that $Z$ be strongly monotone and relatively weak for cMIV-p to hold. The logic we described may help the researcher navigate the potential economic forces in a given application to decide whether cMIV-p is a suitable assumption.

For example, consider estimating the supply elasticity in the market for plane tickets in the early days of Covid-19 pandemic. Suppose $Z$ is an inverse Covid-stringency index for the economy, while $\eta$ may be interpreted as residual cost shocks, defined to be mean-independent of $Z$. It is likely that $\delta_\eta \approx \gamma^s$, i.e. residual cost shocks affect mainly the supply in that sector, and not the demand. It is also likely that either supply is less responsive to $Z$ than demand (so that cMIV is implied by MIV), or the effects are of the same order of magnitude. $Z$ is therefore likely to be a conditionally monotone instrument.

Testing cMIV

One could argue against cMIV conditions whenever $\mathbb{E}[Y(t)|T = t, Z= z]$ fail to be monotone in the data. In general, the power and size of that test are unclear. There is, however, a special case when cMIV can be tested directly, provided that the researcher is willing to assume MIV. In some applications one may conjecture that the potential outcomes' functions $Y(t)$, either in the reduced or in the structural form, are such that the relative effects of $Z$ and the unobserved variable(s) $\eta$, potentially correlated with $Z$, are unchanged across outcome indices $t$.

Researchers often impose even stricter versions of this homogeneity assumption. For example, MP2009 discuss MIV identification under HLR condition: $Y(t) = \beta t + \eta$. Conditions in Proposition (ref) relax HLR to an arbitrary shape of response of a potential outcome to treatment and allow for a generally heteroscedastic/treatment-specific response to unobserved variables and instrument, so long as the relative effects are unchanged across potential outcomes. All functions in the proposition below are assumed to be measurable.

propSuppose that a): i) $Y(t) = g(t, \xi) + h(t) \psi(Z,\eta)$, $h(t) \ne 0$ for all $t \in \mathcal{T}$, with $\xi \rotatebox[origin=c]{90}{$\models$} (T, Z, \eta)$ and ii) MIV holds, strictly for some $z, z' \in \mathcal{Z}$ with $z'>z$; or b): i) $Y(t) = g(t, \xi, T) + h(t) \psi(Z,\eta)$ for all $t \in \mathcal{T}$, with $\xi \rotatebox[origin=c]{90}{$\models$} (T, Z, \eta)$, ii) $\frac{h(t)}{h(d)} > 0 ~ \forall t,d \in \mathcal{T}$; and iii) MIV holds. Then Assumption cMIV-p holds iff $\mathbb{E}[Y(t)|T = t, Z = z]$ are all monotone.

Note that whether or not $h(t) \ne 0$ is observable in the data for case $(a)$ and whether or not $h(t)/h(d)>0$ is also identified for $(b)$.

remarkThe monotonicity of observed outcomes has been routinely used in applied work to motivate the use of MIV condition (e.g. de2017effect). We show that, given that MIV holds and under a homogeneity condition, the observed monotonicity is instead equivalent to cMIV-p.
remarkcMIV is testable in Example (ref) because the reduced form expression has the form $b): i)$. It is testable in Example (ref) if instead of separately observing $\eta, Z$, employers on average observe a mixed signal of ability and effort, $s \equiv a Z + b\eta$ for some $a, b \in \mathbb{R}$.

A test of cMIV-p is thus the test of all $f_t(z) \equiv \mathbb{E}[Y(t)|T = t, Z = z]$ being monotone,

align*[align* omitted — 106 chars of source]

To obtain such a test, we may extend the procedure in chetverikov_2019\footnote{This test is developed for continuous $Z$, which is used in our application. Although the instrument is discretized at the estimation stage, the monotonicity of $\mathbb{E}[Y(t)|T = t, Z = z]$ for continuous $Z$ clearly implies the monotonicity of the discretized moments. The procedure we describe straightforwardly accommodates testing discrete instruments. As noted in chetverikov_2019, for discrete conditioning variable the test is a simple parametric problem, since the conditional moment function can be $\sqrt{n}-$consistently estimated at each point from the support.}. Denote the set of all observations with treatment level $t$ as $\mathcal{I}_t \equiv \{i \in [n]: T_i = t\}$ with $n_t \equiv |\mathcal{I}_t|$. Suppose $\phi^{t}_{n_t}$ is the corresponding Chetverikov's regression monotonicity test (or a corresponding parametric test for discrete $Z$) with the confidence level $\alpha_t \in (0; 0.5)$. We define the joint test as

align*[align* omitted — 72 chars of source]

Denote $\mathcal{P}^C$ to be the set of probability measures, such that for all $P \in \mathcal{P}^C$ and all $t \in \mathcal{T}$ the conditional probability measure given $T=t$ that $P$ generates satisfies the regularity conditions in Theorem 3.1 in chetverikov_2019. Similarly, let $\mathcal{P}^C_t$ be the set of all the conditional probability measures given $T = t$ that measures from $\mathcal{P}^C$ generate.

propIf $\Pi_{t \in \mathcal{T}} (1-\alpha_t) \geq 1-\alpha$, then \begin{align*} \inf_{P \in \mathcal{P}^C \cap \mathcal{H}_0} P[\phi_n = 0] \geq 1-\alpha + o(1), \end{align*} as $n \to \infty$.
proofNotice that each $\phi^t_{n_t}$ is a function of the observations from $\mathcal{I}_t$ only. Since $\mathcal{I}_t$ are mutually exclusive by construction and because the data are i.i.d., we have $P[\phi_n = 0] = \Pi_{t \in \mathcal{T}} P[\phi^t_{n_t} = 0]$. By the standard optimization argument, we have \begin{align*} \Pi_{t \in \mathcal{T}} \underset{P \in \mathcal{P}_t^C}{\inf} P[\phi^t_{n_t} = 0] \leq \underset{P \in \mathcal{P}^C}{\inf} \Pi_{t \in \mathcal{T}} P[\phi^t_{n_t}=0]. \end{align*} Theorem 3.1 from chetverikov_2019 and $\Pi_{t \in \mathcal{T}} (1-\alpha_t) \geq 1-\alpha$ then yield the result.
remarkOne may set $\alpha_t = 1 - (1-\alpha)^{1/N_T}$ as a baseline. If the domain knowledge suggests that for some treatments monotonicity is more likely to hold, one can set a higher $\alpha_t$ for them, so long as $\Pi_{t \in \mathcal{T}} (1-\alpha_t) \geq 1-\alpha$. This may improve the power of the test.

Returns to education in Colombia

Our data is comprised of 861492 observations from Colombian labor force. The sample represents a snapshot of those individuals who could be matched across the educational, formal employment and census datasets in 2021\footnote{Educational dataset was assembled by the testing authority Instituto Colombiano para la Evaluación de la Educación (ICFES), formal employment dataset comes from social security data based on Planilla Integrada de Liquidación de Aportes (PILA), whereas census data is handled by Departamento Administrativo Nacional de Estadística (DANE). The data was merged and anonymized by ICFES.}. For 664633 individuals from this dataset we observe their average lifetime wages, education level and Saber 5 or Saber 11 scores for Mathematics and Spanish language tests\footnote{Saber 5 and 11 tests are taken at different ages, but designed to be comparable between each other, which justifies merging them.}.

The outcome variable we consider ($Y_i$) is a log-wage, and $T_i$ is the education level. We distinguish four education levels: primary, secondary and high school, as well as `university'\footnote{$T_i$ is based on the number of years of schooling, $S_i$. If $S_i < 9$, set $T_i \equiv 0$ meaning the individual only graduated from primary school. $S_i \in [9;11)$ and $T_i \equiv 1$ correspond to completing compulsory education (secondary school), $S_i = 11$ and $T_i \equiv 2$ means that the individual is a high-school graduate, whereas $S_i > 11$ with $T_i \equiv 3$ means university education. Unfortunately, $S_i$ is capped at $17$ years in our sample, making it impossible to distinguish between those who continued to graduate education and those who just finished the $6-$years degree.}. Our measure of ability $Z_i$ is constructed as a CES aggregator of Mathematics and Spanish language test scores. To apply Theorem (ref), we then split $Z_i$ into deciles. Formally, we have

align*[align* omitted — 66 chars of source]
figure[figure omitted — 190 chars of source]

We first test whether cMIV is a reasonable assumption in our setup by implementing the test developed in Section (ref). To that end, we use the parameters and kernel functions recommended by chetverikov_2019 and focus on the theoretically most powerful procedure, the step-down approach. The estimated $p$-value of the test is $0.29$, see Table (ref). We thus conclude that cMIV-p is a credible assumption provided that MIV holds.

table[table omitted — 742 chars of source]

The data we study is rather noisy. One would expect a considerable measurement error in the construction of both treatment levels and the outcome variable\footnote{In particular, age is self-reported when filling an online questionnaire and appears to be of low quality, so we are forced to merge multiple cohorts.}. In line with that, the strongest form of cMIV is not sufficient to provide identification in the absence of further assumptions. While the resulting bounds are tighter than that under MIV, they remain uninformative.

To achieve identification, we supplement our assumptions with the MTR condition, which appears reasonable in our setting.\footnote{See MP2000 for a discussion of MTR in the context of returns to education. From the results in Section (ref) it also follows that MTR need only hold conditionally on all pairs in \(\mathcal{T} \times \mathcal{Z}\)—it does not need to hold almost surely.} The estimated bounds and confidence intervals are presented in Figure (ref). While MIV and cMIV-w remain uninformative, both cMIV-p and cMIV-s yield positive lower bounds on the ATEs. Under cMIV-p, the effect of obtaining a university education is estimated to be at least \(3.54\%\), and at least \(5.5\%\) under cMIV-s. These findings align with previous evidence: causal estimates for the U.S. (card1993using, propensity_score, angrist2011schooling) suggest returns of at least \(10\%\) for a four-year college degree. Recent evidence indicates that this figure may be substantially lower for Colombia gomez2022returns.

figure[figure omitted — 608 chars of source]

We also find significantly positive effects at other education stages, see Figure (ref). Further results and details on the estimation procedure are available in Appendix (ref).

\selectlanguage{english}

\nocite{*}

\addcontentsline{toc}{section}{6 References.}