EconBase
← Back to paper

Estimating Economic Models with Testable Assumptions: Theory and Applications

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

118,921 characters · 25 sections · 46 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Extending Economic Models with Testable Assumptions: Theory and Applications

\affil[1]{Nanjing University}

abstract\begin{spacing}{1.2} {This paper studies the identification and hypothesis testing in complete and incomplete economic models with testable assumptions. A testable assumption ($A$) gives interpretable empirical content to the economic model but it also carry the possibility that some distributions of observed outcomes may reject these assumptions. A way to avoid the data rejection problem is to find a relaxed assumptions ($\tilde{A}$) that cannot be rejected by any distribution of observed outcomes. We also want the identified set for the parameter of interest under $\tilde{A}$ is not changed when the original assumption $A$ is not rejected by the observed data distribution. I characterize the properties of such a relaxed assumption $\tilde{A}$ using a generalized notion of refutability and confirmability. I also propose a general method to construct such $\tilde{A}$. I apply my methodology to the instrument monotonicity assumption in Local Average Treatment Effect (LATE) estimation and to the sector selection assumption in a binary outcome Roy model of employment sector choice. In the LATE application, I use my general method to construct a relaxed assumption $\tilde{A}$ that can never be rejected, and the identified set for LATE is unchanged when $A$ holds. LATE is point identified under my extension $\tilde{A}$ in the application. In the binary outcome Roy model, I use my method to relax Roy's sector selection assumption and characterize the identified set for the binary potential outcomes as a polyhedron.} \end{spacing} {Keywords---} Incomplete Models; Refutability; LATE; Roy Model {JEL---} C12, C13, C18, C51, C52

Introduction

Empirical researchers often make convenient model assumptions in structural estimation. These assumptions usually come from economic theories or intuitions. For example, the `No Defiers' assumption in IA1994 assumes that the instrument has a monotone effect on the decision to take treatment; the `Pure Strategy Nash Equilibrium' assumption in BresnahanReiss1991 assumes only pure strategy Nash Equilibrium is played in a $2\times 2$ entry game; the `Perfect Self Selection' assumption in Roy1951 assumes employees perfectly observe their future earnings and choose a job sector to maximize discounted lifetime earnings. Such assumptions simplify the identification and estimation problems, and make the results easier to interpret. To study structural models and economic assumptions, I generalize the language of econometric structures koopmans1950identification to incomplete structures. A (generalized) econometric structure includes a distribution of some exogenously given latent and observed variables, and a correspondence from the distribution of these exogenous variables to a set of distributions of observed variables. Assumptions are restrictions on the economic structures to reflect empirical researchers' understanding of the economic environment.

Unfortunately, assumptions in the three examples above, when combined with some other reasonable assumptions, can be rejected by some distributions of observables kitagawa2015,mourifie2017testing,mourifie2018roy. When the imposed assumption is refuted by data, the econometrician have an empty identified set for the parameter of interest. As a result, the econometrician cannot give a useful interpretation of the economic environment. A refutable assumption $A$ also imposes challenges to the interpretation of hypothesis testing of parameter values. For example, when we reject the null hypothesis that the parameter equals zero under $A$, it can either be the case that the true parameter equals another value, or the case that $A$ is rejected data. These two cases cannot be distinguished but they have quite different interpretations.

A way to prevent the data rejection problem is to find a relaxed assumption $\tilde{A}$ so that no distributions of observables can reject $\tilde{A}$. We call this the non-refutability criterion. By imposing a non-refutable $\tilde{A}$ before confronting the data, practitioners avoid the ex-post possibility of finding data evidence against their assumption. Therefore, non-refutability is the first criterion for the relaxed assumption $\tilde{A}$ to satisfy. On the other hand, we also do not want to deviate from the old assumption $A$, since it still reflects the economic theory behind it. Specifically, given a parameter of interest $\theta$ as a function of structures, we want the (sharp) identified set under the relaxed assumption $\tilde{A}$ to be equal to the (sharp) identified set under $A$, when $A$ is not rejected by the observed distribution. In other words, we want to preserve the identified set.

This paper aims to do three things: First, I formalize and extend the definition of refutability and confirmability of assumptions in breusch1986 using the language of generalized economic structures. These definitions are useful to characterize properties of a relaxed assumption $\tilde{A}$. I characterize conditions that a relaxed assumption $\tilde{A}$ needs to satisfy so that: (a).No distribution of observables can be rejected under $\tilde{A}$; (b).The identified sets for the parameter of interest under $A$ and $\tilde{A}$ are the same whenever $A$ is not rejected by the observed data distribution. I show that when structures are complete, a relaxed assumption $\tilde{A}$ that satisfies the two properties above always exists. I also characterize a sufficient condition for the existence of $\tilde{A}$ in case of incomplete structures. The possible failure to find $~\tilde{A}$ in incomplete structures encourages researchers to complete the structures, and then find a nice relaxed assumption in the completed structure universe. When the structures are complete, I also provide a general method to construct $\tilde{A}$ from $A$.

Second, I discuss the problem of testing hypothesis on structural parameters' values. I show that any statistical test cannot achieve pointwise size control and test consistency simultaneously when the null hypothesis of structural parameters does not induce a partition of the space of distribution of observables. This is an ill-behaved null hypothesis, and policy decisions based on the result of hypothesis testing can be problematic. Conversely, with a well-behaved null hypothesis, which induces a partition of the space of distribution of observables, I show the existence of statistical tests that achieve pointwise size control and test consistency simultaneously under mild conditions. I also show that when the parameter of interest $\theta$ is point identified, a null hypothesis on the value of $\theta$ is always well-behaved. When structures are complete, I also provide a way to minimally extend (resp. shrink) the null hypothesis set such that the extend (resp. shrink) hypothesis is well-behaved.

Third, I look at 2 applications with complete and incomplete structures respectively. In the complete structure framework, I look at the identification of the local average treatment effects (LATE). kitagawa2015 provides the sharp testable implication of the Imbens and Angrist Monotonicity assumption (IA-M). Therefore, practitioners should anticipate the IA-M to be rejected by some distributions of observables. I provide several relaxations of the IA-M that cannot be rejected by any distribution of observables. The identified sets for LATE under these relaxed assumptions equal the identified set for LATE under the IA-M whenever the IA-M is not rejected by the distributions of observables. One relaxed assumption allows for defiers, and relaxes the independent instrument assumption. The logic of the relaxed assumption is to allow a minimal mass of defiers. It can be shown that the relaxed assumption not only preserves the identified set for LATE, but also preserves the identified set for other parameters of interests such as ATE or ATT. LATE is point identified under this relaxed assumption. I provide an estimator of the LATE which has a normal limit distribution. {To emphasis the fact that a non-refutable relaxation is not unique, I propose two other relaxations, each one of which may be preferable in some contexts.} I apply the method to card1993using, and I show the local average treatment effect of education on earnings. Compared to naively using identification result under the IA-M assumption, my method delivers more reasonable sign and scale for the LATE estimates.

For incomplete structures, I look at a binary outcome job sector selection model with a monotone instrument. In the sector choice model, the Roy assumption does not specify the sector choice rule in case of ties, which may lead to multiple predicted distributions of observables. After completing the structures, I then use a `minimal efficiency loss' criterion to characterize the relaxed assumption. The identified set of job sector potential outcome distribution can be characterized as a polyhedron.

\paragraph{Related Literature} masten2018 propose an ex-post way to salvage a refutable assumption $A$. Their ex-post method characterizes a relaxation of $A$ after the distribution of observables is realized. This paper also relates to the literature that relaxes assumption make model robust to misspecification. In the macroeconomic literature, researchers use robust control to avoid the misspecification issue in their baseline model (see hansen2006robust and hansen2007recursive). The Robust control approach aims to accommodate local perturbations to the baseline model rather than to solve the refutability of the baseline model.\footnote{The perturbation is usually measured by relative entropy the in macroeconomic literature.} It may be true that the baseline model is not refutable by any data distribution. See bonhomme2018minimizing and christensen2019counterfactual for more discussion.

{This paper also contributes to the literature that relaxes the IA-M assumption. chaisemartin2018Defier discusses the economic meaning of the conventional LATE quantity $LATE^{Wald}\equiv ({E[Y_i|Z_i=1]-E[Y_i|Z_i=0]})/({E[D_i|Z_i=1]-E[D_i|Z_i=0]})$ when there are defiers. He shows that the $LATE^{Wald}$ identifies the net average treatment effect of a subgroup of compliers after deducting the average treatment effect of defiers.}

The rest of the paper is organized as follows. Section 2 describes a theory of characterizing a refutable assumption $\tilde{A}$, finding a relaxed assumption $\tilde{A}$ and hypothesis testing. Core definitions in Section 2 are followed by shadowed links where their corresponding illustrations can be found. Section 3 applies complete structure theory to the model in IA1994. Section 4 applies the theory to a binary outcome Roy model. Main proofs are collected in Appendices. Additional proofs and results are collected in the Online Appendices.

Notations

Throughout this paper, I use $X$ to denote the vector of observed variables, and I use $F$ to denote the distribution of $X$. I use $\epsilon$ to denote the vector of latent variables and some observed variables, and I use $G$ to denote the distribution of $\epsilon$. I use $s$ to denote an economic structure. I use $A$ to denote a refutable assumption and use $\tilde{A}$ to denote a relaxed assumption. I use $\mathbb{F}_n$ to denote the empirical distribution of $F$.

A Theory of Identification and Hypothesis Testing

In this section, I develop a theory for identification and refutable assumptions. I will then discuss the problem of hypothesis testing. I start with a definition of the observation space.

definitionThe observation space $\mathcal{F}$ is the collection of all possible distribution of $F(X)$. \colorbox{lightgray}{See (ref) for an illustration.}

The distribution of observables $F(X)$ is generated by some distribution of underlying random vectors $\epsilon$ through some mapping $M$. A pair of a distribution of $\epsilon$ and a mapping $M$ is called an econometric structure. The following definition of econometric structure is a reformulation of the economic structure defined in koopmans1950identification and jovanovic1989. Since in most econometric problems, we focus on the distribution of outcomes $F$ instead of how each $X$ is related to $\epsilon$, I directly define the mapping $M$ as a correspondence from the space of underlying variable distributions to $\mathcal{F}$.

definitionAn econometric structure (Model) $s=(G^s,M^s)$ consists of a distribution $G^s$, and an outcome mapping $M^s$. Let $\mathcal{G}$ denote the space of all possible regular distributions of $G^s(\epsilon)$. The outcome mapping $M^s$ is a correspondence $M^s:\mathcal{G}\rightrightarrows \mathcal{F}$. \colorbox{lightgray}{See ((ref)),((ref)), ((ref)) for illustrations.}
definitionA structure universe $\mathcal{S}$ is a collection of structures such that $\cup_{s\in\mathcal{S}} M^s(G^s)=\mathcal{F}$, and an assumption $A$ is a subset of $\mathcal{S}$. \colorbox{lightgray}{See ((ref)),((ref)),((ref)), ((ref)) for illustrations.}

The mapping $M^s$ of structure $s$ relates the distribution $\epsilon$ to the distributions of $X$. Since the distributions of $X$ are generated by the distributions of $\epsilon$, we call $\epsilon$ the primitive variables. Definition (ref) allows overlap between $X$ and $\epsilon$. A collection of structures is called a structure universe. We want to learn the distribution of $\epsilon$ and the mapping $M$ from the distribution of observables $F$.

Here I explicitly distinguish the structure universe $\mathcal{S}$ and the assumption $A$, though both are just a collection of structures. The structure universe $\mathcal{S}$ is the paradigm that can span different empirical contexts. On the other hand, an assumption $A$ places constraints that are suitable for a particular empirical context, or convenient for empirical analysis.

The condition $\cup_{s\in\mathcal{S}}M^s(G^s)=\mathcal{F}$ requires that all possible distributions of observables can be generated by some structure in the universe. Moreover, no distribution outside $\mathcal{F}$ can be generated by $\mathcal{S}$.

definitionA structure $s:=(M^s,G^s)$ is called complete if $M^s(G^s)$ is a singleton. Otherwise it is called incomplete. A universe $\mathcal{S}$ is called complete if every structure $s$ in $\mathcal{S}$ is complete, otherwise it is incomplete.

The definition of completeness is slightly different from the definition of completeness in tamer2003incomplete. tamer2003incomplete defines a model to be complete if the mapping from $\epsilon$ to $X$ is a singleton and non-empty, and incomplete if the mapping has multiple outputs, and incoherent if the mapping generates no output. Here, my definition of economic structure does not specify the mapping from each $\epsilon$ to $X$. Instead, I consider the mapping from the distribution of $\epsilon$ to the distribution of $X$. If the probability of multiple outcome is non-zero, an incomplete model in tamer2003incomplete implies an incomplete structure in my definition.

My definition of structure, however, does not have a corresponding terminology for incoherent model. This is because I require $\cup_{s\in\mathcal{S}}M^s(G^s)=\mathcal{F}$ to hold. chesher2012simultaneous propose four ways to deal with model incoherence and derive the distribution of observables under the model. My definition of $M^s$ as the mapping between distributions can be viewed as the consequences of chesher2012simultaneous.

definition(Breusch) An assumption $A$ is called refutable if there exists an $F\in \mathcal{F}$ such that $F\notin\cup_{s\in A}M^s(G^s)$. An assumption $A$ is called confirmable if there exists an $F\in \mathcal{F}$ such that $F\notin\cup_{s\in A^c}M^s(G^s)$.

The notions of refutability and confirmability of a complete structure are given in breusch1986. If an assumption is refutable, then there exists some $F$ that can reject $A$. The notions of refutability and confirmability are stated in terms of the observation space $\mathcal{F}$. Equivalently, we can characterize refutability and confirmability in terms of the structure universe $\mathcal{S}$. To do this, I first define the non-refutability and confirmation sets associated with $A$.

definition(non-refutability set) \\ The strong non-refutability set associated with $A$ under $\mathcal{S}$ is defined as \[\mathcal{H}^{snf}_\mathcal{S}(A)=\left\{s\in\mathcal{S}: M^s(G^s)\subseteq \cup_{s^*\in A} M^{s^*}(G^{s^*}) \right\}.\] The weak non-refutability set associated with $A$ under $\mathcal{S}$ is defined as \[\mathcal{H}^{wnf}_\mathcal{S}(A)=\left\{s\in\mathcal{S}: M^s(G^s)\cap \left(\cup_{s^*\in A} M^{s^*}(G^{s^*})\right)\ne \varnothing \right\}.\] \colorbox{lightgray}{See Lemma (ref), ((ref)) and Proposition (ref) for illustrations.}

To accommodate the incomplete structures, we define two types of non-refutability sets. We call $\mathcal{H}_{\mathcal{S}}^{snf}(A)$ the strong non-refutability set associated with $A$, because if the true structure $s$ is in $\mathcal{H}_{\mathcal{S}}^{snf}(A)$, then for any distributions of observables in $M^s(G^s)$, we cannot refute $A$. In contrast, we call $\mathcal{H}_{\mathcal{S}}^{wnf}(A)$ the strong non-refutability set associated with $A$, because if the true structure $s$ is in $\mathcal{H}_{\mathcal{S}}^{wnf}(A)$, then for some distributions of observables in $M^s(G^s)$, we cannot refute $A$.

definition(Confirmation set) \\ The strong confirmation set associated with $A$ under ${S}$ is defined as \[ \mathcal{H}_{S}^{scon}(A)=\left\{s\in{S}: M^s(G^s)\subseteq \cap_{s^*\in A^c} \left(M^{s^*}(G^{s^*})^c\right)\right\}. \] The weak confirmation set associated with $A$ under ${S}$ is defined as \[ \mathcal{H}_{S}^{wcon}(A)=\left\{s\in{S}: M^s(G^s)\cap\left[ \cap_{s^*\in A^c} \left(M^{s^*}(G^{s^*})^c\right)\right]\ne \varnothing\right\}. \] \colorbox{lightgray}{See Proposition (ref) for an illustration.}

The strong confirmation set is the collection of structures that cannot be observationally equivalent to any structures outside $A$ for any observed distribution $F$. Weak confirmation set is the collection of structures that cannot be observationally equivalent to any structures outside $A$ for some observed distribution $F$. I call $\mathcal{H}^{scon}_{\mathcal{S}}(A)$ (resp. $\mathcal{H}^{wcon}_{\mathcal{S}}(A)$) the strong (resp. weak) confirmation set associated with $A$, because if the true structure $s$ is in $\mathcal{H}^{scon}_{\mathcal{S}}(A)$ (resp. $\mathcal{H}^{wcon}_{\mathcal{S}}(A)$), then for all (resp. some) distribution of observables in $M^s(G^s)$, we can confirm that the true structure must lies in $A$. In particular, \[\mathcal{H}_\mathcal{S}^{scon}(A)\subseteq \mathcal{H}_\mathcal{S}^{wcon}(A)\subseteq A\subseteq \mathcal{H}_\mathcal{S}^{snf}(A)\subseteq \mathcal{H}_\mathcal{S}^{wnf}(A).\] When the structure universe $\mathcal{S}$ is complete, $M^s(G^s)$ is always a singleton, and $\mathcal{H}_\mathcal{S}^{scon}(A)= \mathcal{H}_\mathcal{S}^{wcon}(A)$, $\mathcal{H}_\mathcal{S}^{snf}(A)= \mathcal{H}_\mathcal{S}^{wnf}(A)$. The following proposition helps to interpret the confirmation sets associated with $A$ as the non-refutability sets associated with $A^c$. It also shows that the strong non-refutability set and weak confirmation set as operation are idempotent.

propThe following holds: 1. $\left[\mathcal{H}_\mathcal{S}^{snf}(A)\right]^c=\mathcal{H}_\mathcal{S}^{wcon}(A^c)$; 2. $\left[\mathcal{H}_\mathcal{S}^{wnf}(A)\right]^c=\mathcal{H}_\mathcal{S}^{scon}(A^c)$; 3. $\mathcal{H}_\mathcal{S}^{wcon}(\mathcal{H}_\mathcal{S}^{wcon}(A))=\mathcal{H}_\mathcal{S}^{wcon}(A)$; 4. $\mathcal{H}_\mathcal{S}^{snf}(\mathcal{H}_\mathcal{S}^{snf}(A))=\mathcal{H}_\mathcal{S}^{snf}(A)$.

The definition of refutability and confirmability in breusch1986 is defined on the outcome space $\mathcal{F}$, but we can also characterize it on the structure universe $\mathcal{S}$.

propAn assumption $A$ is refutable if and only if $\mathcal{H}_\mathcal{S}^{snf}(A)\ne \mathcal{S}$. An assumption $A$ is confirmable if and only if $\mathcal{H}_\mathcal{S}^{wcon}(A)\ne \varnothing$.

By definition, we should have $\cup_{s\in \mathcal{H}_{\mathcal{S}}^{snf}(A)} M^{s}(G^{s})=\cup_{s\in A} M^{s}(G^{s})$, so $\mathcal{H}_{\mathcal{S}}^{snf}(A)$ is refutable if and only if $A$ is refutable. In many cases, it is easy to check whether $\mathcal{H}_{\mathcal{S}}^{snf}({A})=\mathcal{S}$ in Proposition (ref) than to check Definition (ref).

Identification Problem

In many empirical studies, we want to find the value of a parameter of interest rather than a class of structures that are consistent with data. This parameter can be a moment of unobserved primitive variables, or a counterfactual outcome of the structure. The parameter of interest can also give interpretation on the causal relation between outcome variables and primitive variables. Imposing strong assumptions helps to restrict the set of data-consistent parameter values, but an imposed assumption $A$ as in Definition (ref) may be rejected by some distribution of observables. Therefore, in many empirical studies, researchers often first present some summary statistics that justify the assumption. If the assumption is rejected by the data, researchers can move to another assumption. This is an ex-post way of choosing a relaxed assumption. There are two major problems with this approach. First, such justifications are heuristic pre-testing procedures of assumption $A$, and any subsequent inference on the parameter of interest may have incorrect size control due to pre-testing. Second, researchers do not specify what they will do if $A$ is rejected. Most likely they will choose another assumption that will not be rejected by the data. To avoid the pre-testing issue, I propose to solve the problem from an ex-ante perspective, i.e. impose a non-refutable assumption before any distribution of observables is realized.

I first formalize the definition of an identification system and discuss how to deal with an existing situation, where assumption $A$ may be rejected by the data.

definitionA parameter of interest $\theta$ is a function $\theta: \mathcal{S}\rightarrow \Theta$, where $\Theta$ is the parameter space. The identified set for $\theta$ is a correspondence $\Theta^{ID}_A:\mathcal{F} \rightrightarrows \Theta$ such that \begin{equation} \Theta_A^{ID}(F)=\{\theta(s): s\in A \quad and \quad F\in M^s(G^s) \}. \end{equation} We call $(\mathcal{S},A,\theta,\Theta^{ID}_A)$ an identification system. \colorbox{lightgray}{See ((ref)) for $\theta$ and ((ref)) for an illustration of $\Theta^{ID}_A$.}

A parameter of interest can take a very general form. It can be the structure $s$ itself, or it can be a counterfactual outcome. For example, suppose $M^s$ is known up to a finite dimensional vector: $M^s(\cdot)=M(\cdot\,;(\beta_1,...,\beta_k))$. Further suppose the objective of our counterfactual analysis is to find the predicted distribution of observables when $\beta_1=0$. Then the parameter of interest $\theta$ can be defined as $\theta(s)= M(G^s;(0,...,\beta_k))$. For a parameter of interest $\theta$, $\Theta^{ID}_A(F)$ is the set of parameters that are compatible with the data. For an empirical researcher, the main concern of the partial identification method is the possibility of an empty identified set. Here I characterize the equivalent condition of an empty identified set.

propAn assumption $A$ is non-refutable if and only if for any parameter of interest $\theta$, the associated identified set $\Theta_A^{ID}(F)\ne\varnothing$ holds $\forall F\in \mathcal{F}$.

Here is an intuition of Proposition (ref): If we have a non-empty identified set for an $F$, then there must exist a structure in $A$ that rationalizes $F$. Since this is true for all $F$, $A$ is non-refutable. Conversely, if $A$ is non-refutable, then for any $F$ we can find a structure $s\in A$ to rationalize $F$, and the corresponding $\theta(s)$ must lie in the identified set.

Relaxed Assumption Approach

For a refutable assumption $A$, there exist some $F$ such that $\Theta_A^{ID}(F)=\varnothing$. This can be unsatisfying because empirical researchers cannot directly interpret the distribution of $\epsilon$. To avoid this, before seeing any outcome distribution, a practitioner can impose a relaxed assumption $\tilde{A}$ such that $\mathcal{H}_\mathcal{S}^{snf}(\tilde{A})= \mathcal{S}$ and $A\subseteq \tilde{A}$.

definitionGiven structural universe $\mathcal{S}$, a refutable assumption $A$ and $\theta$, we call $\tilde{A}$ \begin{enumerate} • a well-defined extension, if $A\subseteq \tilde{A}$ and $\mathcal{H}_\mathcal{S}^{snf}(\tilde{A})=\mathcal{S}$; • a $\theta$-consistent extension, if $\tilde{A}$ is a well-defined extension, and $\Theta^{ID}_{\tilde{A}}(F)=\Theta_A^{ID}(F)$ whenever $\Theta_A^{ID}(F)\ne \varnothing$; \colorbox{lightgray}{See Proposition (ref) for an illustration.} • a strong extension, if for any parameter of interest $\theta^*$ defined in Definition (ref), $\tilde{A}$ is a $\theta^*$-consistent extension.\colorbox{lightgray}{See Proposition (ref) and (ref) for an illustration.} \end{enumerate}

These three definitions are nested. A well-defined extension ensures that the identified set will never be empty; A $\theta$-consistent extension preserves the identified set for a parameter of interest $\theta$. A strong extension moreover ensures that the identified set for any parameter of interest will be preserved. In different empirical settings, researchers' parameters of interest can differ. If a strong extension is found, researchers can use this extension across different empirical contexts. The following proposition gives a characterization of whether $\tilde{A}$ is a strong extension.

propGiven $\mathcal{S}$, suppose $A$ is refutable and $\tilde{A}$ is a well-defined extension, then $\tilde{A}$ is a strong consistent extension if and only if $\mathcal{H}_{\mathcal{S}}^{wnf}(A)\cap \tilde{A} =A$.

In a complete structure universe, we can always find a strong extension $\tilde{A}$. This is a major difference between complete and incomplete structure universe.

propIf $\mathcal{S}$ is a complete structure universe, then $\tilde{A}=A\cup [\mathcal{H}_\mathcal{S}^{snf}(A)]^c$ is a strong extension of $A$. Moreover, any strong extension $\tilde{A}'$ is a subset of $\tilde{A}$. We call this $\tilde{A}$ the maximal strong extension of $A$.

In other words, $\tilde{A}=A\cup [\mathcal{H}_\mathcal{S}^{snf}(A)]^c$ is the strong extension that puts the least structural assumption outside $\mathcal{H}_\mathcal{S}^{snf}(A)$. Should we always use $\tilde{A}=A\cup [ \mathcal{H}_\mathcal{S}^{snf}(A)]^c$ as the choice of strong extension when the structural universe $\mathcal{S}$ is complete? Unfortunately, using the maximal strong extension will lead to very badly behaved identified set $\Theta_{\tilde{A}}^{ID}(F)$. Suppose $\mathcal{F}$ is equipped with some metric $d$, and $F_0$ is on some part of the boundary of the set of predicted observable distribution $\cup_{s\in {A}} M^s(G^s)$. It is possible that $\Theta_{\tilde{A}}^{ID}(F_0)$ gives an informative bound (i.e. $\Theta_{\tilde{A}}^{ID}(F_0)\ne \Theta$) on the parameter of interest , but for an $F'\in\left[\cup_{s\in {A}} M^s(G^s) \right]^c$ that is arbitrarily close to $F_0$, the identified set $\Theta_{\tilde{A}}^{ID}(F')$ is uninformative (i.e. $\Theta_{\tilde{A}}^{ID}(F')=\Theta$). See the identification result in Proposition (ref) for an illustration. This raises two concerns. First, at the identification level, the interpretation of an uninformative identified set $\Theta_{\tilde{A}}^{ID}(F')$ under the maximal strong extension $\tilde{A}$ is not very different from an empty identified set $\Theta_{A}^{ID}(F')=\varnothing$ under the original assumption $A$. An uninformative identified set says that any parameter value of $\theta$ is compatible with data, while an empty identified set says that no parameter value of $\theta$ is compatible with data. In either case, the identification result does not help us to interpret the environment. Second, we may get spurious informative inference result due to sampling error. When $F'$ is close to the boundary of $\cup_{s\in {A}} M^s(G^s)$ but not in it, sampling error may lead us to a spurious but informative bound $\Theta_{\tilde{A}}^{ID}(F')$, even if the true identified set should have been uninformative. In other words, estimated identified set is not consistent. The maximal strong extension $\tilde{A}=A\cup [\mathcal{H}_\mathcal{S}^{snf}(A)]^c$ solves the refutability issue, but it imposes too few constraints outside $A$ to generate informative result on $\theta$.

For an incomplete structure universe, there is a gap between the strong and weak non-refutability sets. As a result, we may not be able to find a strong extension of $A$.

propIf { $\left(\cup_{s\in \mathcal{H}_\mathcal{S}^{wnf}(A)}M^s(G^s)\right)\cap \left(\cup_{s\in [\mathcal{H}_\mathcal{S}^{wnf}(A)]^c}M^s(G^s)\right)=\varnothing $} and $\mathcal{H}_{\mathcal{S}}^{wnf}(A)\backslash \mathcal{H}_{\mathcal{S}}^{snf}(A)\ne\varnothing$ hold, then there does not exist a strong extension of $A$.

This situation happens when there is a nesting relation between $A$ and $A^c$: suppose for each structure $s\in A$, we can find an $s^*\in A^c$ such that $M^s(G^s)\subseteq M^{s^*}(G^{s^*})$; and for every $s^*\in A^c$ we can find an $s\in A$ such that $M^s(G^s)\subseteq M^{s^*}(G^{s^*})$. If $A$ is refutable, then $\mathcal{H}_{\mathcal{S}}^{snf}(A)\ne \mathcal{S}$. However the nesting relation implies $\mathcal{H}_{\mathcal{S}}^{wnf}(A)=\mathcal{S}$. Then both conditions in Proposition (ref) hold, and there exists no strong extension of $A$.

exampleFigure (ref) illustrates a simple case where $\mathcal{S}=\{s_1,s_2,s_3\}$ and the assumption set $A=\{s_1\}$. There is a nesting relation between the predicted outcome distributions of $s_2$ and $s_1$, and the predicted outcome distributions of $\{s_1,s_2\}$ and $s_3$ are disjoint. This satisfies the condition in Proposition (ref). The only well-defined extension of $A$ must be $\tilde{A}=\{s_1,s_2,s_3\}$, since we need to include $s_2$ to predict $\tilde{F}$. It is easy to check $\tilde{A}\cap \mathcal{H}_{\mathcal{S}}^{wnf}(A)=\{s_1,s_2\}\ne A$, so $\tilde{A}$ cannot be a strong extension. \begin{figure}[H] \caption{Impossibility to find a strong extension.} \end{figure}

Minimal Deviation Method

In many cases, we can find a function $m:\mathcal{S}\rightarrow \mathbb{R}_+$ such that $m(s)=0$ for all $s\in A$ \footnote{See ((ref)), ((ref)) for example.}. While an assumption $A$ may be refutable, we impose $A$ in the first place because it reflects economic theory suitable in the empirical context. Therefore we would consider a departure from $A$ is abnormal and is against the economic intuition behind $A$. I extend our assumption to allow for a minimal departure from the baseline assumption $A$. This way to relax assumption $A$ is called the minimal deviation method. Formally, suppose the refutable assumption $A$ can be written as an intersection of several larger assumptions: $A=\cap_{j=1}^J A_j$. This representation allows us to consider a departure from a particular $A_j$.

definitionFix an index $j\in \{1,2,...J\}$. A relaxation measure of departure from $A_j$ with respect to $\{A_l\}_{l\ne j}$ is a function $m_j: \mathcal{S}\rightarrow \mathbb{R}_+\cup\{+\infty\}$ such that $m_j(s)=0$ for all $s\in A_j$. We say $m_j$ is well-behaved if for any $F\in \mathcal{F}$, there exists a structure $s^*\in \cap_{l\ne j}A_l$ such that \[ m_j(s^*)=\inf\{m_j(s): F\in M^s(G^s) \quad and \quad s\in \cap_{l\ne j}A_l\}, \] and $m_j(s^*)<\infty$. \colorbox{lightgray}{See ((ref)) and ((ref)) for illustrations.}

We want the relaxation measure to be well-behaved such that when we push the deviation to infinity, we can generate any distributions in $\mathcal{F}$. The well-behaved condition ensures that there exists a structure in $\cap_{l\ne j}A_l$ that can achieve the minimal measure. This is essential for the construction of a extension using $m_j$. See Proposition (ref) is not well behaved if use full independence} for examples of ill-behaved measures. Now I construct the minimal deviation extension $\tilde{A}$.

definitionFix an index $j\in \{1,2,...J\}$ and a well-behaved relaxation measure $m_j$. For any $F$, let $m^{min}(F)\equiv \inf\{m_j(s): F\in M^s(G^s) \,\, and \,\, s\in \cap_{l\ne j}A_l\}$ be the minimal deviation for the observable distribution $F$. We call $\tilde{A}=\cup_{F\in \mathcal{F}} \left\{s\in \cap_{l\ne j}A_l: m_j(s)=m^{min}(F),\,\,F\in M^s(G^s)\right\}$ the minimal deviation extension of $A$ under $m_j$. \colorbox{lightgray}{See the constructions of Assumptions (ref) and (ref) for illustrations.}

The construction in Definition (ref) only relaxes the assumption $A_j$ and keeps other assumptions unchanged. The following proposition shows a way to check whether a minimal deviation extension is $\theta$-consistent or strong consistent.

propLet $A=\cap_{l=1}^J A_l$ and fix an index $j\in \{1,2,...,J\}$. Suppose $m_j$ is a well-behaved relaxation measure with respect to $\{A_l\}_{l\ne j}$. Then a minimal deviation extension $\tilde{A}$ under $m_j$ is \begin{enumerate} • $\theta$-consistent if $\Theta_A^{ID}(F)$ defined in ((ref)) satisfies \[\Theta_A^{ID}(F)=\{\theta(s):\quad F\in M^s(G^s)\quad and \quad s\in \left(\cap_{l\ne j}A_l\right) \quad and \quad m_j(s)=0\};\] • a strong extension if $A_j=\{s\in \cap_{l\ne j}A_l: \,\,m_j(s)=0\}$. \end{enumerate}

Multiplicity of Extensions and the Extension Choice

To facilitate the discussion of criteria of choosing among multiple extensions, I assume $\mathcal{F}$ is endowed with a metric $d_F$. Moreover, I fix the parameter of interest $\theta$, and assume $\Theta$ is endowed with a metric $d_\theta$.

In some cases, there can be multiple ways to write an assumption, i.e. $A=A_j\cap(\cap_{l\ne j} A_l)=A_j\cap(\cap_{l\ne j} A_l^\prime)$. Fix a $j$, even if we use the same measure $m_j$, since the minimal deviation is defined with respect to $\{A_l\}_{l\ne j}$, the extension can differ when using a different representation. It should also be noted that to check whether $m_j$ is a well-defined relaxation measure, we need to look at $\{A_l\}_{l\ne j}$. This means $m_j$ can be a well-defined relaxation measure with respect to $\{A_l\}_{l\ne j}$ but not $\{A_l^\prime\}_{l\ne j}$. Here I leave the choice of representation of the assumption to researchers and discuss the issue of multiple extensions that arises from two aspect: which assumption to relax and the choice of relaxation measure.

Each minimal deviation extension $\tilde{A}$ corresponds to an index $j$ and a relaxation measure $m_j$. Given a representation of $A=\cap_{j=1}^J A_j$, we can choose which sub-assumption $A_j$ to relax, and we can also choose the relaxation measure $m_j$. Different choices of which assumption $j$ to relax, and different relaxation measures $m_j$ will result in different relaxed assumptions. Moreover, relaxed assumptions constructed from different relaxation measures can be non-nested with each other. Here, I discuss two criteria to choose an assumption among non-nested relaxed assumptions.

First, the relaxed assumption $\tilde{A}$ should be suitable for the empirical context. Given the empirical context, if we can find an economic story such that the $j$-th assumption in $A=\cap_{j=1}^J A_j$ fails, we will focus on finding a well-behaved relaxation measure $m_j$ corresponding to $A_j$. Second, we want some continuity property of the identified set with respect to the outcome distribution $F$.

property[Identified Set Continuity] The identified set $\Theta_{\tilde{A}}^{ID}(F)$ under $\tilde{A}$ is a continuous correspondence\footnote{Recall that a correspondence $\Gamma :A\rightrightarrows B$ is called upper hemicontinuous at the point $a$ if for any open neighborhood $V$ of $\Gamma(a)$ there exists a neighborhood $U$ of $a$ such that for all $a'\in U$, $\Gamma(a')$ is a subset of $V$. A correspondence $\Gamma :A\rightrightarrows B$ is called lower hemicontinuous at the point $a$ if for any open set $V$ intersecting $\Gamma(a)$ there exists a neighborhood $U$ of a such that $\Gamma(a')$ intersects $V$ for all $a'\in U$. A continuous correspondence is both upper and lower hemicontinuous.} from $\mathcal{F}$ to $\Theta$.

Recall that we fix the parameter of interest $\theta$ in the beginning of this section. An $\tilde{A}$ that satisfies Property (ref) for $\theta$ may fail Property (ref) for a different parameter of interest. Without the continuity property, a consistent estimator of the identified set may not exist, and the identified set can be spuriously informative due to sampling error. Examples of discontinuous and continuous identified set correspondences can be found in Proposition (ref). Sufficient conditions to check Property (ref) and the further reasoning of Property (ref) can be found in Appendix (ref).

Structure Completion

We have seen that for an incomplete structure universe, a refutable assumption $A$ may not have a strong extension. This is because in incomplete structures, we are agnostic about how distributions of outcomes are selected. If a structure $s$ has two predicted outcome distribution $M^s(G^s)=\{F_1,F_2\}$, we consider a completion procedure that separates $s$ into two complete structures $s_1^*$ and $s_2^*$ such that $M^{s_1^*}(G^{s_1^*})=\{F_1\}$ and $M^{s_2^*}(G^{s_2^*})=\{F_2\}$. The completion procedure then allows us to distinguish $s_1^*$ from $s_2^*$ by observing either $F_1$ or $F_2$.

definitionGiven a structure $s=(M^s,G^s)$, let $\mathcal{C}(s)$ be the collection of all single-valued correspondences from $\mathcal{G}$ to $\mathcal{F}$ such that $M^*\in \mathcal{C}(s)$ implies $M^*(G^s)\subseteq M^s(G^s)$. We call \begin{equation} {\mathcal{S}}^*=\{({M}^{*},G): s\in \mathcal{S},\quad G=G^s,\quad M^{*}\in \mathcal{C}(s)\} \end{equation} the completion of $\mathcal{S}$. \colorbox{lightgray}{See ((ref)) for an illustration of $\mathcal{C}(s)$.}

Definition (ref) considers all possible completions $C$. The cardinality of $\mathcal{C}(s)$ is the same as that of $M^s(G^s)$. The completion procedure is without loss of generality, since all possible selections are considered. The key property is that for any parameter of interest, the identified set is not changed if the parameter of interest in the completed structure is properly defined in the following way.

propLet $\mathcal{S}$ be an incomplete structural universe, let $A$ be any assumption and $\theta$ be any parameter parameter of interest. Let $\mathcal{S}^*,A^*,\theta^*,{\Theta}_{A^*}^{*ID}$ be defined in the following way: \begin{equation} \begin{split} &{\mathcal{S}}^* \,\, is the completion of \,\,\mathcal{S},\\ &{A}^*=\{s^*:s^*=({M}^{s^*},G^{s^*}),\quad s\in \mathcal{S}\cap A,\quad G^{s^*}=G^s, \quad {M}^{s^*}\in \mathcal{C}(s)\},\\ &{\theta}^*(s^*)=\theta(s)\quad for\,\,all\,\, s=(M^s,G^s),\, s^*=({M}^{s^*},G^{s^*}) such that M^{s*}\in \mathcal{C}(s)\,\,and\,\, G^{s^*}=G^{s}\\ &\Theta^{*ID}_{A^*}(F)= \{\theta^*(s^*): F\in M^{s^*}(G^s),\,\, s^*\in A^*\}. \end{split} \end{equation} Then $\Theta_A^{ID}(F)={\Theta}^{*ID}(F)$ for all $F$.

In many cases, finding a strong extension is not feasible for an incomplete structure universe, but feasible for its completion. See Proposition (ref) for an illustration.

The Hypothesis Testing Problem

In empirical research, a commonly asked question is whether we can tell if the true value of the parameter of interest lies in a set, which can be written as a hypothesis $H$ on the parameter value. In this section, I consider the following formulation of a hypothesis $H$ on a structural parameter $\theta$ under a non-refutable assumption $\tilde{A}$: $H=\{s\in \tilde{A}: \theta(s)\in \Theta^0\}$, where $ \Theta^0$ is a parameter value set. The implicit alternative is $H^c\cap \tilde{A}$. Here I only consider non-refutable assumption $\tilde{A}$. If an assumption $A$ is refutable and cannot generate all distributions of observables, then for some distributions of observables, we cannot make say at least one of $H$ and $H^c\cap A$ holds true.

Policy makers sometimes use the result of hypothesis testing of a parameter value to guide their policy decisions. This decision procedure is called the `inference-based' approach in manski2019econometrics and is a conventional practice in medical treatment policy decision manski2020covid. However, the `inference-based' policy decision approach can be problematic if the hypothesis on parameter value $H$ does not induce a partition on the observation space $\mathcal{F}$: if both $H$ and $H^c\cap \tilde{A}$ can generate some observed distribution $F_0$, then we cannot tell whether $H$ holds by observing $F_0$. To formally discuss this issue, I first discuss the `hypothesis testing' problem assuming that I know the distribution of observables. I call this the binary decision problem \footnote{The same problem is called the binary choice problem in manski2019econometrics. To avoid the confusion with concepts in the discrete choice literature, I slightly change the name.}.

definitionWe say a hypothesis $H$ can be decided by $F$ under $\tilde{A}$ if either of the following conditions holds: \begin{enumerate} • $F\notin\cup_{s^*\in [H^c\cap \tilde{A}] } M^{s^*}(G^{s^*})$; • $F\notin \cup_{s^*\in H } M^{s^*}(G^{s^*}).$ \end{enumerate} $H$ is called weakly binary decidable under $\tilde{A}$ if there exists an $F\in \mathcal{F}$ such that $H$ can be decided by $F$. $H$ is called strongly binary decidable under $\tilde{A}$, if for all $F\in\mathcal{F}$, $H$ can be decided by $F$.

If condition 1 in Definition (ref) holds, it implies that the true structure $s$ that generates $F$ must be in $H$, since $F$ cannot be predicted by $H^c\cap \tilde{A}$, and this confirms $s\in H$; if condition 2 in Definition (ref) holds, it implies that the true structure $s$ cannot be in $H$, since $F$ cannot be predicted by $H$, and this refutes $s\in H$. If both conditions fail, it means $F$ can be predicted by structures both inside and outside $H$, which creates an ambiguity in the binary decision problem. If $H$ can be decided by any $F$, we say it is strongly binary decidable.

exampleConsider a simple linear regression model \begin{equation} Y_i=\beta_0+\beta_1 Z_i+\eta_i, \end{equation} where primitive variables are $(Z_i,\eta_i)$, observed variables are $(Y_i,Z_i)$. The correspondence $M^s$ is determined by ((ref)) \footnote{ The image of mapping $M^s$ is the push-forward measure of $(Y_i,Z_i)$ under the linear function.}. $M^s$ is determined by two parameters $\beta^s_0,\beta^s_1$. Assumption $\tilde{A}$ is the classical zero conditional mean restriction: $ \tilde{A}=\{s: E_{G^s}[\eta_i|Z_i]=0,\,\, (\beta^s_0,\beta^s_1)\in \mathbb{R}^2\}.$ We can show that the hypothesis $H=\{s\in \tilde{A}: \beta^s_1\ge 0\}$ is strongly binary decidable. Indeed, we have $\cup_{s^*\in H } M^{s^*}(G^{s^*})=\{F: Cov_F(Y_i,Z_i)\ge 0\}$ and $\cup_{s^*\in {H^c\cap\tilde{A}} } M^{s^*}(G^{s^*})=\{F: Cov_F(Y_i,Z_i)< 0\}$. These two sets do not intersect, so conditions in Definition (ref) can be verified for all $F$. We will see another strongly binary decidable hypothesis in an interval data example later.

The following lemma provides an equivalent condition to check whether $H$ is strongly binary decidable. The lemma below uses the definition of non-refutability set (Definition (ref)) and confirmation set (Definition (ref)) under $\tilde{A}$ instead of $\mathcal{S}$.

lemA hypothesis $H$ is strongly binary decidable under $\tilde{A}$ if and only if $\mathcal{H}_{\tilde{A}}^{scon}(H)=\mathcal{H}_{\tilde{A}}^{wnf}(H)$.

Intuitively, Lemma (ref) says that if we can confirm that the true structure $s$ is in $H$ for all distributions of observables, then we can refute $H^c\cap \tilde{A}$ for all distributions of observables.

Finite Sample Testing

Now I consider statistical testing of $H$ based on a finite sample. We want to test the null hypothesis that the true structure $s_0$, which generates the outcome distribution $F$, satisfies hypothesis $H$ against its complement in $\tilde{A}$: \[ \mathcal{H}_0: s_0\in H\quad \quad v.s. \quad \quad \mathcal{H}_1: s_0\in H^c\cap \tilde{A}. \] We have a finite sample of $i.i.d$ realizations from $F$ with empirical distribution $\mathbb{F}_n$ that converges weakly to $F$. A statistical test $T_n$ is a binary function that maps the empirical distribution and some random vector $\mathbf{\eta}$ to $\{0,1\}$: \[ T_n(\mathbb{F}_n,\mathbf{\eta})=

cases&1\quad \quad means we faile to reject \mathcal{H}_0,\\ &0\quad \quad means we reject \mathcal{H}_0.

\]

definitionWe say a test statistic $T_n$ achieves pointwise structural size control at level $\alpha$ if: \begin{equation} \inf_{F\in\cup_{s\in H} M^s(G^s)}{\lim\inf}_{n\rightarrow \infty} Pr(T_n(\mathbb{F}_n,\mathbf{\eta})=1)\ge 1-\alpha, \end{equation} and achieves structural test consistency if: \begin{equation} \inf_{F\in\cup_{s\in [H^c\cap \tilde{A}]} M^s(G^s)} {\lim \sup}_{n\rightarrow \infty } Pr(T_n(\mathbb{F}_n,\mathbf{\eta})=0) =1. \end{equation}

The names `structural size' and `structural test consistency' come from the fact that we construct the criteria ((ref)), ((ref)) through a partition of the assumption $\tilde{A}=H\cup[H^c\cap\tilde{A}]$ rather than a partition of the observation space $\mathcal{F}$. Structural size and structural power are what we care about since we aim to make a statement on the true structural parameter value. In particular, we may want to make binary decision on counterfactual outcomes. As we discuss after Definition (ref), a counterfactual analysis can be written as a parameter of interest.

The following proposition shows strongly that binary decidability is closely related to structural size control and structural test consistency.

propIf $H$ is not strongly binary decidable under $\tilde{A}$, then no statistic can simultaneously achieve pointwise structural size control ((ref)) for $\alpha<1$ and structural test consistency ((ref)).

The converse of this proposition also holds under further regularity conditions: if $H$ is strongly binary decidable under $\tilde{A}$, we can always find a test statistic that achieves pointwise size control and test consistency. Let \[

split[split omitted — 250 chars of source]

\] be the collection of all empirical distributions supported on a finite subset of $supp(X)$, and $Pr_{\mathbb{F}_n}(X_i=x)$ can be written as a fraction.

assumptionLet $F$ be the true distribution of outcomes and $\Theta_A^{ID}(F)$ is the identified set. Let $d_{\tilde{\mathcal{F}}}$ be a metric on $\tilde{\mathcal{F}}=\mathcal{F}\cup\mathcal{F}^d$. The following two conditions hold: \begin{enumerate} • $\Theta_{\tilde{A}}^{ID}(F)$ is upper hemicontinuous at $F$; • There exists a sequence of $a_n$ such that $\mathcal{C}_n= \sqrt{a_n}d_{\tilde{\mathcal{F}}}(\mathbb{F}_n,F)=O_p(1)$, and a sequence of constant $c_n$ such that $c_n\ge \mathcal{C}_n$ holds with probability converging to 1, and $c_n/\sqrt{a_n}\rightarrow 0$. \end{enumerate}
propLet Assumption (ref) hold. If $\Theta^0$ is a closed set, then there exists a test statistic $T_1(\mathbb{F}_n,\eta)$ that achieves pointwise structural size control ((ref)) for any $\alpha\ge 0$ and structural test consistency ((ref)) simultaneously.

Point Identified and Partially Identified Models

The following proposition shows that hypotheses about a point identified parameter of interest are always strongly binary decidable.

propLet $\tilde{A}$ be non-refutable. If $\theta$ is point identified under $\tilde{A}$, i.e. $\Theta_{\tilde{A}}^{ID}(F)$ is a singleton for all $F\in \mathcal{F}$, then $H= \{s\in \tilde{A}: \theta(s)\in \Theta_0\}$ is strongly binary decidable for any parameter value set $\Theta_0\subseteq \Theta$. Conversely, suppose $\theta$ is partially identified under $\tilde{A}$, and there exist $F$ and $F'$ such that $\Theta_{\tilde{A}}^{ID}(F)\cap \Theta_{\tilde{A}}^{ID}(F')\ne \emptyset$, $\Theta_{\tilde{A}}^{ID}(F)\ne \Theta_{\tilde{A}}^{ID}(F')$. Then there exists a parameter value set $\Theta_0$ such that $H= \{s\in \tilde{A}: \theta(s)\in \Theta_0\}$ is not strongly binary decidable.

Proposition (ref) shows that the traditional hypothesis testing approach works in a point identified model, regardless of the hypotheses on the parameter of interest. However, for a partially identified model, the formulation of a hypothesis is crucial. Let's consider the following policy decision rule: `we implement a policy $P$ if and only if the true structure is in $H$. When $H$ is not strongly binary decidable, we have size and power issue for any test statistic $T_n(\mathbb{F},\eta)$. If we decide to implement $P$ if and only if $T_n(\mathbb{F},\eta)=1$, we also know that the testing procedure cannot reject structures in $H^c$ consistently. If the policy $P$ is harmful when the true structure $s$ does not satisfy the parameter constraint of $H$, and we implement $P$ when $T_n(\mathbb{F},\eta)=1$, the policy $P$ can be harmful to the economy.

The problem does not arise from the sampling error but arises from the intrinsic inability to distinguish $H$ and $H^c$ by the distribution of observables. If we want to use a decision rule based on a hypothesis $\tilde{H}$ such that `we implement the policy $P$ if and only if the true structure is in $\tilde{H}$, the hypothesis $\tilde{H}$ must be strongly binary decidable. \footnote{An alternative approach is to formulate the hypothesis testing problem as a statistical decision problem, see Section 2.3 in manski2019econometrics for discussion.}

Extended and Subset Hypotheses

The next question is whether we can find a strongly binary decidable extended set or subset.

definitionA strongly binary decidable extension $H^{ext}$ is a strongly binary decidable set such that $H\subseteq H^{ext}\subseteq \tilde{A}$. A strongly binary decidable subset $H^{sub}$ is a strongly binary decidable set such that $H^{sub}\subseteq H$.

If the benefit to correctly implement a policy $P$ when $H$ is true is large, and the cost of mistakenly implementing $P$ when the true structure is in $H^{ext}\backslash H$ is small, we may want to test $H^{ext}$. Conversely, if there is a huge cost when we implement $P$ if $H^c$ is true, we may want to test $H^{sub}$. In this case, we sacrifice the benefit when the true structure is in $H\backslash H^{sub}$ to avoid the risk of mistakenly implementing $P$. The following proposition provides the minimal (resp. maximal) strongly binary decidable extension (resp. subset set).

propIf $\mathcal{H}_{\tilde{A}}^{snf}(H)=\mathcal{H}_{\tilde{A}}^{wnf}(H)$, then $\mathcal{H}_{\tilde{A}}^{snf}(H)$ is the smallest strongly binary decidable extension. If $\mathcal{H}_{\tilde{A}}^{scon}(H)=\mathcal{H}_{\tilde{A}}^{wcon}(H)$, then $\mathcal{H}_{\tilde{A}}^{wcon}(H)$ is the largest strongly binary decidable subset set.

For complete structure universes, $\mathcal{H}_{\tilde{A}}^{snf}(H)=\mathcal{H}_{\tilde{A}}^{wnf}(H)$ and $\mathcal{H}_{\tilde{A}}^{scon}(H)=\mathcal{H}_{\tilde{A}}^{wcon}(H)$ hold automatically, so we can always find a non-trivial strongly binary decidable extension (subset set). In the following, I present an example with a complete structure universe.

example[label=exa: Interval Data] (Interval Data) Consider a classical missing data problem where $Y_i^*$ is the unobserved real random variable, bounded above and below by observed variables $Y_i^u$ and $Y_i^l$. In this case, we can consider two primitive random variables $\epsilon_i^u$ and $\epsilon_i^l$, such that $\epsilon_i^u$ is supported on $[0,\infty)$ and $\epsilon_i^l$ is supported on $(-\infty,0]$. Observed variables $Y_i^u$ and $Y_i^l$ are generated through: \begin{equation} Y_i^u=Y_i^*+\epsilon_i^u\quad \quad and\quad \quad Y_i^l=Y_i^*+\epsilon_i^l. \end{equation} A structure consists of a joint distribution $G^s$ of $(Y_i^*,\epsilon_i^u,\epsilon_i^l)$ that satisfies the support conditions, and the mapping ((ref)). The structure universe $\mathcal{S}$ contains all structures with a distribution of $(Y_i^*,\epsilon_i^u,\epsilon_i^l)$ and the mapping ((ref)). We impose no further assumption, so $\tilde{A}=\mathcal{S}$. Since the mapping outcome in ((ref)) is unique, the structure universe is complete. Our hypothesis set is $H=\{s: E_{G^s}(Y_i^*)\in[a,b]\}$ and the corresponding hypothesis testing problem is: \[ \mathcal{H}_0: \, E[Y_i^*]\in [a,b] \quad \quad v.s.\quad \quad \mathcal{H}_1: \, E[Y_i^*]\notin [a,b]. \] The non-refutability set associated with $H$ is \begin{equation} \mathcal{H}_{\mathcal{S}}^{snf}(H)=\mathcal{H}_{\mathcal{S}}^{wnf}(H)=\left\{s:\, \left[E_{G^s}(Y_i^*+\epsilon_i^l),E_{G^s}(Y_i^*+\epsilon_i^u)\right]\cap [a,b]\ne \varnothing\right\}. \end{equation} Indeed, for any $s$ that satisfies the intersection condition above, suppose without loss of generality that $a\in\left[E_{G^s}(Y_i^*+\epsilon_i^l),E_{G^s}(Y_i^*+\epsilon_i^u)\right]$. We can construct $\tilde{s}$ such that \[ \begin{split} \tilde{Y}_i^*=Y_i^*+\epsilon_i^u,\quad \tilde{\epsilon}_i^u=0, \quad \text{and}\quad \tilde{\epsilon}_i^l=\epsilon_i^l-\epsilon_i^u\quad a.s., \end{split} \] and $G^{\tilde{s}}$ is the distribution of $(\tilde{Y}_i^*,\tilde{\epsilon}_i^u,\tilde{\epsilon}_i^l)$. It is easy to see $E_{G^{\tilde{s}}}(\tilde{Y}_i^*)=a$ and $\tilde{\epsilon}_i^u\ge 0$, $\tilde{\epsilon}_i^l\le 0$ almost surely, so support conditions of $(\epsilon_i^u,\epsilon_i^l)$ are satisfied. This implies $\tilde{s}\in \mathcal{H}_{\mathcal{S}}^{snf}(H)$. Conversely, for any $s$ that fails the intersection condition ((ref)), for example $E_{G^s}(Y_i^*+\epsilon_i^u)<a$, then for any $\tilde{s}$ such that $M^{\tilde{s}}(G^{\tilde{s}})=M^{{s}}(G^{{s}})$, we have $E_{\tilde{s}}[Y_i^*]\le E_{\tilde{s}}[Y_i^u]=E_{G^{\tilde{s}}}(Y_i^*+\epsilon_i^u)<a$. As a result, $\tilde{s}\notin H$. The confirmation set can be derived through Proposition (ref): \[ \mathcal{H}_{\mathcal{S}}^{scon}(H)=\mathcal{H}_{\mathcal{S}}^{wcon}(H)=\left\{s:\, \left[E_{G^s}(Y_i^*+\epsilon_i^l),E_{G^s}(Y_i^*+\epsilon_i^u)\right]\subseteq [a,b]\right\}. \] If we want to test $\mathcal{H}_{\mathcal{S}}^{snf}(H)$, a natural statistic is \[ T_n^{nf}(\mathbb{F}_n)= \sqrt{n}\left[\left(\frac{1}{n}\sum_{i=1}^n Y_i^u-a\right)_-^2+\left(b-\frac{1}{n}\sum_{i=1}^n Y_i^l\right)_-^2\right], \] where $(x)_-=\min(0,x)$. If we want to test $\mathcal{H}_{\mathcal{S}}^{scon}(H)$, a natural statistic is \[ T_n^{con}(\mathbb{F}_n)= \sqrt{n}\left[\left(b-\frac{1}{n}\sum_{i=1}^n Y_i^u\right)_-^2+\left(\frac{1}{n}\sum_{i=1}^n Y_i^l-a\right)_-^2\right]. \]

In the example above, the non-refutability set and confirmation set associated with $H$ are easy to find, while in more complicated structural models, the non-refutable and confirmation sets can be hard to characterize. In a complete structure universe, if $ H$ is not a strongly binary decidable, and $H$ is refutable (resp. confirmable), we want to instead test $s_0\in \mathcal{H}_{\tilde{A}}^{snf}(H)$ (resp. $s_0\in \mathcal{H}_{\tilde{A}}^{wcon}(H)$), which is strongly binary decidable. The following proposition shows that in a complete structure universe, testing $s_0\in \mathcal{H}_{\tilde{A}}^{snf}(H)$ can be equivalently written as a test of the existence of a structure that rationalizes data.

prop(Equivalent Decision) Let $s^0$ be the true structure that generates $F$. If $\mathcal{H}_{\tilde{A}}^{wnf}(H)= \mathcal{H}_{\tilde{A}}^{snf}(H)$, then the following two conditions are equivalent : \begin{enumerate} • $\exists s\in H$ such that $F\in M^s(G^s)$. • The true structure $s_0\in \mathcal{H}_{\tilde{A}}^{snf}(H)$. \end{enumerate} If $\mathcal{H}_{\tilde{A}}^{wcon}(H)= \mathcal{H}_{\tilde{A}}^{scon}(H)$, then the following two conditions are equivalent : \begin{enumerate} • $\{s\in \tilde{A}: \,\,F\in M^s(G^s)\}\subseteq H$. • The true structure $s_0\in \mathcal{H}_{\tilde{A}}^{scon}(H)$. \end{enumerate}
example[continues=exa: Interval Data] In the interval data example above, if there exists a structure $s$ such that $F\in M^s(G^s)$, and $E_{G^s}(Y_i^*)\in [a,b]$, the structure $s$ implies \[ E_{F}(Y_i^u)\ge E_{G^s}(Y_i^*)\ge a\quad \quad and \quad \quad E_{F}(Y_i^l)\le E_{G^s}(Y_i^*)\le b. \] A possible test statistic to test this implication is to use $T_n^{nf}(\mathbb{F}_n)$ defined above. On the other hand, if any struture $s$ that can generate $F$ is contained in $H$, the following two extreme cases: \[ \begin{split} Y_i^*= Y_i^u \quad \epsilon_i^u\equiv 0 \quad \epsilon_i^l=Y_i^l-Y_i^u,\\ Y_i^*= Y_i^l \quad \epsilon_i^l\equiv 0 \quad \epsilon_i^u=Y_i^u-Y_i^l, \end{split} \] must also be included in $H$, which means $E_F[Y_i^u]\le b$ and $E_F[Y_i^l]\ge a $ must hold. A possible statistic to test this implication is to use $T_n^{con}(\mathbb{F}_n)$ defined above.

Application to Treatment Effects

In this section, I apply the method to IA1994 with a binary treatment and a binary instrument. The observed outcome variable $Y_i$ and treatment decision $D_i$ are generated through

align[align omitted — 185 chars of source]

where $D_i(1),D_i(0)$ are potential treatment decisions, $Y_i(d,z)$ are the potential outcome and $Z_i$ is a binary instrument.

Primitive variables are $\epsilon_i=(D_i(1),D_i(0),Y_i(0,0),Y_i(1,0),Y_i(0,1),Y_i(1,1),Z_i)$ and observed variables are $X_i=(Y_i,D_i,Z_i)$. Let $\mathcal{Y}$ be the space of $Y_i$ and let $\mathcal{B}$ be a Borel-sigma algebra on $\mathcal{Y}$. The observation space is

equation[equation omitted — 103 chars of source]

and space of potential distribution

equation[equation omitted — 209 chars of source]

All structures agrees on the functional relation between $X_i$ and $\epsilon_i$ specified in ((ref)). The mapping\footnote{See Definition (ref)} $M^s$ is defined as:

equation[equation omitted — 227 chars of source]

$M^s$ contains exactly one predicted distribution of observables and all structures are complete. The structure universe $\mathcal{S}$ is:

equation[equation omitted — 170 chars of source]

Following kitagawa2015, I define the following two quantities for all $B\in \mathcal{B}$ and $d\in \{0,1\}$:

equation[equation omitted — 187 chars of source]

The Imbens-Angrist Monotonicity assumption (IA-M) assumes exogeneity, exclusion and monotonicity of the instrument $Z_i$:

equation[equation omitted — 304 chars of source]

kitagawa2015 derives the sharp testable implications of the IA-M assumption ((ref)). I reformulate the result in the language of non-refutability sets in the following lemma.

lemLet $P(\cdot,d)$ and $Q(\cdot,d)$, $d\in\{0,1\}$, be absolutely continuous with respect to some measure $\mu_F$.\footnote{Such dominating measure always exists, for example define $\mu_F(B)=P(B,1)+Q(B,1)+P(B,0)+Q(B,0)$ for all $B\in \mathcal{B}$.} The non-refutability set associated with IA-M assumption $\mathcal{H}_\mathcal{S}^{snf}(A)$ is the collection of structures $s$ such that if $F\in M^s(G^s)$, then for all Borel set $B$: \begin{equation} \begin{split} P(B,1)&\ge Q(B,1),\\ Q(B,0)&\ge P(B,0). \end{split} \end{equation}

kitagawa2015 proposes using the core determining class galichon2011set such as the class of closed intervals to test ((ref)). Alternatively, ((ref)) can be equivalently formulated using Radon-Nikodym densities. We will see the advantage of densities when we construct extensions.

thmLet $\mu_F$ be the common dominating measure in Lemma (ref). Let $p(y,d)$ and $q(y,d)$ be defined as: \begin{equation} p(y,d)=\frac{dP(B,d)}{d\mu_F}\quad\quad q(y,d)=\frac{dQ(B,d)}{d\mu_F}. \end{equation} Then the testable implication ((ref)) holds if and only if \begin{equation} \begin{split} p(y,1)-q(y,1)\ge 0 \quad \quad &\mu_F-a.s.\,,\\ q(y,0)-p(y,0)\ge 0 \quad \quad &\mu_F-a.s.\, . \end{split} \end{equation}

Our main parameter of interest $\theta$ is the local average treatment effect for compliers:

equation[equation omitted — 93 chars of source]

Under the IA-M assumption $A$, the identified set for LATE is characterized by

equation[equation omitted — 256 chars of source]

As shown in Lemma (ref), the IA-M assumption $A$ is refutable. In most empirical applications, researchers do not test this implication, neither do they specify what should be done when the testable implication is rejected. In the next section, I use the relaxed assumption approach to find relaxed assumptions $\tilde{A}$ such that $\tilde{A}$ is non-refutable, find the identified set under $\tilde{A}$, and discuss the estimation and inference on LATE under $\tilde{A}$.

Extensions of the IA-M Assumption

In this section, I will first show that the IA-M assumption have an alternative representation. As discussed in Section (ref), different representations of an assumption can result in different extensions. The canonical representation ((ref)) and the alternative representation will be used to constructed different extensions. I will then look at the maximal relaxation in Definition (ref) and show that the identified set for LATE under the maximal relaxation does not satisfy Property (ref). Then I proceed to construct extensions using the minimal deviation method.

The following is an alternative representation of the IA-M assumption that will be used throughout this section.

lemThe IA-M assumption defined in ((ref)) can be equivalently written as the intersection: $A=A^{ER}\cap A^{TI}\cap A^{EM-NTAT}\cap A^{ND}$ where:\begin{enumerate} • $A^{ER}=\left\{s\big|Y_i(1,1)=Y_i(1,0) \quad and \quad Y_i(0,1)=Y_i(0,0) \right\}$ is the exclusion restriction; • $A^{TI}=\left\{s\big| Z_i\perp \left(Y_i(1,1),Y_i(0,1),Y_i(1,0),Y_i(0,0)\right)|D_i(1),D_i(0) \right\}$ is the type independent instrument assumption; • Assumption $A^{EM-NTAT}$ is the set of structures $s$ such that the measures of always takers and never takers are independent of $Z_i$, i.e. \[ \begin{split} E_{G^s}[\mathbbm{1}(D_i(1)=D_i(0)=1)|Z_i=1]&=E_{G^s}[\mathbbm{1}(D_i(1)=D_i(0)=1)|Z_i=0],\\ E_{G^s}[\mathbbm{1}(D_i(1)=D_i(0)=0)|Z_i=1]&=E_{G^s}[\mathbbm{1}(D_i(1)=D_i(0)=0)|Z_i=0]. \end{split}\]$A^{ND}=\left\{ s\big| G^s\,\,satisfies: D_i(1)\ge D_i(0)\right\}$ is the no defiers assumption. \end{enumerate}

The Maximal Extension with Exclusion Restriction and `No Defiers'

To fix the idea, let's consider the case that $A^{ER}$ and $A^{ND}$ holds but we relax the independent instrument assumption. Moreover we consider the extension set $\tilde{A}^{max}\equiv \left(A\cup[{\mathcal{H}_{\mathcal{S}}^{snf}(A)}]^c\right) \cap A^{ER}\cap A^{ND}$. $\tilde{A}^{max}$ is the maximal strong extension defined in Proposition (ref) intersected with the exclusion restriction and the `No Defiers' assumption.

prop$\tilde{A}^{max}\equiv \left(A\cup[{\mathcal{H}_{\mathcal{S}}^{snf}(A)}]^c\right) \cap A^{ER}\cap A^{ND}$ is a strong extension. The closure of the identified set for LATE under $\tilde{A}^{max}$ is \[ \overline{{LATE}^{ID}_{\tilde{A}^{max}}(F)}=\begin{cases} \frac{E[Y_i|Z_i=1]-E[Y_i|Z_i=0]}{E[D_i|Z_i=1]-E[D_i|Z_i=0]}\quad &\text{if (\ref{eq: testable implication in density form}) holds for } \,\,F,\\ \left[ \underline{\mathcal{Y}}_{P(B,1)}-\bar{\mathcal{Y}}_{Q(B,0)},\bar{\mathcal{Y}}_{P(B,1)}-\underline{\mathcal{Y}}_{Q(B,0)}\right] &\text{otherwise}, \end{cases} \] where for $V\in\{P,Q\}$, $\underline{\mathcal{Y}}_{V(B,0)}$ is the lower bound of the support of $Y_i$ under measure $V(B,0)$, and $\bar{\mathcal{Y}}_{V(B,1)}$ is the upper bound of the support of $Y_i$ under measure $V(B,1)$.

The $\tilde{A}^{max}$ above allows arbitrary dependence of the instrument on the potential outcomes whenever the testable implication ((ref)) fails. First, we should note that the identified set for LATE is very unstable when $F$ satisfies $p(y,1)-q(y,1)=0$ for some $y\in \mathcal{Y}$. Whenever we perturb $F$ slightly such that $p(y,1)-q(y,1)<0$ and ((ref)) fails, the identified set for LATE explodes. Second, the identification set ${LATE}^{ID}_{\tilde{A}^{max}}(F)$ is not any better than the $LATE^{ID}_A(F)$ in equation ((ref)). In terms of interpretation, an uninformative identified set\footnote{Note that the identification result in Proposition (ref) contains only support information when ((ref)) fails.} for LATE is not different from an empty identified set. This is because whenever $F$ fails ((ref)), we give up the `Independent Instrument' assumption. Therefore the remaining assumptions $A^{ER}$ and $A^{ND}$ cannot generate any restrictions on the parameter of interest. As a result, I focus on deriving extensions using minimal deviation method in Definition (ref). I will relax the `No Defiers' and the independent instrument assumption in the following. An extension that relaxes the independent instrument assumption and an extension that relaxes the exclusion restriction are given Appendix (ref).

The Minimal Defiers Extension

Recall that $\mathcal{S}$ is a complete structure universe. I consider a strong extension that use measure of defiers as deviation from the no defiers assumption. The extension relaxes the independent instrument to a type independent instrument assumption. I first define the measure of defiers in $G^s$ as

equation[equation omitted — 106 chars of source]
assumptionLet $m^{min}(F)\equiv \inf\{m^d(s): F\in M^s(G^s) \,\, and \,\, s\in A^{ER}\cap A^{TI}\cap A^{EM-NTAT}\}$ be the minimal defier amount under $F$. We call \[\tilde{A}=\cup_{F\in \mathcal{F}} \left\{s\in A^{ER}\cap A^{TI}\cap A^{EM-NTAT}: m^d(s)=m^{min}(F),\,\,F\in M^s(G^s)\right\}\] the minimal defiers extension with type independent instrument.

In the above extension, I also relax the independent instrument condition. The type independence assumption is also used in other empirical contexts to study LATE (e.g. see kedagni2019). This is because, by kitagawa2009identification, exclusion restrictions (ER) and instrument condition (IV) has testable implication, so any non-refutable relaxation should relax either ER or IV condition. The second condition in Assumption (ref) requires the measure of always takers (AT) and never takers (NT) to be independent of the instrument.

propThe extension $\tilde{A}$ defined in Assumption (ref) is a strong extension of $A$.
proofSuffice to check conditions in Proposition (ref), see Appendix (ref).
remarkTo emphasize that extensions $\tilde{A}$ constructed by minimal deviation method may not be strong extensions, I provide two examples of extensions in Appendix (ref) that are LATE-consistent extension but are not strong extensions.

Identified Set under Different Extensions

This section describes the identified set for LATE under $\tilde{A}$ in Assumption (ref). Let $ \mathcal{Y}_d\equiv\{y\in \mathcal{Y}: (-1)^{d}[q(y,d)-p(y,d)]\ge 0\} $ be the collection of $y\in \mathcal{Y}$ such that the density differences are positive.

assumptionThere exists a constant $c\ge 0$ such that: (i) $Pr_F(Z_i=1)\in (c,1-c)$; (ii) $Q(\mathcal{Y}_0,0)-P(\mathcal{Y}_0,0)>c$ and $P(\mathcal{Y}_1,1)-Q(\mathcal{Y}_1,1)>c$.

Assumption (ref) is a regularity assumption. For the identification result, we only need it to hold with $c=0$ so that LATE is well-defined. For inference purpose, I require $c>0$ to avoid weak instrument issue.

propLet the extension $\tilde{A}$ satisfies Assumptions (ref). If Assumption (ref) holds for $c=0$, then the identified $ {LATE}^{ID}_{\tilde{A}}$ satisfies \begin{equation} {LATE}_{\tilde{A}}^{ID}(F)=\frac{\int_{\mathcal{Y}_1}{y (p(y,1)-q(y,1))}d\mu_F(y)}{P(\mathcal{Y}_1,1)-Q(\mathcal{Y}_1,1)} -\frac{\int_{\mathcal{Y}_0}{y (q(y,0)-p(y,0))}d\mu_F(y)}{Q(\mathcal{Y}_0,0)-P(\mathcal{Y}_0,0)}, \end{equation} where $\mathcal{Y}_1=\{y\in\mathcal{Y}: p(y,1)-q(y,1)\ge 0\}$ and $\mathcal{Y}_0=\{y\in\mathcal{Y}: q(y,0)-p(y,0)\ge 0\}$. Note that LATE is point identified.

In Assumptions (ref), I require that potential outcomes are independent of the instrument conditional on compliers, i.e. $\{Y_i(d,z)\}_{d,z\in\{0,1\}}\perp Z_i\big| (D_i(1)=1,D_i(0)=0)$. As a result, the identified probability of compliers conditioned on $Z_i=1$ is $P(\mathcal{Y}_1,1)-Q(\mathcal{Y}_1,1)$ and the identified probability of compliers conditioned on $Z_i=0$ is $Q(\mathcal{Y}_0,0)-P(\mathcal{Y}_0,0)$.

As we discussed in Section 2, we want the identified set for LATE under the chosen extensions to be a continuous correspondence with respect to the distribution of observables $F$. If the identified set for LATE is discontinuous under the extension, we may get spuriously informative bound for LATE due to sampling error. I compare the identified set for LATE under the maximal extension in Proposition (ref) and the identified set for LATE under the minimal defiers extension in Proposition (ref) in terms of Property (ref). I equip $\mathcal{F}$ with the Sobolev norm: $||F||_{1,\infty}\equiv \max_{i=0,1} ||F^{(i)}||_{\infty}$, where $F^{(i)}$ is the $i$-th Radon-Nikodym density of $F$ with respect to $\mu_F$.

propLet $\tilde{A}_1$ be the maximal extension and ${LATE}^{ID}_{\tilde{A}_1}(F)$ be the corresponding identified set defined in Proposition (ref). Let $\tilde{A}_2$ be the minimal defiers extension defined in Assumption (ref) and ${LATE}^{ID}_{\tilde{A}_2}(F)$ be the corresponding identified set defined in ((ref)). Suppose $\forall F\in \mathcal{F}$, the support of $Y_i$ is bounded above by $M^u_s$ and bounded below by $M^l_s$, then ${LATE}^{ID}_{\tilde{A}_1}(F)$ is not upper hemicontinuous with respect to the Sobolev norm $||\cdot ||_{1,\infty}$, and ${LATE}^{ID}_{\tilde{A}_2}(F)$ is continuous with respect to $||\cdot||_{1,\infty}$.

Proposition (ref) shows that the maximal extension is not a good choice if the parameter of interest is LATE.

Estimation and Inference

The identification result in Proposition (ref) relies on the sets $\mathcal{Y}_1,\mathcal{Y}_0$. Throughout this section, I focus on the estimation and inference problem when $Y_i$ is continuously distributed on, and $\mu_F$ is the Lebesgue measure.

assumption$Y_i$ is continuously distributed with unbounded support and the measure $P(B,d)$, $Q(B,d)$ is absolutely continuous with respect to the Lebesgue measure.

To estimate $\hat{\mathcal{Y}}_0,\hat{\mathcal{Y}}_1$, I estimate the density $p(y,d)$ and $q(y,d)$ using kernel density estimators:

equation[equation omitted — 582 chars of source]
assumptionThere exist constants $M_l$ and $M_u$ such that for $d=0,1$ such that $\mathcal{Y}_d\cap [M_u,\infty) \in \{\varnothing, [M_u,\infty)\}$ and $\mathcal{Y}_d\cap (-\infty,M_l] \in \{\varnothing, (-\infty,M_l]\}$. Moreover, we know $\mathcal{Y}_d\cap [M_u,\infty)$ and $\mathcal{Y}_d\cap (-\infty,M_l]$.

The assumption above assumes that the sign of $p(y,d)-q(y,d)$ is known and fixed in the large value of $y$. As a result, we only need to estimate the set $\mathcal{Y}_d\cap [M_l,M_u]$.

example[Gaussian tail dominance.] Suppose $p(y,d)$ and $q(y,d)$ have Gaussian tails: $p(y,d)=C_pe^{-y^2/\sigma_p(d)^2}$ and $q(y,d)=C_qe^{-y^2/\sigma_q(d)^2}$ for $|y|>C^{tail}>0$. If $\sigma_p(1)>\sigma_q(1)$, then $\mathcal{Y}_1\cap [C^{tail},\infty)=[C^{tail},\infty)$ and $\mathcal{Y}_1\cap (-\infty,-C^{tail}]=(-\infty,-C^{tail}]$.

Define the upper tail set $\mathcal{Y}_d^{ut}= \mathcal{Y}_d\cap [M_u,\infty)$ and the lower tail set $\mathcal{Y}_d^{lt}= \mathcal{Y}_d\cap (-\infty,M_l]$ and we estimate \[ \hat{\mathcal{Y}}_d(b_n)= \{y\in (M_l,M_u): f_h(y,d)\ge b_n\}\cup \mathcal{Y}_d^{ut}\cup \mathcal{Y}_d^{lt}, \] where $b_n$ is a sequence of positive constants that converges to zero. The estimated set above only uses density $f_h(y,d)$ to distinguish whether $y\in \mathcal{Y}_d$ in the range $(M_l,M_u)$, and uses the known tail sign in Assumption (ref) directly. When the relaxed assumption is defined in Assumption (ref), I construct an estimator of $LATE^{ID}_{\tilde{A}}(F)$ ((ref)) as:

equation[equation omitted — 1,068 chars of source]

Limit Distribution of $\widehat{LATE}$

I present the limit distribution of $\widehat{LATE}$ defined in ((ref)). The following assumptions are sufficient to guarantee $\widehat{LATE}$ in ((ref)) will converge to a normal distribution.

assumptionThe kernel function $K$ satisfies: (i) $K(u)$ is continuous and supported on $[-A,A]$ and $\int_u K(u) du=1$; (ii) $\int_u uK(u)du=0$; (iii) $\int u^2K(u)du<\infty$.
assumptionThe conditional distribution $F(y|D_i=k,Z_i=l)$ has a density $f(y|k,l)$ for all $k,l\in\{0,1\}$, and $f''(y|k,l)$ exists and is uniformly bounded by a constant $c_f$; (iii) $E(Y_i^{2+\delta})<\infty$ for some $\delta>0$.

The above two assumptions are standard in literature and guarantee the density difference estimator $f_h(y,d)$ will converges uniformly in probability to its limit $(-1)^{1-d}(p(y,d)-q(y,d))$ at polynomial rate.

assumptionLet $f(y,1)=p(y,1)-q(y,1)$ and $f(y,0)=q(y,0)-p(y,0)$. The following condition holds for any sequence $b_n\rightarrow 0_+$: $\int_{M_l}^{M_u} |f(y,d)| \mathbbm{1}(-b_n\le f(y,d)\le b_n)dy =O(b_n^2).$

Assumption (ref) controls the bias from trimming $\{y\in[M_l,M_u]:0<f_h(y,d)<b_n\}$. Essentially, we rule out all outcome distributions such that $\{y:f(y,d)=0\}$ has a positive measure. This assumption is imposed to remove the bias from sampling error in kernel estimator $f_h(y,d)$. Assumption (ref) can be replaced by a sufficient primitive condition.

assumptionLet $M_0<\infty$ be a positive integer. For $d=0,1$, the set $\mathcal{C}_d=\{y: f(y,d)=0,\,\,y\in[M_l,M_u]\}$ has at most $M_0$ points. Let $B(\mathcal{C}_d,\delta)=\cup_{y\in\mathcal{C}_d} B(y,\delta)$ be the $\delta$-neighborhood of $\mathcal{C}_d$ for $d=0,1$. For both $d=0,1$, we have $\sup_{y\in B(\mathcal{C}_d,\delta)}|\frac{d(f(y,d))}{dy}|>1/C$ for some $C,\delta>0$.
lemAssumption (ref) implies Assumption (ref).
proofBy the bounded density condition, the Lebesgue measure of set $\{y:\mathbbm{1}(-b_n\le f(y,d)\le b_n)\}$ is less than $CM_0b_n$. Therefore \[\int_{M_l}^{M_u} |f(y,d)| \mathbbm{1}(-b_n\le f(y,d)\le b_n)dy\le CM_0b_n^2.\] So Assumption (ref) implies Assumption (ref).
thmLet $\widehat{LATE}$ be defined in ((ref)) and ${LATE}^{ID}_{\tilde{A}}(F)$ be defined in ((ref)). Suppose Assumption (ref) holds for $c>0$ and Assumption (ref) -(ref) hold. Let $b_n\asymp n^{-1/4}/\log n$ and $h_n\asymp n^{-1/5}$, then $\sqrt{n}(\widehat{LATE}-{LATE}_{\tilde{A}}^{ID}(F))\rightarrow_d N(0,\Pi'\Gamma\Sigma\Gamma'\Pi)$, where \begin{equation*} \Sigma=Var\begin{pmatrix} \mathbbm{1}(Z_i=0)\\ \mathbbm{1}(Z_i=1)\\ Y_i \mathbbm{1}(D_i=1,Z_i=1)\mathbbm{1}(Y_i\in \mathcal{Y}_1)\\ Y_i \mathbbm{1}(D_i=1,Z_i=0)\mathbbm{1}(Y_i\in \mathcal{Y}_1)\\ Y_i \mathbbm{1}(D_i=0,Z_i=0)\mathbbm{1}(Y_i\in \mathcal{Y}_0)\\ Y_i \mathbbm{1}(D_i=0,Z_i=1)\mathbbm{1}(Y_i\in \mathcal{Y}_0)\\ \mathbbm{1}(D_i=1,Z_i=1)\mathbbm{1}(Y_i\in \mathcal{Y}_1)\\ \mathbbm{1}(D_i=1,Z_i=0)\mathbbm{1}(Y_i\in \mathcal{Y}_1)\\ \mathbbm{1}(D_i=0,Z_i=0)\mathbbm{1}(Y_i\in \mathcal{Y}_0)\\ \mathbbm{1}(D_i=0,Z_i=1)\mathbbm{1}(Y_i\in \mathcal{Y}_0)\end{pmatrix}, \end{equation*} and matrices $\Pi$ and $\Gamma$ are specified as by { \[ \Pi=\begin{pmatrix} \frac{1}{\pi_3}\\ -\frac{1}{\pi_4}\\ -\frac{\pi_1}{\pi_3^2}\\ \frac{\pi_2}{\pi_4^2} \end{pmatrix},\quad\pi\equiv\begin{pmatrix} \pi_1\\ \pi_2\\ \pi_3\\ \pi_4 \end{pmatrix}= \begin{pmatrix} \int_{\mathcal{Y}_1} y(p(y,1)-q(y,1))dy\\ \int_{\mathcal{Y}_0} y(q(y,0)-p(y,0))dy\\ \int_{\mathcal{Y}_1} (p(y,1)-q(y,1))dy\\ \int_{\mathcal{Y}_0} (q(y,0)-p(y,0))dy \end{pmatrix}, \quad \Gamma=\frac{1}{Pr(Z_i=1)Pr(Z_i=0)}\begin{pmatrix} \Gamma_1 &\Gamma_3 &\bm{0}_{2\times 4}\\ \Gamma_2 &\bm{0}_{2\times 4} &\Gamma_3 \end{pmatrix}, \]where \begin{equation*} \Gamma_1= \begin{pmatrix} E[Y_i\mathbbm{1}(D_i=1,Z_i=1)\mathbbm{1}(Y_i\in \mathcal{Y}_1)] & -E[Y_i\mathbbm{1}(D_i=1,Z_i=0)\mathbbm{1}(Y_i\in \mathcal{Y}_1)]\\ -E[Y_i\mathbbm{1}(D_i=0,Z_i=1)\mathbbm{1}(Y_i\in \mathcal{Y}_0)] & E[Y_i\mathbbm{1}(D_i=0,Z_i=0)\mathbbm{1}(Y_i\in \mathcal{Y}_0)] \end{pmatrix}, \end{equation*} \begin{equation*} \Gamma_2=\begin{pmatrix} E[\mathbbm{1}(D_i=1,Z_i=1)\mathbbm{1}(Y_i\in \mathcal{Y}_1)] & -E[\mathbbm{1}(D_i=1,Z_i=0)\mathbbm{1}(Y_i\in \mathcal{Y}_1)]\\ -E[\mathbbm{1}(D_i=0,Z_i=1)\mathbbm{1}(Y_i\in \mathcal{Y}_0)] & E[\mathbbm{1}(D_i=0,Z_i=0)\mathbbm{1}(Y_i\in \mathcal{Y}_0)] \end{pmatrix}, \end{equation*} \begin{equation*} \Gamma_3=\begin{pmatrix} Pr(Z_i=0) & -Pr(Z_i=1) &0 &0\\ 0 &0 &Pr(Z_i=1) & -Pr(Z_i=0) \end{pmatrix}. \end{equation*} }
CorollaryLet $(\hat{\Gamma},\hat{\Pi},\hat{\Sigma})\rightarrow_p ({\Gamma},{\Pi},{\Sigma})$, and let $\hat{\sigma}=\sqrt{\hat{\Pi}'\hat{\Gamma}\hat{\Sigma}\hat{\Gamma}'\hat{\Pi}}$. Then the set \begin{equation} \left[\widehat{LATE}-\frac{\hat{\sigma}}{\sqrt{n}}\Phi(\frac{\alpha}{2}),\widehat{LATE}+\frac{\hat{\sigma}}{\sqrt{n}}\Phi(1-\frac{\alpha}{2})\right] \end{equation} is a valid $\alpha$-confidence interval for ${LATE}_{\tilde{A}}^{ID}(F)$, where $\Phi$ is the normal CDF function.

Theorem (ref) shows that the LATE estimator in ((ref)) is $\sqrt{n}$ consistent. Once the matrices $\Pi$, $\Gamma$ and $\Sigma$ are estimated by consistent estimators, we can test hypothesis such as $H_0:{LATE}_{\tilde{A}}^{ID}(F)=0$. Since LATE is point identified, conventional hypothesis testing method can achieve structural size control and test consistency simultaneously. However, Assumption (ref) requires the econometrician to know the sign of tail behavior of $p(y,1)-q(y,1)$ and $q(y,0)-p(y,0)$. In some empirical application, we may want to be agnostic about tail signs or only impose less restrictive conditions on tail signs. In this case, we can calculate the confidence interval for each possible tail condition, and then take the union, but this confidence interval will be conservative.

Simulation

This section illustrates the finite sample performance of the proposed inference method. I consider two simulation settings. In the first setting, the IA-M assumption is violated, and the goal of the simulation is to see how the inference method works under the known and unknown tail signs. In the second setting, the IA-M assumption is not violated, and the goal of the simulation is to compare the numerical difference of the 2SLS estimator of ((ref)) and estimator ((ref)).

Simulation Setting I

Instead of simulating the primitive variable $Y_i(d,z),D_i(z)$, I directly simulate the distribution of observed variable such that $Pr(Z_i=1)=0.6$ and \[

split[split omitted — 138 chars of source]

\] In this simulation, Assumptions (ref), (ref) and (ref) are satisfied. The trimming band is $[M_l,M_u]=[-2.5,7]$ and Assumption (ref) is satisfied since $\mathcal{Y}^{ut}_1=\mathcal{Y}^{lt}_1=\varnothing$, $\mathcal{Y}^{ut}_0=[M_u,\infty)$, and $\mathcal{Y}^{lt}_0=(-\infty,M_l]$. Assumption (ref) is satisfied since $p(y,1)-q(y,1)=0$ and $q(y,0)-p(y,1)=0$ have two solutions in interval $[M_l,M_u]$ and the derivatives are bounded away from zero (see Figure (ref)).

figure[figure omitted — 221 chars of source]

The true identified value of ${LATE}^{ID}_{\tilde{A}}(F)$ is $1.7385$. Simulation results are given in Table (ref). Coverage probability are calculated from $1000$ replications, and I compare the coverage probability under different sample size $n$ in each replication and the choice of trimming constant $b$. The row of `known tail' in Table (ref) corresponds to the constraints $\mathcal{Y}^{ut}_1=\mathcal{Y}^{lt}_1=\varnothing$, $\mathcal{Y}^{ut}_0=[M_u,\infty)$, and $\mathcal{Y}^{lt}_0=(-\infty,M_l]$ as in the simulation design. The row of `Conservative' in Table (ref) corresponds to the case that the tail set is unknown, and I take the union of confidence intervals under all $16$ possible tail conditions.

table[table omitted — 347 chars of source]

The simulation result shows that if we can correctly impose the tail condition as in the known tail case, inference on the true LATE value based on ((ref)) is asymptotically exact but can be sensitive to the choice of trimming sequence $b_n$. If we want to be agnostic about the true tail condition, the union method is conservative.

Simulation Setting II

When the IA-M assumption holds, by vytlacil2002independence, the potential outcome model is equivalent to the latent index model. In this simulation setting, let $U_i\sim U[0,1]$ and $D_i=\mathbbm{1}(0.2+0.6Z_i>U_i)$, where $U[0,1]$ is a uniform distribution on interval $[0,1]$, and $Z_i\sim \mathtt{Bernoulli}(0.5)$ is independent of $U_i$. The potential outcome $(Y_i(1),Y_i(0))\sim N(\mu,\Sigma)$, where $\mu=[2,1.5]$ and $Var(Y_i(1))=2$, $Var(Y_i(0))=1.5$, $Corr(Y_i(1),Y_i(0))=0.7$. We can show $LATE^{ID}(F)=0.5$.

Let $\widehat{LATE}^{Wald}$ denote the Wald estimator of $LATE^{ID}(F)$, and let $\widehat{LATE}^{New}$ denote the estimator in equation ((ref)). Table (ref) shows some summary statistics of the numerical difference between these two estimators over $m=1000$ simulation replications. The first row shows that when the sample size increase, the numerical difference of the two estimator converges to zero in the second moment. The second row in Table (ref) reports the efficiency loss of $\widehat{LATE}^{New}$. We see that at finite sample, not imposing IA-M assumption when it holds can lead to efficiency loss, but the efficiency loss decreases with sample size.

table[table omitted — 593 chars of source]

Empirical Illustration

In this section, I apply my results in Proposition (ref) and Theorem (ref) to card1993using, who studied the causal effect of college attendance on earnings. In this application, the outcome variable $Y_i$ is an individual $i$'s log wage in 1976, $D_i=1$ means individual $i$ attended a four-year college, and $Z_i=1$ means the individual was born near a four-year college. This data set has been used by both kitagawa2015 and mourifie2017testing to test the IA-M assumption, and they both reject the IA-M assumption. If a child grew up near a college, he or she may hear more stories of heavy tuition burden, which may discourage him or her from attending college. On the other hand, if this child grew up far away from a college, he or she may instead choose to attend college. Therefore, we would expect defiers to exist in this empirical setting. Moreover, it is unclear why this instrument is fully independent of the potential income, since the choice of residence may depend on parents' potential income, which may be correlated with their children's income.

I conditioned $(Y_i,D_i,Z_i)$ on three characteristics: living in the south (S/NS), living in a metropolitan area (M/NM), and ethnic group (B/NB). I follow mourifie2017testing in excluding subgroup NS/NM/B due to the small sample size, and also exclude subgroup NS/M/B due to the high frequency of $Z=1$. I conduct estimation and inference on each of the remaining 6 subgroups and the pooled sample. The choices of trimming sequence $b_n$, kernel bandwidth $h$, upper and lower band $M_u,M_l$, tail set $\mathcal{Y}_d^{ut},\mathcal{Y}_d^{lt}$ are specified in Appendix (ref).

Estimation results are reported in Table (ref). I also report the LATE estimates when we directly use the IA-M assumption and Wald statistics. The estimated measure of compliers under $\tilde{A}$ satisfying (ref) conditioned on $Z_i=1$ and $Z_i=0$ are reported as $P(\mathcal{Y}_1,1)-Q(\mathcal{Y}_1,1)$ and $Q(\mathcal{Y}_0,0)-P(\mathcal{Y}_0,0)$, while the estimated measure of compliers under IA-M assumption is $E[D_i|Z_i=1]-E[D_i|Z_i=0]$. The estimates of ${LATE}^{ID}_{\tilde{A}}$ and $LATE^{Wald}$ differ the most for three groups: S/NM/NB, S/M/NB and S/M/B. It should be noted that for all these three groups, estimated $E[D_i|Z_i=1]-E[D_i|Z_i=0]$ differs from $P(\mathcal{Y}_1,1)-Q(\mathcal{Y}_1,1)$ and $Q(\mathcal{Y}_0,0)-P(\mathcal{Y}_0,0)$. If we blindly use the identification result under the IA-M assumption and use the standard LATE Wald estimator, the `identified' local average treatment effect can be negative (subgroups S/NM/NB and S/M/NB), or be unrealistically large (subgroup S/M/B). Once we use a strong extension, the estimated LATE for each of the 6 subgroups is positive, and the value of LATE is all between zero and one. We fail to reject education decrease future earning for the complier group ($LATE<0$) for all 6 subgroups for the $LATE^{wald}$. On the other hand, my method can reject the hypothesis $LATE<0$ for the NS/NM/NB and the S/NM/B group at 95% confidence level. When I look at the African-American only, while $LATE^{wald}$ is large, the hypothesis fail to reject that education is harmful to their earning, while my method will reject the hypothesis $LATE^{wald}$ for the African-American is negative.

{

table[table omitted — 2,543 chars of source]

}

Application to Binary Outcome Sector Choice

In this section, I apply the method to a binary outcome sector choice model with a binary instrument mourifie2018roy. This model is incomplete. Observed outcome $Y_i\in \{0,1\}$ is binary and the observed job sector choice $D_i\in\{0,1\}$ is binary. A binary instrument $Z_i$ is observed. The $\epsilon$ variables include $Y_i(1)$ and $Y_i(0)$, which are the potential outcome in job sector 1 and 0 respectively, and the instrument variable $Z_i$. Observed sector outcome $Y_i$ is generated through:

equation[equation omitted — 83 chars of source]

Without imposing further assumptions, equation ((ref)) does not specify how sector choice $D_i$ is determined, and $D_i$ can take values in one of the three sets $\{0\},\{1\},\{0,1\}$. If we impose the classical Roy sector selection rule, for all $Z_i=z\in\{0,1\}$ we have:

equation[equation omitted — 216 chars of source]

The Roy sector selection rule ((ref)) is just a special case of how $D_i$ is determined. To specify the structure universe $\mathcal{S}$, we consider all possible sector selection rules. \footnote{For each $Y_i(0)=y_0,Y_i(1)=y_1,Z_i=z$, $D_i$ can take values in three sets $\{0\},\{1\},\{0,1\}$. Therefore, there are $3^8$ ways to specify the sector selection rule.}

definitionA sector selection rule is a set-valued function \[ \begin{split} D^{sel}: \{0,1\}^3&\rightarrow \{\{0\},\{1\},\{0,1\}\},\\ (y_1,y_0,z)&\rightarrow D^{sel}(y_1,y_0,z). \end{split} \] Let $\mathcal{D}^{sel}$ be the collection of all possible sector selection rules. Let $D^{s,sel}\in \mathcal{D}^{sel}$ be a sector selection rule. Given a distribution $G^s(\epsilon)$ and a $D^{s,sel}$, the associated correspondence $M^s$ is defined as: { \begin{equation} \begin{split} M^s(G^s)\equiv \bigg\{ F\in \mathcal{F}: &Pr_F(Y_i=y,D_i=d,Z_i=z)= (C^{y0z}_1+C^{y1z}_1)\mathbbm{1}(d=1)+(C^{0yz}_0+C^{1yz}_0)\mathbbm{1}(d=0),\\ &holds for some vector \left(C^{ykz}_{d}\right)_{y,k,z,d\in\{0,1\}} such that \\ &C^{ykz}_1+C^{ykz}_0=Pr_{G^s}(Y_i(1)=y,Y_i(0)=k,Z_i=z),\\ &C^{ykz}_d = 0 \quad if \quad D^{s,sel}(y,k,z)=\{1-d\}\quad and \quad C^{ykz}_d\ge 0 \quad \quad \forall y,k,z,d\in\{0,1\} \bigg\}. \end{split} \end{equation}}

In Definition (ref), $C^{ykz}_d$ is the probability of choosing sector $d$ when $Y_i(1)=y,Y_i(0)=k,Z_i=z$. When $D^{s,sel}(y,k,z)$ is the set $\{0,1\}$, a structure $s$ associated with $D^{s,sel}$ does not specify how a sector choice is determined, so the only constraint is $C^{ykz}_1+C^{ykz}_0=Pr_{G^s}(Y_i(1)=y,Y_i(1)=k,Z_i=z)$. When $D^{s,sel}(y,k,z)=\{d\}$, the constraint $C^{ykz}_{1-d}=0$ in the last row of ((ref)) requires the probability of choosing sector $1-d$ is zero. Given the set of sector selection rules $\mathcal{D}^{sel}$, we can specify the structure universe $\mathcal{S}$ in this application as follows:

equation[equation omitted — 264 chars of source]

Instead of imposing the strong independent instrument condition $\left(Y_i(1),Y_i(0)\right)\perp Z_i$, we require the instrument to have monotone effects on the potential outcomes:

definitionWe say $(Y_i(1),Y_i(0))|Z_i=1$ dominates $(Y_i(1),Y_i(0))|Z_i=0$ at the best and worst outcomes if \begin{equation} \begin{split} Pr(Y_i(1)=Y_i(0)=1|Z_i=1)&\ge Pr(Y_i(1)=Y_i(0)=1|Z_i=0),\\ Pr(Y_i(1)=Y_i(0)=0|Z_i=1)&\le Pr(Y_i(1)=Y_i(0)=0|Z_i=0). \end{split} \end{equation}

Definition (ref) only requires the instrument to generate the best (resp. worst) potential outcome $Y_i(1)=Y_i(0)=1$ (resp. $Y_i(1)=Y_i(0)=0$) with higher (resp. lower) probability at $Z_i=1$ than that at $Z_i=0$. This requirement is weaker than Assumption 5 in mourifie2018roy, where they require $Pr(Y_i(d)=1|Z_i=1)\ge Pr(Y_i(d)=1|Z_i=0)$ for $d\in \{0,1\}$ in addition to ((ref)). \footnote{ Condition ((ref)) along with the additional requirement $Pr(Y_i(1)=1|Z_i=1)\ge Pr(Y_i(1)=1|Z_i=0)$ will imply $1-Pr_F(Y_i=0,D_i=1|Z_i=1)\ge Pr_F(Y_i=1,D_i=1|Z_i=0)$ for all outcome distributions $F$. This implication holds even in the absence of the Roy selection condition. On the other hand, equation ((ref)) alone does not imply any constraints on $F$.}

Dominance at the best and worst outcome in Definition (ref) can accommodate broader empirical scenarios compared with Assumption 5 in mourifie2018roy. For example, suppose $Y_i(d)=1$ means individual $i$ gets tenure in sector $d$, and $Z_i=1$ means individual $i$ participates in a job training program. If the skill obtained from the training program can be applied to both sectors, we would expect $Pr(Y_i(1)=Y_i(0)=1|Z_i=1)\ge Pr(Y_i(1)=Y_i(0)=1|Z_i=0)$. On the other hand, each job sector may require specific skill that cannot be obtained from the job training program. If the training program is costly and prevents an individual $i$ from developing skills specific to the sector $d$, it is possible that $Pr(Y_i(d)=1|Z_i=1)\le Pr(Y_i(d)=1|Z_i=0)$. Such a scenario violates Assumption 5 in mourifie2018roy but not ((ref)). Lastly, $Pr(Y_i(1)=Y_i(0)=0|Z_i=1)\le Pr(Y_i(1)=Y_i(0)=0|Z_i=0)$ ensures the training program is beneficial: it increases the probability of success in at least one sector.

However, when ((ref)) is combined with Roy's selection assumption, they are jointly refutable. Formally, the assumption with Roy's selection rule and a monotone instrument satisfying ((ref)) is defined in the following.

assumptionThe assumption of Roy's selection rule with instrument condition ((ref)), denoted as $A^{Roy}$, is the collection of structures such that \begin{equation} \begin{split} A^{Roy}=\bigg\{s\in \mathcal{S}:\,\,& M^s\,\,is associated with the \,\,D^{sel}\,\,in\,\,((ref)),\\ &M^s(G^s)\,\,satisfies ((ref)),\quad G^s\,\, satisfies ((ref)) \bigg\}. \end{split} \end{equation}

The following proposition characterizes the non-refutability and confirmation sets associated with $A^{Roy}$.

propLet $\mathcal{F}^{nf}$ be the collection of outcome distributions such that \begin{equation} \begin{split} Pr_F(Y_i=0|Z_i=1)\le Pr_F(Y_i=0|Z_i=0). \end{split} \end{equation} The non-refutability and confirmation sets of $A^{Roy}$ are \[ \begin{split} \mathcal{H}_{\mathcal{S}}^{snf}(A^{Roy})&= \{s:\, M^s(G^s)\subseteq \mathcal{F}^{nf}\},\\ \mathcal{H}_{\mathcal{S}}^{wnf}(A^{Roy})&= \{s:\, M^s(G^s)\cap \mathcal{F}^{nf}\ne \varnothing \},\\ \mathcal{H}_{\mathcal{S}}^{scon}(A^{Roy})&=\mathcal{H}_{\mathcal{S}}^{wcon}(A^{Roy})=\varnothing . \end{split} \]

Proposition (ref) reveals several things: first $\mathcal{H}_{\mathcal{S}}^{snf}(A^{Roy})\ne \mathcal{S}$, so the assumption $A^{Roy}$ is refutable; second, $\mathcal{H}_{\mathcal{S}}^{wnf}(A^{Roy})\ne \mathcal{H}_{\mathcal{S}}^{snf}(A^{Roy})$ so we cannot find a strong extension of $A^{Roy}$. Therefore, I look at the completion of the binary sector choice model. The completion given in Definition (ref) is abstract and in the following I give an explicit form of $\mathcal{C}(s)$.

Recall that each $s$ in the incomplete space is associated with a sector decision rule, denoted by $D^{s,sel}$. We consider a tie breaking rule associated with $s$: $\left(C^{s^*,tb}_{d}(y,k,z)\right)_{y,k,z,d\in\{0,1\}}$, which specifies the probability of choosing sector $d$ for different values of $Y_i(0)=y,Y_i(1)=k,Z_i=z$. Different values of the vector $\left(C^{s^*,tb}_{d}(y,k,z)\right)_{y,k,z,d\in\{0,1\}}$ correspond to different selections. The completion $\mathcal{C}(s)$ is the collection of $M^{s^*}(\cdot)$ such that {

equation[equation omitted — 576 chars of source]

} Compared with the incomplete mapping in ((ref)), the tie breaking rule $C_{d}^{s^*,tb}$ makes $M^{s^*}(G^{s})$ a singleton. The completed universe is given in Definition (ref) and the corresponding Roy assumption set in the completed structural universe is given in Definition Proposition (ref).

Minimal Efficiency Loss and the Corresponding Strong Extension

I now consider a strong extension under the completed structure universe $\mathcal{S}^*$. I first define the efficiency loss of a structure $s^*$, which can be viewed as a deviation from Roy's sector selection assumption.

definitionThe efficiency loss of a structure $s^*\in\mathcal{S}^*$ \begin{equation} m^{EL}(s^*)=E_{G^{s^*}}[\max\{Y_i(1),Y_i(0)\}]-E_{F}[Y_i] \quad \quad for \quad F\in M^{s^*}(G^{s^*}) \end{equation} is the difference between the expected optimal sector selection outcome and the expected predicted outcome.

The efficiency loss is a function since $M^{s^*}(G^{s^*})$ is a singleton. It is easy to see that when the Roy sector selection rule holds, $m^{EL}(s^*)=0$. Conversely, by ((ref)), $m^{EL}(s^*)=0$ implies \[ D_i=

cases1 \quad \quad if \quad Y_i(1)>Y_i(0)\\ 0\quad \quad if \quad Y_i(1)<Y_i(0)\\

\] with probability 1, so the Roy sector selection rule holds. Once we verify $m^{EL}(s^*)$ is a well-behaved minimal deviation measure (see Definition (ref)), we can use minimal efficiency loss to construct a strong extension.

assumptionLet $m^{min}(F)\equiv \inf\{m^{EL}(s^*): F\in M^{s^*}(G^{s^*}) \,\, and \,\, G^{s^*}\text{ satisfies }(\ref{eq: roy model, dominating instrument at best and worst outcome}) \}$ be the minimal efficiency loss under $F$. We call \[\tilde{A}=\cup_{F\in \mathcal{F}} \left\{s^*: m^{EL}(s^*)=m^{min}(F),\,G^{s^*}\text{ satisfies }(\ref{eq: roy model, dominating instrument at best and worst outcome})\right\}\] the minimal efficiency loss extension of $A^{Roy}$.
prop$\tilde{A}^{*Roy}$ is a strong extension of $A^{*Roy}$. Moreover, given an observed distribution $F$, the identified set for $G^{s^*}$ under $\tilde{A}^{*Roy}$ is { \begin{equation} \begin{split} \Bigg\{ G^{s^*}\bigg| & there exists \{C_d^{ykz}\} \quad \forall {y,k,z,d\in\{0,1\}}\,\,s.t.\,\, C_d^{ykz}\ge 0,\\ &Pr_F(Y_i=y,D_i=d,Z_i=z)= (C^{y0z}_d+C^{y1z}_d)\mathbbm{1}(d=1)+(C^{0yz}_d+C^{1yz}_d)\mathbbm{1}(d=0),\\ &C^{ykz}_1+C^{ykz}_0=Pr_{G^{s*}}(Y_i(1)=y,Y_i(0)=k,Z_i=z),\\ & C^{010}_1=C^{100}_0=0,\quad\frac{C^{110}_1+C^{110}_0}{Pr_F(Z_i=0)}\le \frac{C^{111}_1+C^{111}_0}{Pr_F(Z_i=1)},\\ &C^{101}_0+C^{011}_1=\max\left\{Pr_F(Y_i=0,Z_i=1)-\frac{Pr_F(Y_i=0,Z_i=0)Pr_F(Z_i=1)}{Pr_F(Z_i=0)},0\right\}\Bigg\}. \end{split} \end{equation}}

Proposition (ref) characterizes the sharp identified set of distributions of $(Y_i(1),Y_i(0),Z_i)$. The identified set ((ref)) under $\tilde{A}^{*Roy}$ satisfies: (1). There is no efficiency loss when $Z_i=0$ ($C^{010}_1=C^{100}_0=0$); (2). The minimal efficiency loss is $\max\left\{Pr_F(Y_i=0,Z_i=1)-\frac{Pr_F(Y_i=0,Z_i=0)Pr_F(Z_i=1)}{Pr_F(Z_i=0)},0\right\}$; (3). Condition ((ref)) holds as long as $(C^{110}_1+C^{110}_0){Pr_F(Z_i=1)}\le (C^{111}_1+C^{111}_0){Pr_F(Z_i=0)}$ holds. The identified set of $G^{s^*}$ is a polyhedron characterized by the 16-dimensional vector $(C_d^{ykz})_{y,k,z,d\in\{0,1\}}$. Many parameters of interest are linear functions of $(C^{jkz}_d)$, and linear-programming can be used to find the identified set. One example is given in Corollary (ref).

CorollaryThe identified set for $Pr(Y_i(1)=1|Z_i=z)$ under $\tilde{A}^{*Roy}$ is \[ \begin{split} Pr(Y_i=1,D_i=1|Z_i=0)\le &Pr(Y_i(1)=1|Z_i=0)\le Pr(Y_i=1|Z_i=0),\\ Pr(Y_i=1,D_i=1|Z_i=1)\le &Pr(Y_i(1)=1|Z_i=1)\le Pr(Y_i=1|Z_i=1)+ \frac{m^{EL,min}(F)}{Pr(Z_i=1)},\\ \end{split} \] where $ m^{EL,min}(F)=\max\left\{Pr_F(Y_i=0,Z_i=1)-\frac{Pr_F(Y_i=0,Z_i=0)Pr_F(Z_i=1)}{Pr_F(Z_i=0)},0\right\}$.

It is worth noticing that if $A^{*Roy}$ cannot be rejected by $F$, the identified set for $\theta\equiv Pr(Y_i(1)=1|Z_i=1)$ is given by $[Pr(Y_i=1,D_i=1|Z_i=1),Pr(Y_i=1|Z_i=1)]$. However, suppose we ignore the testable implication of $A^{*Roy}$ and directly use $[Pr(Y_i=1,D_i=1|Z_i=1),Pr(Y_i=1|Z_i=1)]$ as the identified set for $\theta$, we get a spuriously informative identified set when $m^{EL,min}(F)>0$. This happens when the true structure that generates the data lies in $\tilde{A}^{*Roy}\backslash {A}^{*Roy}$, and the spuriously informative identified set is a proper subset of the true identified set.

The upper bound of the identified set for $Pr(Y_i(1)=1|Z_i=1)$ is not Fr\'echet differentiable with respect to $F$ due to the $\max$ operator. However, by Example 2 in fang2019inference, the upper bound $Pr(Y_i=1|Z_i=1)+ \frac{m^{EL,min}(F)}{Pr(Z_i=1)}$ is directionally differentiable in $F$. The bootstrap method in fang2019inference can be used to construct confidence interval for $Pr(Y_i(1)=1|Z_i=1)$. However, since $Pr(Y_i(1)=1|Z_i=1)$ is partially identified and satisfies the conditions in Proposition (ref), we cannot directly test a hypothesis on the value of $Pr(Y_i(1)=1|Z_i=1)$. Instead, we should test the equivalent existence hypothesis as in Proposition (ref).