Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
118,921 characters · 25 sections · 46 citation commands
Extending Economic Models with Testable Assumptions: Theory and Applications
\affil[1]{Nanjing University}
Empirical researchers often make convenient model assumptions in structural estimation. These assumptions usually come from economic theories or intuitions. For example, the `No Defiers' assumption in IA1994 assumes that the instrument has a monotone effect on the decision to take treatment; the `Pure Strategy Nash Equilibrium' assumption in BresnahanReiss1991 assumes only pure strategy Nash Equilibrium is played in a $2\times 2$ entry game; the `Perfect Self Selection' assumption in Roy1951 assumes employees perfectly observe their future earnings and choose a job sector to maximize discounted lifetime earnings. Such assumptions simplify the identification and estimation problems, and make the results easier to interpret. To study structural models and economic assumptions, I generalize the language of econometric structures koopmans1950identification to incomplete structures. A (generalized) econometric structure includes a distribution of some exogenously given latent and observed variables, and a correspondence from the distribution of these exogenous variables to a set of distributions of observed variables. Assumptions are restrictions on the economic structures to reflect empirical researchers' understanding of the economic environment.
Unfortunately, assumptions in the three examples above, when combined with some other reasonable assumptions, can be rejected by some distributions of observables kitagawa2015,mourifie2017testing,mourifie2018roy. When the imposed assumption is refuted by data, the econometrician have an empty identified set for the parameter of interest. As a result, the econometrician cannot give a useful interpretation of the economic environment. A refutable assumption $A$ also imposes challenges to the interpretation of hypothesis testing of parameter values. For example, when we reject the null hypothesis that the parameter equals zero under $A$, it can either be the case that the true parameter equals another value, or the case that $A$ is rejected data. These two cases cannot be distinguished but they have quite different interpretations.
A way to prevent the data rejection problem is to find a relaxed assumption $\tilde{A}$ so that no distributions of observables can reject $\tilde{A}$. We call this the non-refutability criterion. By imposing a non-refutable $\tilde{A}$ before confronting the data, practitioners avoid the ex-post possibility of finding data evidence against their assumption. Therefore, non-refutability is the first criterion for the relaxed assumption $\tilde{A}$ to satisfy. On the other hand, we also do not want to deviate from the old assumption $A$, since it still reflects the economic theory behind it. Specifically, given a parameter of interest $\theta$ as a function of structures, we want the (sharp) identified set under the relaxed assumption $\tilde{A}$ to be equal to the (sharp) identified set under $A$, when $A$ is not rejected by the observed distribution. In other words, we want to preserve the identified set.
This paper aims to do three things: First, I formalize and extend the definition of refutability and confirmability of assumptions in breusch1986 using the language of generalized economic structures. These definitions are useful to characterize properties of a relaxed assumption $\tilde{A}$. I characterize conditions that a relaxed assumption $\tilde{A}$ needs to satisfy so that: (a).No distribution of observables can be rejected under $\tilde{A}$; (b).The identified sets for the parameter of interest under $A$ and $\tilde{A}$ are the same whenever $A$ is not rejected by the observed data distribution. I show that when structures are complete, a relaxed assumption $\tilde{A}$ that satisfies the two properties above always exists. I also characterize a sufficient condition for the existence of $\tilde{A}$ in case of incomplete structures. The possible failure to find $~\tilde{A}$ in incomplete structures encourages researchers to complete the structures, and then find a nice relaxed assumption in the completed structure universe. When the structures are complete, I also provide a general method to construct $\tilde{A}$ from $A$.
Second, I discuss the problem of testing hypothesis on structural parameters' values. I show that any statistical test cannot achieve pointwise size control and test consistency simultaneously when the null hypothesis of structural parameters does not induce a partition of the space of distribution of observables. This is an ill-behaved null hypothesis, and policy decisions based on the result of hypothesis testing can be problematic. Conversely, with a well-behaved null hypothesis, which induces a partition of the space of distribution of observables, I show the existence of statistical tests that achieve pointwise size control and test consistency simultaneously under mild conditions. I also show that when the parameter of interest $\theta$ is point identified, a null hypothesis on the value of $\theta$ is always well-behaved. When structures are complete, I also provide a way to minimally extend (resp. shrink) the null hypothesis set such that the extend (resp. shrink) hypothesis is well-behaved.
Third, I look at 2 applications with complete and incomplete structures respectively. In the complete structure framework, I look at the identification of the local average treatment effects (LATE). kitagawa2015 provides the sharp testable implication of the Imbens and Angrist Monotonicity assumption (IA-M). Therefore, practitioners should anticipate the IA-M to be rejected by some distributions of observables. I provide several relaxations of the IA-M that cannot be rejected by any distribution of observables. The identified sets for LATE under these relaxed assumptions equal the identified set for LATE under the IA-M whenever the IA-M is not rejected by the distributions of observables. One relaxed assumption allows for defiers, and relaxes the independent instrument assumption. The logic of the relaxed assumption is to allow a minimal mass of defiers. It can be shown that the relaxed assumption not only preserves the identified set for LATE, but also preserves the identified set for other parameters of interests such as ATE or ATT. LATE is point identified under this relaxed assumption. I provide an estimator of the LATE which has a normal limit distribution. {To emphasis the fact that a non-refutable relaxation is not unique, I propose two other relaxations, each one of which may be preferable in some contexts.} I apply the method to card1993using, and I show the local average treatment effect of education on earnings. Compared to naively using identification result under the IA-M assumption, my method delivers more reasonable sign and scale for the LATE estimates.
For incomplete structures, I look at a binary outcome job sector selection model with a monotone instrument. In the sector choice model, the Roy assumption does not specify the sector choice rule in case of ties, which may lead to multiple predicted distributions of observables. After completing the structures, I then use a `minimal efficiency loss' criterion to characterize the relaxed assumption. The identified set of job sector potential outcome distribution can be characterized as a polyhedron.
\paragraph{Related Literature} masten2018 propose an ex-post way to salvage a refutable assumption $A$. Their ex-post method characterizes a relaxation of $A$ after the distribution of observables is realized. This paper also relates to the literature that relaxes assumption make model robust to misspecification. In the macroeconomic literature, researchers use robust control to avoid the misspecification issue in their baseline model (see hansen2006robust and hansen2007recursive). The Robust control approach aims to accommodate local perturbations to the baseline model rather than to solve the refutability of the baseline model.\footnote{The perturbation is usually measured by relative entropy the in macroeconomic literature.} It may be true that the baseline model is not refutable by any data distribution. See bonhomme2018minimizing and christensen2019counterfactual for more discussion.
{This paper also contributes to the literature that relaxes the IA-M assumption. chaisemartin2018Defier discusses the economic meaning of the conventional LATE quantity $LATE^{Wald}\equiv ({E[Y_i|Z_i=1]-E[Y_i|Z_i=0]})/({E[D_i|Z_i=1]-E[D_i|Z_i=0]})$ when there are defiers. He shows that the $LATE^{Wald}$ identifies the net average treatment effect of a subgroup of compliers after deducting the average treatment effect of defiers.}
The rest of the paper is organized as follows. Section 2 describes a theory of characterizing a refutable assumption $\tilde{A}$, finding a relaxed assumption $\tilde{A}$ and hypothesis testing. Core definitions in Section 2 are followed by shadowed links where their corresponding illustrations can be found. Section 3 applies complete structure theory to the model in IA1994. Section 4 applies the theory to a binary outcome Roy model. Main proofs are collected in Appendices. Additional proofs and results are collected in the Online Appendices.
Throughout this paper, I use $X$ to denote the vector of observed variables, and I use $F$ to denote the distribution of $X$. I use $\epsilon$ to denote the vector of latent variables and some observed variables, and I use $G$ to denote the distribution of $\epsilon$. I use $s$ to denote an economic structure. I use $A$ to denote a refutable assumption and use $\tilde{A}$ to denote a relaxed assumption. I use $\mathbb{F}_n$ to denote the empirical distribution of $F$.
In this section, I develop a theory for identification and refutable assumptions. I will then discuss the problem of hypothesis testing. I start with a definition of the observation space.
The distribution of observables $F(X)$ is generated by some distribution of underlying random vectors $\epsilon$ through some mapping $M$. A pair of a distribution of $\epsilon$ and a mapping $M$ is called an econometric structure. The following definition of econometric structure is a reformulation of the economic structure defined in koopmans1950identification and jovanovic1989. Since in most econometric problems, we focus on the distribution of outcomes $F$ instead of how each $X$ is related to $\epsilon$, I directly define the mapping $M$ as a correspondence from the space of underlying variable distributions to $\mathcal{F}$.
The mapping $M^s$ of structure $s$ relates the distribution $\epsilon$ to the distributions of $X$. Since the distributions of $X$ are generated by the distributions of $\epsilon$, we call $\epsilon$ the primitive variables. Definition (ref) allows overlap between $X$ and $\epsilon$. A collection of structures is called a structure universe. We want to learn the distribution of $\epsilon$ and the mapping $M$ from the distribution of observables $F$.
Here I explicitly distinguish the structure universe $\mathcal{S}$ and the assumption $A$, though both are just a collection of structures. The structure universe $\mathcal{S}$ is the paradigm that can span different empirical contexts. On the other hand, an assumption $A$ places constraints that are suitable for a particular empirical context, or convenient for empirical analysis.
The condition $\cup_{s\in\mathcal{S}}M^s(G^s)=\mathcal{F}$ requires that all possible distributions of observables can be generated by some structure in the universe. Moreover, no distribution outside $\mathcal{F}$ can be generated by $\mathcal{S}$.
The definition of completeness is slightly different from the definition of completeness in tamer2003incomplete. tamer2003incomplete defines a model to be complete if the mapping from $\epsilon$ to $X$ is a singleton and non-empty, and incomplete if the mapping has multiple outputs, and incoherent if the mapping generates no output. Here, my definition of economic structure does not specify the mapping from each $\epsilon$ to $X$. Instead, I consider the mapping from the distribution of $\epsilon$ to the distribution of $X$. If the probability of multiple outcome is non-zero, an incomplete model in tamer2003incomplete implies an incomplete structure in my definition.
My definition of structure, however, does not have a corresponding terminology for incoherent model. This is because I require $\cup_{s\in\mathcal{S}}M^s(G^s)=\mathcal{F}$ to hold. chesher2012simultaneous propose four ways to deal with model incoherence and derive the distribution of observables under the model. My definition of $M^s$ as the mapping between distributions can be viewed as the consequences of chesher2012simultaneous.
The notions of refutability and confirmability of a complete structure are given in breusch1986. If an assumption is refutable, then there exists some $F$ that can reject $A$. The notions of refutability and confirmability are stated in terms of the observation space $\mathcal{F}$. Equivalently, we can characterize refutability and confirmability in terms of the structure universe $\mathcal{S}$. To do this, I first define the non-refutability and confirmation sets associated with $A$.
To accommodate the incomplete structures, we define two types of non-refutability sets. We call $\mathcal{H}_{\mathcal{S}}^{snf}(A)$ the strong non-refutability set associated with $A$, because if the true structure $s$ is in $\mathcal{H}_{\mathcal{S}}^{snf}(A)$, then for any distributions of observables in $M^s(G^s)$, we cannot refute $A$. In contrast, we call $\mathcal{H}_{\mathcal{S}}^{wnf}(A)$ the strong non-refutability set associated with $A$, because if the true structure $s$ is in $\mathcal{H}_{\mathcal{S}}^{wnf}(A)$, then for some distributions of observables in $M^s(G^s)$, we cannot refute $A$.
The strong confirmation set is the collection of structures that cannot be observationally equivalent to any structures outside $A$ for any observed distribution $F$. Weak confirmation set is the collection of structures that cannot be observationally equivalent to any structures outside $A$ for some observed distribution $F$. I call $\mathcal{H}^{scon}_{\mathcal{S}}(A)$ (resp. $\mathcal{H}^{wcon}_{\mathcal{S}}(A)$) the strong (resp. weak) confirmation set associated with $A$, because if the true structure $s$ is in $\mathcal{H}^{scon}_{\mathcal{S}}(A)$ (resp. $\mathcal{H}^{wcon}_{\mathcal{S}}(A)$), then for all (resp. some) distribution of observables in $M^s(G^s)$, we can confirm that the true structure must lies in $A$. In particular, \[\mathcal{H}_\mathcal{S}^{scon}(A)\subseteq \mathcal{H}_\mathcal{S}^{wcon}(A)\subseteq A\subseteq \mathcal{H}_\mathcal{S}^{snf}(A)\subseteq \mathcal{H}_\mathcal{S}^{wnf}(A).\] When the structure universe $\mathcal{S}$ is complete, $M^s(G^s)$ is always a singleton, and $\mathcal{H}_\mathcal{S}^{scon}(A)= \mathcal{H}_\mathcal{S}^{wcon}(A)$, $\mathcal{H}_\mathcal{S}^{snf}(A)= \mathcal{H}_\mathcal{S}^{wnf}(A)$. The following proposition helps to interpret the confirmation sets associated with $A$ as the non-refutability sets associated with $A^c$. It also shows that the strong non-refutability set and weak confirmation set as operation are idempotent.
The definition of refutability and confirmability in breusch1986 is defined on the outcome space $\mathcal{F}$, but we can also characterize it on the structure universe $\mathcal{S}$.
By definition, we should have $\cup_{s\in \mathcal{H}_{\mathcal{S}}^{snf}(A)} M^{s}(G^{s})=\cup_{s\in A} M^{s}(G^{s})$, so $\mathcal{H}_{\mathcal{S}}^{snf}(A)$ is refutable if and only if $A$ is refutable. In many cases, it is easy to check whether $\mathcal{H}_{\mathcal{S}}^{snf}({A})=\mathcal{S}$ in Proposition (ref) than to check Definition (ref).
In many empirical studies, we want to find the value of a parameter of interest rather than a class of structures that are consistent with data. This parameter can be a moment of unobserved primitive variables, or a counterfactual outcome of the structure. The parameter of interest can also give interpretation on the causal relation between outcome variables and primitive variables. Imposing strong assumptions helps to restrict the set of data-consistent parameter values, but an imposed assumption $A$ as in Definition (ref) may be rejected by some distribution of observables. Therefore, in many empirical studies, researchers often first present some summary statistics that justify the assumption. If the assumption is rejected by the data, researchers can move to another assumption. This is an ex-post way of choosing a relaxed assumption. There are two major problems with this approach. First, such justifications are heuristic pre-testing procedures of assumption $A$, and any subsequent inference on the parameter of interest may have incorrect size control due to pre-testing. Second, researchers do not specify what they will do if $A$ is rejected. Most likely they will choose another assumption that will not be rejected by the data. To avoid the pre-testing issue, I propose to solve the problem from an ex-ante perspective, i.e. impose a non-refutable assumption before any distribution of observables is realized.
I first formalize the definition of an identification system and discuss how to deal with an existing situation, where assumption $A$ may be rejected by the data.
A parameter of interest can take a very general form. It can be the structure $s$ itself, or it can be a counterfactual outcome. For example, suppose $M^s$ is known up to a finite dimensional vector: $M^s(\cdot)=M(\cdot\,;(\beta_1,...,\beta_k))$. Further suppose the objective of our counterfactual analysis is to find the predicted distribution of observables when $\beta_1=0$. Then the parameter of interest $\theta$ can be defined as $\theta(s)= M(G^s;(0,...,\beta_k))$. For a parameter of interest $\theta$, $\Theta^{ID}_A(F)$ is the set of parameters that are compatible with the data. For an empirical researcher, the main concern of the partial identification method is the possibility of an empty identified set. Here I characterize the equivalent condition of an empty identified set.
Here is an intuition of Proposition (ref): If we have a non-empty identified set for an $F$, then there must exist a structure in $A$ that rationalizes $F$. Since this is true for all $F$, $A$ is non-refutable. Conversely, if $A$ is non-refutable, then for any $F$ we can find a structure $s\in A$ to rationalize $F$, and the corresponding $\theta(s)$ must lie in the identified set.
For a refutable assumption $A$, there exist some $F$ such that $\Theta_A^{ID}(F)=\varnothing$. This can be unsatisfying because empirical researchers cannot directly interpret the distribution of $\epsilon$. To avoid this, before seeing any outcome distribution, a practitioner can impose a relaxed assumption $\tilde{A}$ such that $\mathcal{H}_\mathcal{S}^{snf}(\tilde{A})= \mathcal{S}$ and $A\subseteq \tilde{A}$.
These three definitions are nested. A well-defined extension ensures that the identified set will never be empty; A $\theta$-consistent extension preserves the identified set for a parameter of interest $\theta$. A strong extension moreover ensures that the identified set for any parameter of interest will be preserved. In different empirical settings, researchers' parameters of interest can differ. If a strong extension is found, researchers can use this extension across different empirical contexts. The following proposition gives a characterization of whether $\tilde{A}$ is a strong extension.
In a complete structure universe, we can always find a strong extension $\tilde{A}$. This is a major difference between complete and incomplete structure universe.
In other words, $\tilde{A}=A\cup [\mathcal{H}_\mathcal{S}^{snf}(A)]^c$ is the strong extension that puts the least structural assumption outside $\mathcal{H}_\mathcal{S}^{snf}(A)$. Should we always use $\tilde{A}=A\cup [ \mathcal{H}_\mathcal{S}^{snf}(A)]^c$ as the choice of strong extension when the structural universe $\mathcal{S}$ is complete? Unfortunately, using the maximal strong extension will lead to very badly behaved identified set $\Theta_{\tilde{A}}^{ID}(F)$. Suppose $\mathcal{F}$ is equipped with some metric $d$, and $F_0$ is on some part of the boundary of the set of predicted observable distribution $\cup_{s\in {A}} M^s(G^s)$. It is possible that $\Theta_{\tilde{A}}^{ID}(F_0)$ gives an informative bound (i.e. $\Theta_{\tilde{A}}^{ID}(F_0)\ne \Theta$) on the parameter of interest , but for an $F'\in\left[\cup_{s\in {A}} M^s(G^s) \right]^c$ that is arbitrarily close to $F_0$, the identified set $\Theta_{\tilde{A}}^{ID}(F')$ is uninformative (i.e. $\Theta_{\tilde{A}}^{ID}(F')=\Theta$). See the identification result in Proposition (ref) for an illustration. This raises two concerns. First, at the identification level, the interpretation of an uninformative identified set $\Theta_{\tilde{A}}^{ID}(F')$ under the maximal strong extension $\tilde{A}$ is not very different from an empty identified set $\Theta_{A}^{ID}(F')=\varnothing$ under the original assumption $A$. An uninformative identified set says that any parameter value of $\theta$ is compatible with data, while an empty identified set says that no parameter value of $\theta$ is compatible with data. In either case, the identification result does not help us to interpret the environment. Second, we may get spurious informative inference result due to sampling error. When $F'$ is close to the boundary of $\cup_{s\in {A}} M^s(G^s)$ but not in it, sampling error may lead us to a spurious but informative bound $\Theta_{\tilde{A}}^{ID}(F')$, even if the true identified set should have been uninformative. In other words, estimated identified set is not consistent. The maximal strong extension $\tilde{A}=A\cup [\mathcal{H}_\mathcal{S}^{snf}(A)]^c$ solves the refutability issue, but it imposes too few constraints outside $A$ to generate informative result on $\theta$.
For an incomplete structure universe, there is a gap between the strong and weak non-refutability sets. As a result, we may not be able to find a strong extension of $A$.
This situation happens when there is a nesting relation between $A$ and $A^c$: suppose for each structure $s\in A$, we can find an $s^*\in A^c$ such that $M^s(G^s)\subseteq M^{s^*}(G^{s^*})$; and for every $s^*\in A^c$ we can find an $s\in A$ such that $M^s(G^s)\subseteq M^{s^*}(G^{s^*})$. If $A$ is refutable, then $\mathcal{H}_{\mathcal{S}}^{snf}(A)\ne \mathcal{S}$. However the nesting relation implies $\mathcal{H}_{\mathcal{S}}^{wnf}(A)=\mathcal{S}$. Then both conditions in Proposition (ref) hold, and there exists no strong extension of $A$.
In many cases, we can find a function $m:\mathcal{S}\rightarrow \mathbb{R}_+$ such that $m(s)=0$ for all $s\in A$ \footnote{See ((ref)), ((ref)) for example.}. While an assumption $A$ may be refutable, we impose $A$ in the first place because it reflects economic theory suitable in the empirical context. Therefore we would consider a departure from $A$ is abnormal and is against the economic intuition behind $A$. I extend our assumption to allow for a minimal departure from the baseline assumption $A$. This way to relax assumption $A$ is called the minimal deviation method. Formally, suppose the refutable assumption $A$ can be written as an intersection of several larger assumptions: $A=\cap_{j=1}^J A_j$. This representation allows us to consider a departure from a particular $A_j$.
We want the relaxation measure to be well-behaved such that when we push the deviation to infinity, we can generate any distributions in $\mathcal{F}$. The well-behaved condition ensures that there exists a structure in $\cap_{l\ne j}A_l$ that can achieve the minimal measure. This is essential for the construction of a extension using $m_j$. See Proposition (ref) is not well behaved if use full independence} for examples of ill-behaved measures. Now I construct the minimal deviation extension $\tilde{A}$.
The construction in Definition (ref) only relaxes the assumption $A_j$ and keeps other assumptions unchanged. The following proposition shows a way to check whether a minimal deviation extension is $\theta$-consistent or strong consistent.
To facilitate the discussion of criteria of choosing among multiple extensions, I assume $\mathcal{F}$ is endowed with a metric $d_F$. Moreover, I fix the parameter of interest $\theta$, and assume $\Theta$ is endowed with a metric $d_\theta$.
In some cases, there can be multiple ways to write an assumption, i.e. $A=A_j\cap(\cap_{l\ne j} A_l)=A_j\cap(\cap_{l\ne j} A_l^\prime)$. Fix a $j$, even if we use the same measure $m_j$, since the minimal deviation is defined with respect to $\{A_l\}_{l\ne j}$, the extension can differ when using a different representation. It should also be noted that to check whether $m_j$ is a well-defined relaxation measure, we need to look at $\{A_l\}_{l\ne j}$. This means $m_j$ can be a well-defined relaxation measure with respect to $\{A_l\}_{l\ne j}$ but not $\{A_l^\prime\}_{l\ne j}$. Here I leave the choice of representation of the assumption to researchers and discuss the issue of multiple extensions that arises from two aspect: which assumption to relax and the choice of relaxation measure.
Each minimal deviation extension $\tilde{A}$ corresponds to an index $j$ and a relaxation measure $m_j$. Given a representation of $A=\cap_{j=1}^J A_j$, we can choose which sub-assumption $A_j$ to relax, and we can also choose the relaxation measure $m_j$. Different choices of which assumption $j$ to relax, and different relaxation measures $m_j$ will result in different relaxed assumptions. Moreover, relaxed assumptions constructed from different relaxation measures can be non-nested with each other. Here, I discuss two criteria to choose an assumption among non-nested relaxed assumptions.
First, the relaxed assumption $\tilde{A}$ should be suitable for the empirical context. Given the empirical context, if we can find an economic story such that the $j$-th assumption in $A=\cap_{j=1}^J A_j$ fails, we will focus on finding a well-behaved relaxation measure $m_j$ corresponding to $A_j$. Second, we want some continuity property of the identified set with respect to the outcome distribution $F$.
Recall that we fix the parameter of interest $\theta$ in the beginning of this section. An $\tilde{A}$ that satisfies Property (ref) for $\theta$ may fail Property (ref) for a different parameter of interest. Without the continuity property, a consistent estimator of the identified set may not exist, and the identified set can be spuriously informative due to sampling error. Examples of discontinuous and continuous identified set correspondences can be found in Proposition (ref). Sufficient conditions to check Property (ref) and the further reasoning of Property (ref) can be found in Appendix (ref).
We have seen that for an incomplete structure universe, a refutable assumption $A$ may not have a strong extension. This is because in incomplete structures, we are agnostic about how distributions of outcomes are selected. If a structure $s$ has two predicted outcome distribution $M^s(G^s)=\{F_1,F_2\}$, we consider a completion procedure that separates $s$ into two complete structures $s_1^*$ and $s_2^*$ such that $M^{s_1^*}(G^{s_1^*})=\{F_1\}$ and $M^{s_2^*}(G^{s_2^*})=\{F_2\}$. The completion procedure then allows us to distinguish $s_1^*$ from $s_2^*$ by observing either $F_1$ or $F_2$.
Definition (ref) considers all possible completions $C$. The cardinality of $\mathcal{C}(s)$ is the same as that of $M^s(G^s)$. The completion procedure is without loss of generality, since all possible selections are considered. The key property is that for any parameter of interest, the identified set is not changed if the parameter of interest in the completed structure is properly defined in the following way.
In many cases, finding a strong extension is not feasible for an incomplete structure universe, but feasible for its completion. See Proposition (ref) for an illustration.
In empirical research, a commonly asked question is whether we can tell if the true value of the parameter of interest lies in a set, which can be written as a hypothesis $H$ on the parameter value. In this section, I consider the following formulation of a hypothesis $H$ on a structural parameter $\theta$ under a non-refutable assumption $\tilde{A}$: $H=\{s\in \tilde{A}: \theta(s)\in \Theta^0\}$, where $ \Theta^0$ is a parameter value set. The implicit alternative is $H^c\cap \tilde{A}$. Here I only consider non-refutable assumption $\tilde{A}$. If an assumption $A$ is refutable and cannot generate all distributions of observables, then for some distributions of observables, we cannot make say at least one of $H$ and $H^c\cap A$ holds true.
Policy makers sometimes use the result of hypothesis testing of a parameter value to guide their policy decisions. This decision procedure is called the `inference-based' approach in manski2019econometrics and is a conventional practice in medical treatment policy decision manski2020covid. However, the `inference-based' policy decision approach can be problematic if the hypothesis on parameter value $H$ does not induce a partition on the observation space $\mathcal{F}$: if both $H$ and $H^c\cap \tilde{A}$ can generate some observed distribution $F_0$, then we cannot tell whether $H$ holds by observing $F_0$. To formally discuss this issue, I first discuss the `hypothesis testing' problem assuming that I know the distribution of observables. I call this the binary decision problem \footnote{The same problem is called the binary choice problem in manski2019econometrics. To avoid the confusion with concepts in the discrete choice literature, I slightly change the name.}.
If condition 1 in Definition (ref) holds, it implies that the true structure $s$ that generates $F$ must be in $H$, since $F$ cannot be predicted by $H^c\cap \tilde{A}$, and this confirms $s\in H$; if condition 2 in Definition (ref) holds, it implies that the true structure $s$ cannot be in $H$, since $F$ cannot be predicted by $H$, and this refutes $s\in H$. If both conditions fail, it means $F$ can be predicted by structures both inside and outside $H$, which creates an ambiguity in the binary decision problem. If $H$ can be decided by any $F$, we say it is strongly binary decidable.
The following lemma provides an equivalent condition to check whether $H$ is strongly binary decidable. The lemma below uses the definition of non-refutability set (Definition (ref)) and confirmation set (Definition (ref)) under $\tilde{A}$ instead of $\mathcal{S}$.
Intuitively, Lemma (ref) says that if we can confirm that the true structure $s$ is in $H$ for all distributions of observables, then we can refute $H^c\cap \tilde{A}$ for all distributions of observables.
Now I consider statistical testing of $H$ based on a finite sample. We want to test the null hypothesis that the true structure $s_0$, which generates the outcome distribution $F$, satisfies hypothesis $H$ against its complement in $\tilde{A}$: \[ \mathcal{H}_0: s_0\in H\quad \quad v.s. \quad \quad \mathcal{H}_1: s_0\in H^c\cap \tilde{A}. \] We have a finite sample of $i.i.d$ realizations from $F$ with empirical distribution $\mathbb{F}_n$ that converges weakly to $F$. A statistical test $T_n$ is a binary function that maps the empirical distribution and some random vector $\mathbf{\eta}$ to $\{0,1\}$: \[ T_n(\mathbb{F}_n,\mathbf{\eta})=
\]
The names `structural size' and `structural test consistency' come from the fact that we construct the criteria ((ref)), ((ref)) through a partition of the assumption $\tilde{A}=H\cup[H^c\cap\tilde{A}]$ rather than a partition of the observation space $\mathcal{F}$. Structural size and structural power are what we care about since we aim to make a statement on the true structural parameter value. In particular, we may want to make binary decision on counterfactual outcomes. As we discuss after Definition (ref), a counterfactual analysis can be written as a parameter of interest.
The following proposition shows strongly that binary decidability is closely related to structural size control and structural test consistency.
The converse of this proposition also holds under further regularity conditions: if $H$ is strongly binary decidable under $\tilde{A}$, we can always find a test statistic that achieves pointwise size control and test consistency. Let \[
\] be the collection of all empirical distributions supported on a finite subset of $supp(X)$, and $Pr_{\mathbb{F}_n}(X_i=x)$ can be written as a fraction.
The following proposition shows that hypotheses about a point identified parameter of interest are always strongly binary decidable.
Proposition (ref) shows that the traditional hypothesis testing approach works in a point identified model, regardless of the hypotheses on the parameter of interest. However, for a partially identified model, the formulation of a hypothesis is crucial. Let's consider the following policy decision rule: `we implement a policy $P$ if and only if the true structure is in $H$. When $H$ is not strongly binary decidable, we have size and power issue for any test statistic $T_n(\mathbb{F},\eta)$. If we decide to implement $P$ if and only if $T_n(\mathbb{F},\eta)=1$, we also know that the testing procedure cannot reject structures in $H^c$ consistently. If the policy $P$ is harmful when the true structure $s$ does not satisfy the parameter constraint of $H$, and we implement $P$ when $T_n(\mathbb{F},\eta)=1$, the policy $P$ can be harmful to the economy.
The problem does not arise from the sampling error but arises from the intrinsic inability to distinguish $H$ and $H^c$ by the distribution of observables. If we want to use a decision rule based on a hypothesis $\tilde{H}$ such that `we implement the policy $P$ if and only if the true structure is in $\tilde{H}$, the hypothesis $\tilde{H}$ must be strongly binary decidable. \footnote{An alternative approach is to formulate the hypothesis testing problem as a statistical decision problem, see Section 2.3 in manski2019econometrics for discussion.}
The next question is whether we can find a strongly binary decidable extended set or subset.
If the benefit to correctly implement a policy $P$ when $H$ is true is large, and the cost of mistakenly implementing $P$ when the true structure is in $H^{ext}\backslash H$ is small, we may want to test $H^{ext}$. Conversely, if there is a huge cost when we implement $P$ if $H^c$ is true, we may want to test $H^{sub}$. In this case, we sacrifice the benefit when the true structure is in $H\backslash H^{sub}$ to avoid the risk of mistakenly implementing $P$. The following proposition provides the minimal (resp. maximal) strongly binary decidable extension (resp. subset set).
For complete structure universes, $\mathcal{H}_{\tilde{A}}^{snf}(H)=\mathcal{H}_{\tilde{A}}^{wnf}(H)$ and $\mathcal{H}_{\tilde{A}}^{scon}(H)=\mathcal{H}_{\tilde{A}}^{wcon}(H)$ hold automatically, so we can always find a non-trivial strongly binary decidable extension (subset set). In the following, I present an example with a complete structure universe.
In the example above, the non-refutability set and confirmation set associated with $H$ are easy to find, while in more complicated structural models, the non-refutable and confirmation sets can be hard to characterize. In a complete structure universe, if $ H$ is not a strongly binary decidable, and $H$ is refutable (resp. confirmable), we want to instead test $s_0\in \mathcal{H}_{\tilde{A}}^{snf}(H)$ (resp. $s_0\in \mathcal{H}_{\tilde{A}}^{wcon}(H)$), which is strongly binary decidable. The following proposition shows that in a complete structure universe, testing $s_0\in \mathcal{H}_{\tilde{A}}^{snf}(H)$ can be equivalently written as a test of the existence of a structure that rationalizes data.
In this section, I apply the method to IA1994 with a binary treatment and a binary instrument. The observed outcome variable $Y_i$ and treatment decision $D_i$ are generated through
where $D_i(1),D_i(0)$ are potential treatment decisions, $Y_i(d,z)$ are the potential outcome and $Z_i$ is a binary instrument.
Primitive variables are $\epsilon_i=(D_i(1),D_i(0),Y_i(0,0),Y_i(1,0),Y_i(0,1),Y_i(1,1),Z_i)$ and observed variables are $X_i=(Y_i,D_i,Z_i)$. Let $\mathcal{Y}$ be the space of $Y_i$ and let $\mathcal{B}$ be a Borel-sigma algebra on $\mathcal{Y}$. The observation space is
and space of potential distribution
All structures agrees on the functional relation between $X_i$ and $\epsilon_i$ specified in ((ref)). The mapping\footnote{See Definition (ref)} $M^s$ is defined as:
$M^s$ contains exactly one predicted distribution of observables and all structures are complete. The structure universe $\mathcal{S}$ is:
Following kitagawa2015, I define the following two quantities for all $B\in \mathcal{B}$ and $d\in \{0,1\}$:
The Imbens-Angrist Monotonicity assumption (IA-M) assumes exogeneity, exclusion and monotonicity of the instrument $Z_i$:
kitagawa2015 derives the sharp testable implications of the IA-M assumption ((ref)). I reformulate the result in the language of non-refutability sets in the following lemma.
kitagawa2015 proposes using the core determining class galichon2011set such as the class of closed intervals to test ((ref)). Alternatively, ((ref)) can be equivalently formulated using Radon-Nikodym densities. We will see the advantage of densities when we construct extensions.
Our main parameter of interest $\theta$ is the local average treatment effect for compliers:
Under the IA-M assumption $A$, the identified set for LATE is characterized by
As shown in Lemma (ref), the IA-M assumption $A$ is refutable. In most empirical applications, researchers do not test this implication, neither do they specify what should be done when the testable implication is rejected. In the next section, I use the relaxed assumption approach to find relaxed assumptions $\tilde{A}$ such that $\tilde{A}$ is non-refutable, find the identified set under $\tilde{A}$, and discuss the estimation and inference on LATE under $\tilde{A}$.
In this section, I will first show that the IA-M assumption have an alternative representation. As discussed in Section (ref), different representations of an assumption can result in different extensions. The canonical representation ((ref)) and the alternative representation will be used to constructed different extensions. I will then look at the maximal relaxation in Definition (ref) and show that the identified set for LATE under the maximal relaxation does not satisfy Property (ref). Then I proceed to construct extensions using the minimal deviation method.
The following is an alternative representation of the IA-M assumption that will be used throughout this section.
To fix the idea, let's consider the case that $A^{ER}$ and $A^{ND}$ holds but we relax the independent instrument assumption. Moreover we consider the extension set $\tilde{A}^{max}\equiv \left(A\cup[{\mathcal{H}_{\mathcal{S}}^{snf}(A)}]^c\right) \cap A^{ER}\cap A^{ND}$. $\tilde{A}^{max}$ is the maximal strong extension defined in Proposition (ref) intersected with the exclusion restriction and the `No Defiers' assumption.
The $\tilde{A}^{max}$ above allows arbitrary dependence of the instrument on the potential outcomes whenever the testable implication ((ref)) fails. First, we should note that the identified set for LATE is very unstable when $F$ satisfies $p(y,1)-q(y,1)=0$ for some $y\in \mathcal{Y}$. Whenever we perturb $F$ slightly such that $p(y,1)-q(y,1)<0$ and ((ref)) fails, the identified set for LATE explodes. Second, the identification set ${LATE}^{ID}_{\tilde{A}^{max}}(F)$ is not any better than the $LATE^{ID}_A(F)$ in equation ((ref)). In terms of interpretation, an uninformative identified set\footnote{Note that the identification result in Proposition (ref) contains only support information when ((ref)) fails.} for LATE is not different from an empty identified set. This is because whenever $F$ fails ((ref)), we give up the `Independent Instrument' assumption. Therefore the remaining assumptions $A^{ER}$ and $A^{ND}$ cannot generate any restrictions on the parameter of interest. As a result, I focus on deriving extensions using minimal deviation method in Definition (ref). I will relax the `No Defiers' and the independent instrument assumption in the following. An extension that relaxes the independent instrument assumption and an extension that relaxes the exclusion restriction are given Appendix (ref).
Recall that $\mathcal{S}$ is a complete structure universe. I consider a strong extension that use measure of defiers as deviation from the no defiers assumption. The extension relaxes the independent instrument to a type independent instrument assumption. I first define the measure of defiers in $G^s$ as
In the above extension, I also relax the independent instrument condition. The type independence assumption is also used in other empirical contexts to study LATE (e.g. see kedagni2019). This is because, by kitagawa2009identification, exclusion restrictions (ER) and instrument condition (IV) has testable implication, so any non-refutable relaxation should relax either ER or IV condition. The second condition in Assumption (ref) requires the measure of always takers (AT) and never takers (NT) to be independent of the instrument.
This section describes the identified set for LATE under $\tilde{A}$ in Assumption (ref). Let $ \mathcal{Y}_d\equiv\{y\in \mathcal{Y}: (-1)^{d}[q(y,d)-p(y,d)]\ge 0\} $ be the collection of $y\in \mathcal{Y}$ such that the density differences are positive.
Assumption (ref) is a regularity assumption. For the identification result, we only need it to hold with $c=0$ so that LATE is well-defined. For inference purpose, I require $c>0$ to avoid weak instrument issue.
In Assumptions (ref), I require that potential outcomes are independent of the instrument conditional on compliers, i.e. $\{Y_i(d,z)\}_{d,z\in\{0,1\}}\perp Z_i\big| (D_i(1)=1,D_i(0)=0)$. As a result, the identified probability of compliers conditioned on $Z_i=1$ is $P(\mathcal{Y}_1,1)-Q(\mathcal{Y}_1,1)$ and the identified probability of compliers conditioned on $Z_i=0$ is $Q(\mathcal{Y}_0,0)-P(\mathcal{Y}_0,0)$.
As we discussed in Section 2, we want the identified set for LATE under the chosen extensions to be a continuous correspondence with respect to the distribution of observables $F$. If the identified set for LATE is discontinuous under the extension, we may get spuriously informative bound for LATE due to sampling error. I compare the identified set for LATE under the maximal extension in Proposition (ref) and the identified set for LATE under the minimal defiers extension in Proposition (ref) in terms of Property (ref). I equip $\mathcal{F}$ with the Sobolev norm: $||F||_{1,\infty}\equiv \max_{i=0,1} ||F^{(i)}||_{\infty}$, where $F^{(i)}$ is the $i$-th Radon-Nikodym density of $F$ with respect to $\mu_F$.
Proposition (ref) shows that the maximal extension is not a good choice if the parameter of interest is LATE.
The identification result in Proposition (ref) relies on the sets $\mathcal{Y}_1,\mathcal{Y}_0$. Throughout this section, I focus on the estimation and inference problem when $Y_i$ is continuously distributed on, and $\mu_F$ is the Lebesgue measure.
To estimate $\hat{\mathcal{Y}}_0,\hat{\mathcal{Y}}_1$, I estimate the density $p(y,d)$ and $q(y,d)$ using kernel density estimators:
The assumption above assumes that the sign of $p(y,d)-q(y,d)$ is known and fixed in the large value of $y$. As a result, we only need to estimate the set $\mathcal{Y}_d\cap [M_l,M_u]$.
Define the upper tail set $\mathcal{Y}_d^{ut}= \mathcal{Y}_d\cap [M_u,\infty)$ and the lower tail set $\mathcal{Y}_d^{lt}= \mathcal{Y}_d\cap (-\infty,M_l]$ and we estimate \[ \hat{\mathcal{Y}}_d(b_n)= \{y\in (M_l,M_u): f_h(y,d)\ge b_n\}\cup \mathcal{Y}_d^{ut}\cup \mathcal{Y}_d^{lt}, \] where $b_n$ is a sequence of positive constants that converges to zero. The estimated set above only uses density $f_h(y,d)$ to distinguish whether $y\in \mathcal{Y}_d$ in the range $(M_l,M_u)$, and uses the known tail sign in Assumption (ref) directly. When the relaxed assumption is defined in Assumption (ref), I construct an estimator of $LATE^{ID}_{\tilde{A}}(F)$ ((ref)) as:
I present the limit distribution of $\widehat{LATE}$ defined in ((ref)). The following assumptions are sufficient to guarantee $\widehat{LATE}$ in ((ref)) will converge to a normal distribution.
The above two assumptions are standard in literature and guarantee the density difference estimator $f_h(y,d)$ will converges uniformly in probability to its limit $(-1)^{1-d}(p(y,d)-q(y,d))$ at polynomial rate.
Assumption (ref) controls the bias from trimming $\{y\in[M_l,M_u]:0<f_h(y,d)<b_n\}$. Essentially, we rule out all outcome distributions such that $\{y:f(y,d)=0\}$ has a positive measure. This assumption is imposed to remove the bias from sampling error in kernel estimator $f_h(y,d)$. Assumption (ref) can be replaced by a sufficient primitive condition.
Theorem (ref) shows that the LATE estimator in ((ref)) is $\sqrt{n}$ consistent. Once the matrices $\Pi$, $\Gamma$ and $\Sigma$ are estimated by consistent estimators, we can test hypothesis such as $H_0:{LATE}_{\tilde{A}}^{ID}(F)=0$. Since LATE is point identified, conventional hypothesis testing method can achieve structural size control and test consistency simultaneously. However, Assumption (ref) requires the econometrician to know the sign of tail behavior of $p(y,1)-q(y,1)$ and $q(y,0)-p(y,0)$. In some empirical application, we may want to be agnostic about tail signs or only impose less restrictive conditions on tail signs. In this case, we can calculate the confidence interval for each possible tail condition, and then take the union, but this confidence interval will be conservative.
This section illustrates the finite sample performance of the proposed inference method. I consider two simulation settings. In the first setting, the IA-M assumption is violated, and the goal of the simulation is to see how the inference method works under the known and unknown tail signs. In the second setting, the IA-M assumption is not violated, and the goal of the simulation is to compare the numerical difference of the 2SLS estimator of ((ref)) and estimator ((ref)).
Instead of simulating the primitive variable $Y_i(d,z),D_i(z)$, I directly simulate the distribution of observed variable such that $Pr(Z_i=1)=0.6$ and \[
\] In this simulation, Assumptions (ref), (ref) and (ref) are satisfied. The trimming band is $[M_l,M_u]=[-2.5,7]$ and Assumption (ref) is satisfied since $\mathcal{Y}^{ut}_1=\mathcal{Y}^{lt}_1=\varnothing$, $\mathcal{Y}^{ut}_0=[M_u,\infty)$, and $\mathcal{Y}^{lt}_0=(-\infty,M_l]$. Assumption (ref) is satisfied since $p(y,1)-q(y,1)=0$ and $q(y,0)-p(y,1)=0$ have two solutions in interval $[M_l,M_u]$ and the derivatives are bounded away from zero (see Figure (ref)).
The true identified value of ${LATE}^{ID}_{\tilde{A}}(F)$ is $1.7385$. Simulation results are given in Table (ref). Coverage probability are calculated from $1000$ replications, and I compare the coverage probability under different sample size $n$ in each replication and the choice of trimming constant $b$. The row of `known tail' in Table (ref) corresponds to the constraints $\mathcal{Y}^{ut}_1=\mathcal{Y}^{lt}_1=\varnothing$, $\mathcal{Y}^{ut}_0=[M_u,\infty)$, and $\mathcal{Y}^{lt}_0=(-\infty,M_l]$ as in the simulation design. The row of `Conservative' in Table (ref) corresponds to the case that the tail set is unknown, and I take the union of confidence intervals under all $16$ possible tail conditions.
The simulation result shows that if we can correctly impose the tail condition as in the known tail case, inference on the true LATE value based on ((ref)) is asymptotically exact but can be sensitive to the choice of trimming sequence $b_n$. If we want to be agnostic about the true tail condition, the union method is conservative.
When the IA-M assumption holds, by vytlacil2002independence, the potential outcome model is equivalent to the latent index model. In this simulation setting, let $U_i\sim U[0,1]$ and $D_i=\mathbbm{1}(0.2+0.6Z_i>U_i)$, where $U[0,1]$ is a uniform distribution on interval $[0,1]$, and $Z_i\sim \mathtt{Bernoulli}(0.5)$ is independent of $U_i$. The potential outcome $(Y_i(1),Y_i(0))\sim N(\mu,\Sigma)$, where $\mu=[2,1.5]$ and $Var(Y_i(1))=2$, $Var(Y_i(0))=1.5$, $Corr(Y_i(1),Y_i(0))=0.7$. We can show $LATE^{ID}(F)=0.5$.
Let $\widehat{LATE}^{Wald}$ denote the Wald estimator of $LATE^{ID}(F)$, and let $\widehat{LATE}^{New}$ denote the estimator in equation ((ref)). Table (ref) shows some summary statistics of the numerical difference between these two estimators over $m=1000$ simulation replications. The first row shows that when the sample size increase, the numerical difference of the two estimator converges to zero in the second moment. The second row in Table (ref) reports the efficiency loss of $\widehat{LATE}^{New}$. We see that at finite sample, not imposing IA-M assumption when it holds can lead to efficiency loss, but the efficiency loss decreases with sample size.
In this section, I apply my results in Proposition (ref) and Theorem (ref) to card1993using, who studied the causal effect of college attendance on earnings. In this application, the outcome variable $Y_i$ is an individual $i$'s log wage in 1976, $D_i=1$ means individual $i$ attended a four-year college, and $Z_i=1$ means the individual was born near a four-year college. This data set has been used by both kitagawa2015 and mourifie2017testing to test the IA-M assumption, and they both reject the IA-M assumption. If a child grew up near a college, he or she may hear more stories of heavy tuition burden, which may discourage him or her from attending college. On the other hand, if this child grew up far away from a college, he or she may instead choose to attend college. Therefore, we would expect defiers to exist in this empirical setting. Moreover, it is unclear why this instrument is fully independent of the potential income, since the choice of residence may depend on parents' potential income, which may be correlated with their children's income.
I conditioned $(Y_i,D_i,Z_i)$ on three characteristics: living in the south (S/NS), living in a metropolitan area (M/NM), and ethnic group (B/NB). I follow mourifie2017testing in excluding subgroup NS/NM/B due to the small sample size, and also exclude subgroup NS/M/B due to the high frequency of $Z=1$. I conduct estimation and inference on each of the remaining 6 subgroups and the pooled sample. The choices of trimming sequence $b_n$, kernel bandwidth $h$, upper and lower band $M_u,M_l$, tail set $\mathcal{Y}_d^{ut},\mathcal{Y}_d^{lt}$ are specified in Appendix (ref).
Estimation results are reported in Table (ref). I also report the LATE estimates when we directly use the IA-M assumption and Wald statistics. The estimated measure of compliers under $\tilde{A}$ satisfying (ref) conditioned on $Z_i=1$ and $Z_i=0$ are reported as $P(\mathcal{Y}_1,1)-Q(\mathcal{Y}_1,1)$ and $Q(\mathcal{Y}_0,0)-P(\mathcal{Y}_0,0)$, while the estimated measure of compliers under IA-M assumption is $E[D_i|Z_i=1]-E[D_i|Z_i=0]$. The estimates of ${LATE}^{ID}_{\tilde{A}}$ and $LATE^{Wald}$ differ the most for three groups: S/NM/NB, S/M/NB and S/M/B. It should be noted that for all these three groups, estimated $E[D_i|Z_i=1]-E[D_i|Z_i=0]$ differs from $P(\mathcal{Y}_1,1)-Q(\mathcal{Y}_1,1)$ and $Q(\mathcal{Y}_0,0)-P(\mathcal{Y}_0,0)$. If we blindly use the identification result under the IA-M assumption and use the standard LATE Wald estimator, the `identified' local average treatment effect can be negative (subgroups S/NM/NB and S/M/NB), or be unrealistically large (subgroup S/M/B). Once we use a strong extension, the estimated LATE for each of the 6 subgroups is positive, and the value of LATE is all between zero and one. We fail to reject education decrease future earning for the complier group ($LATE<0$) for all 6 subgroups for the $LATE^{wald}$. On the other hand, my method can reject the hypothesis $LATE<0$ for the NS/NM/NB and the S/NM/B group at 95% confidence level. When I look at the African-American only, while $LATE^{wald}$ is large, the hypothesis fail to reject that education is harmful to their earning, while my method will reject the hypothesis $LATE^{wald}$ for the African-American is negative.
{
}
In this section, I apply the method to a binary outcome sector choice model with a binary instrument mourifie2018roy. This model is incomplete. Observed outcome $Y_i\in \{0,1\}$ is binary and the observed job sector choice $D_i\in\{0,1\}$ is binary. A binary instrument $Z_i$ is observed. The $\epsilon$ variables include $Y_i(1)$ and $Y_i(0)$, which are the potential outcome in job sector 1 and 0 respectively, and the instrument variable $Z_i$. Observed sector outcome $Y_i$ is generated through:
Without imposing further assumptions, equation ((ref)) does not specify how sector choice $D_i$ is determined, and $D_i$ can take values in one of the three sets $\{0\},\{1\},\{0,1\}$. If we impose the classical Roy sector selection rule, for all $Z_i=z\in\{0,1\}$ we have:
The Roy sector selection rule ((ref)) is just a special case of how $D_i$ is determined. To specify the structure universe $\mathcal{S}$, we consider all possible sector selection rules. \footnote{For each $Y_i(0)=y_0,Y_i(1)=y_1,Z_i=z$, $D_i$ can take values in three sets $\{0\},\{1\},\{0,1\}$. Therefore, there are $3^8$ ways to specify the sector selection rule.}
In Definition (ref), $C^{ykz}_d$ is the probability of choosing sector $d$ when $Y_i(1)=y,Y_i(0)=k,Z_i=z$. When $D^{s,sel}(y,k,z)$ is the set $\{0,1\}$, a structure $s$ associated with $D^{s,sel}$ does not specify how a sector choice is determined, so the only constraint is $C^{ykz}_1+C^{ykz}_0=Pr_{G^s}(Y_i(1)=y,Y_i(1)=k,Z_i=z)$. When $D^{s,sel}(y,k,z)=\{d\}$, the constraint $C^{ykz}_{1-d}=0$ in the last row of ((ref)) requires the probability of choosing sector $1-d$ is zero. Given the set of sector selection rules $\mathcal{D}^{sel}$, we can specify the structure universe $\mathcal{S}$ in this application as follows:
Instead of imposing the strong independent instrument condition $\left(Y_i(1),Y_i(0)\right)\perp Z_i$, we require the instrument to have monotone effects on the potential outcomes:
Definition (ref) only requires the instrument to generate the best (resp. worst) potential outcome $Y_i(1)=Y_i(0)=1$ (resp. $Y_i(1)=Y_i(0)=0$) with higher (resp. lower) probability at $Z_i=1$ than that at $Z_i=0$. This requirement is weaker than Assumption 5 in mourifie2018roy, where they require $Pr(Y_i(d)=1|Z_i=1)\ge Pr(Y_i(d)=1|Z_i=0)$ for $d\in \{0,1\}$ in addition to ((ref)). \footnote{ Condition ((ref)) along with the additional requirement $Pr(Y_i(1)=1|Z_i=1)\ge Pr(Y_i(1)=1|Z_i=0)$ will imply $1-Pr_F(Y_i=0,D_i=1|Z_i=1)\ge Pr_F(Y_i=1,D_i=1|Z_i=0)$ for all outcome distributions $F$. This implication holds even in the absence of the Roy selection condition. On the other hand, equation ((ref)) alone does not imply any constraints on $F$.}
Dominance at the best and worst outcome in Definition (ref) can accommodate broader empirical scenarios compared with Assumption 5 in mourifie2018roy. For example, suppose $Y_i(d)=1$ means individual $i$ gets tenure in sector $d$, and $Z_i=1$ means individual $i$ participates in a job training program. If the skill obtained from the training program can be applied to both sectors, we would expect $Pr(Y_i(1)=Y_i(0)=1|Z_i=1)\ge Pr(Y_i(1)=Y_i(0)=1|Z_i=0)$. On the other hand, each job sector may require specific skill that cannot be obtained from the job training program. If the training program is costly and prevents an individual $i$ from developing skills specific to the sector $d$, it is possible that $Pr(Y_i(d)=1|Z_i=1)\le Pr(Y_i(d)=1|Z_i=0)$. Such a scenario violates Assumption 5 in mourifie2018roy but not ((ref)). Lastly, $Pr(Y_i(1)=Y_i(0)=0|Z_i=1)\le Pr(Y_i(1)=Y_i(0)=0|Z_i=0)$ ensures the training program is beneficial: it increases the probability of success in at least one sector.
However, when ((ref)) is combined with Roy's selection assumption, they are jointly refutable. Formally, the assumption with Roy's selection rule and a monotone instrument satisfying ((ref)) is defined in the following.
The following proposition characterizes the non-refutability and confirmation sets associated with $A^{Roy}$.
Proposition (ref) reveals several things: first $\mathcal{H}_{\mathcal{S}}^{snf}(A^{Roy})\ne \mathcal{S}$, so the assumption $A^{Roy}$ is refutable; second, $\mathcal{H}_{\mathcal{S}}^{wnf}(A^{Roy})\ne \mathcal{H}_{\mathcal{S}}^{snf}(A^{Roy})$ so we cannot find a strong extension of $A^{Roy}$. Therefore, I look at the completion of the binary sector choice model. The completion given in Definition (ref) is abstract and in the following I give an explicit form of $\mathcal{C}(s)$.
Recall that each $s$ in the incomplete space is associated with a sector decision rule, denoted by $D^{s,sel}$. We consider a tie breaking rule associated with $s$: $\left(C^{s^*,tb}_{d}(y,k,z)\right)_{y,k,z,d\in\{0,1\}}$, which specifies the probability of choosing sector $d$ for different values of $Y_i(0)=y,Y_i(1)=k,Z_i=z$. Different values of the vector $\left(C^{s^*,tb}_{d}(y,k,z)\right)_{y,k,z,d\in\{0,1\}}$ correspond to different selections. The completion $\mathcal{C}(s)$ is the collection of $M^{s^*}(\cdot)$ such that {
} Compared with the incomplete mapping in ((ref)), the tie breaking rule $C_{d}^{s^*,tb}$ makes $M^{s^*}(G^{s})$ a singleton. The completed universe is given in Definition (ref) and the corresponding Roy assumption set in the completed structural universe is given in Definition Proposition (ref).
I now consider a strong extension under the completed structure universe $\mathcal{S}^*$. I first define the efficiency loss of a structure $s^*$, which can be viewed as a deviation from Roy's sector selection assumption.
The efficiency loss is a function since $M^{s^*}(G^{s^*})$ is a singleton. It is easy to see that when the Roy sector selection rule holds, $m^{EL}(s^*)=0$. Conversely, by ((ref)), $m^{EL}(s^*)=0$ implies \[ D_i=
\] with probability 1, so the Roy sector selection rule holds. Once we verify $m^{EL}(s^*)$ is a well-behaved minimal deviation measure (see Definition (ref)), we can use minimal efficiency loss to construct a strong extension.
Proposition (ref) characterizes the sharp identified set of distributions of $(Y_i(1),Y_i(0),Z_i)$. The identified set ((ref)) under $\tilde{A}^{*Roy}$ satisfies: (1). There is no efficiency loss when $Z_i=0$ ($C^{010}_1=C^{100}_0=0$); (2). The minimal efficiency loss is $\max\left\{Pr_F(Y_i=0,Z_i=1)-\frac{Pr_F(Y_i=0,Z_i=0)Pr_F(Z_i=1)}{Pr_F(Z_i=0)},0\right\}$; (3). Condition ((ref)) holds as long as $(C^{110}_1+C^{110}_0){Pr_F(Z_i=1)}\le (C^{111}_1+C^{111}_0){Pr_F(Z_i=0)}$ holds. The identified set of $G^{s^*}$ is a polyhedron characterized by the 16-dimensional vector $(C_d^{ykz})_{y,k,z,d\in\{0,1\}}$. Many parameters of interest are linear functions of $(C^{jkz}_d)$, and linear-programming can be used to find the identified set. One example is given in Corollary (ref).
It is worth noticing that if $A^{*Roy}$ cannot be rejected by $F$, the identified set for $\theta\equiv Pr(Y_i(1)=1|Z_i=1)$ is given by $[Pr(Y_i=1,D_i=1|Z_i=1),Pr(Y_i=1|Z_i=1)]$. However, suppose we ignore the testable implication of $A^{*Roy}$ and directly use $[Pr(Y_i=1,D_i=1|Z_i=1),Pr(Y_i=1|Z_i=1)]$ as the identified set for $\theta$, we get a spuriously informative identified set when $m^{EL,min}(F)>0$. This happens when the true structure that generates the data lies in $\tilde{A}^{*Roy}\backslash {A}^{*Roy}$, and the spuriously informative identified set is a proper subset of the true identified set.
The upper bound of the identified set for $Pr(Y_i(1)=1|Z_i=1)$ is not Fr\'echet differentiable with respect to $F$ due to the $\max$ operator. However, by Example 2 in fang2019inference, the upper bound $Pr(Y_i=1|Z_i=1)+ \frac{m^{EL,min}(F)}{Pr(Z_i=1)}$ is directionally differentiable in $F$. The bootstrap method in fang2019inference can be used to construct confidence interval for $Pr(Y_i(1)=1|Z_i=1)$. However, since $Pr(Y_i(1)=1|Z_i=1)$ is partially identified and satisfies the conditions in Proposition (ref), we cannot directly test a hypothesis on the value of $Pr(Y_i(1)=1|Z_i=1)$. Instead, we should test the equivalent existence hypothesis as in Proposition (ref).