EconBase
← Back to paper

Pairwise Valid Instruments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

115,178 characters · 0 sections · 120 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Pairwise Valid Instruments

bibunit\onehalfspacing \begin{abstract} Finding valid instruments is difficult. We propose Validity Set Instrumental Variable (VSIV) estimation, a method for estimating local average treatment effects (LATEs) in heterogeneous causal effect models when the instruments are partially invalid. We consider settings with pairwise valid instruments, that is, instruments that are valid for a subset of instrument value pairs. VSIV estimation exploits testable implications of instrument validity to remove invalid pairs and provides estimates of the LATEs for all remaining pairs, which can be aggregated into a single parameter of interest using researcher-specified weights. We show that the proposed VSIV estimators are asymptotically normal under weak conditions and remove or reduce the asymptotic bias relative to standard LATE estimators (that is, LATE estimators that do not use testable implications to remove invalid variation). We evaluate the finite sample properties of VSIV estimation in application-based simulations and apply our method to estimate the returns to college education using parental education as an instrument. Keywords: Invalid instruments, local average treatment effects, identification, instrumental variable estimation, asymptotic bias reduction \end{abstract} \section{Introduction} Instrumental variable (IV) methods based on the local average treatment effect (LATE) framework imbens1994identification,angrist1995two,angrist1996identification rely on three assumptions:\footnote{See, for example, imbens2014instrumental, melly2017local, and huber2018local for recent reviews, and angrist2008mostly, angrist2014mastering, and imbens2015causal for textbook treatments.} (i) exclusion (the instrument does not have a direct effect on the outcome), (ii) random assignment (the instrument is independent of potential outcomes and treatments), and (iii) monotonicity (the instrument has a monotonic impact on treatment take-up).\footnote{Some papers also include the instrument first-stage assumption as part of the LATE assumptions. We will maintain suitable first-stage assumptions.} In many applications, some of these assumptions are likely to be violated or at least questionable. This has motivated the derivation of testable restrictions and tests for IV validity in various settings balke1997bounds,imbens1997estimating,heckman2005structural,huber2015testing,kitagawa2015test,mourifie2016testing,kedagni2020generalized,carr2021testing,farbmacher2022instrument,frandsen2023judging,jiang2023testing,sun2021ivvalidity.\footnote{There is a related literature on inference with invalid instruments in linear IV models conley2012plausibly,nevo2012identification,armstrong2021sensitivity,goh2022causal.} The main contribution of this paper is to propose a method for exploiting the information available in the testable restrictions of IV validity to remove or reduce the asymptotic bias when estimating LATE parameters.\footnote{We define the asymptotic bias as the probability limit of the $\ell^2$ difference between an estimator and the true value.} We consider settings where the available instruments are partially invalid. A leading example of such a setting is when there is a multivalued instrument for which only some pairs of instrument values satisfy the IV assumptions. In Section (ref), we revisit the analysis of the causal effect of college education on earnings using parental education as an instrument. This instrument is likely partially invalid due to parental education having a positive effect on future earnings, at least up to a certain education level kedagni2020discordant. Another example is the quarter of birth (QOB) instrument of angrist1991does. A potential concern with this instrument is that the seasonality in birth patterns renders the QOB instrument partially invalid bound1995problems,buckles2013season, motivating some studies to only consider a subset of QOBs as instruments dahl2017s. Another empirically relevant setting where partially invalid instruments may arise is when there are multiple instruments.\footnote{Settings with multiple instruments are common in empirical research mogstad2021causal.} In applications with multiple instruments, the validity of a subset of the instruments may be questionable, or the instruments may be partially invalid because the heterogeneity in individual choice behavior renders standard monotonicity assumptions invalid mogstad2021causal. As an example of the latter, consider the study by thornton2008demand, who estimates the causal effect of knowing HIV status on the likelihood of buying condoms using two randomly assigned instruments: Monetary incentives and distance to results centers. In this application, monotonicity is likely to fail due to differences in individual preferences over monetary incentives and distance mogstad2021causal.\footnote{mogstad2021causal propose a weaker version of monotonicity, referred to as partial monotonicity, that they argue is more plausible in this application. We discuss the connection between our assumptions and partial monotonicity in Section (ref). jiang2023testing develop formal tests for partial monotonicity and apply these tests in the context of the thornton2008demand application.} The proposed method, which we refer to as Validity Set IV (VSIV) estimation, uses testable implications of IV validity to remove invalid variation in the instruments and provides LATE estimates based on the remaining variation in the instruments. We establish the asymptotic normality of the proposed VSIV estimators and show that they always remove or reduce the asymptotic bias relative to standard LATE estimators, that is, LATE estimators that do not exploit testable implications of IV validity to remove invalid pairs. Thus, VSIV estimation constitutes a data-driven approach for removing or reducing the asymptotic bias of LATE estimators, given all the information about IV validity available in the data. The use of the testable implications of IV validity in VSIV estimation is more constructive than the standard practice where researchers first test for IV validity, discard the instruments if they reject IV validity, and proceed with standard IV analyses if they do not reject IV validity. VSIV estimation uses the testable implications to remove invalid information in the instruments. Consequently, it can be used to estimate causal effects in settings where the instruments are only partially invalid so that existing tests reject the null of full IV validity.\footnote{See Appendix (ref) for a comparison between VSIV estimation and pairwise pretests based on existing tests for IV validity.} VSIV estimation salvages falsified instruments by exploiting the variation in the instruments not refuted by the data and thereby contributes to the literature on salvaging falsified models masten2021salvaging,kedagni2020discordant. Our goal is to estimate the causal effect of an endogenous treatment $D$ on an outcome of interest $Y$, using a potentially vector-valued discrete instrument $Z$. We consider binary treatments in the main text and multivalued ordered or unordered treatments in the Appendix. In the ideal case, $Z$ is fully valid, that is, the LATE assumptions hold for all instrument values (the instrument is valid for the whole population). However, full IV validity is questionable in many applications, especially when there are many instruments or instrument values. To this end, we introduce the notion of pairwise valid instruments. Pairwise valid instruments are only valid for a subset of all pairs of instrument values, which we refer to as \emph{validity pair set}. Intuitively, the instruments are valid for some subpopulations but invalid for the others. Pairwise validity separates the instrument value pairs into two groups: Valid pairs for which all LATE assumptions hold and invalid pairs for which at least one of the LATE assumptions is violated. Pairwise validity does not require researchers to specify which LATE assumptions are violated for the invalid pairs and to which extent; it allows for failures of exclusion, random assignment, monotonicity, or combinations thereof. As a result, there is no information about the LATE for the invalid instrument value pairs absent additional restrictions (see Appendix (ref)). Pairwise validity is motivated by the fact that it is often difficult to determine why exactly specific instrument pairs are invalid based on contextual knowledge (that is, which combinations of assumptions are violated and how), especially when there are many potentially invalid pairs. If additional information is available on which LATE assumptions fail and how, we can exploit such information for partial identification and sensitivity analysis huber2014sensitivity,noack2021sensitivity,kedagni2023identifying,cui2024robust or focus on target parameters that are identified under relaxations of the LATE assumptions de2017tolerating,frandsen2023judging. Even in settings where such information is available, pairwise validity provides a useful benchmark and starting point. VSIV estimation provides estimates of the LATEs for all pairs of instrument values that satisfy the testable restrictions for IV validity. Specifically, we obtain an estimator, $\widehat{\mathscr{Z}_0}$, of the set of pairs of instrument values that satisfy the testable restrictions in kitagawa2015test, mourifie2016testing, kedagni2020generalized, and sun2021ivvalidity and estimate LATEs for all pairs of instrument values in $\widehat{\mathscr{Z}_0}$. These LATEs can then be aggregated into a single parameter of interest based on user-specified weights. We study the theoretical properties of VSIV estimation under two scenarios. First, we assume that the estimated validity pair set, $\widehat{\mathscr{Z}_0}$, is consistent for the largest validity pair set $\mathscr{Z}_{\bar{M}}$ (that is, the union of all validity pair sets) in the sense that $\mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_{\bar{M}})\to 1$. In this case, VSIV estimation is asymptotically unbiased and normal under standard conditions. Second, since the estimator of the validity pair set, $\widehat{\mathscr{Z}_0}$, is typically constructed based on necessary (but not necessarily sufficient) conditions for IV validity, it could converge to a \emph{pseudo-validity pair set} $\mathscr{Z}_{0}$ that is larger than $\mathscr{Z}_{\bar{M}}$, that is, $\mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_{0})\to 1$.\footnote{kitagawa2015test shows that there exist no sufficient conditions for IV validity when $D$ and $Z$ are both binary.} Let $\mathscr{Z}_P$ be a presumed set of valid pairs of instrument values, incorporating prior information about instrument validity ($\mathscr{Z}_P$ is equal to the set of all pairs if no such prior information is available). We prove that VSIV estimation based on $\widehat{\mathscr{Z}_0}\cap \mathscr{Z}_P$ leads to a smaller asymptotic bias than standard LATE estimators based on $\mathscr{Z}_P$. Taken together, our theoretical results show that, irrespective of whether the largest validity pair set can be estimated consistently or not, VSIV estimation leads to asymptotically normal LATE estimators with reduced asymptotic bias. Finally, we use VSIV estimation to revisit the estimation of the causal effect of college education on earnings using parental education as an instrument. We evaluate the finite sample performance of VSIV estimation in a simulation study calibrated to this application and use these simulations to determine the choice of the tuning parameter required for VSIV estimation. Based on this choice of tuning parameter, VSIV estimation screens out the pairs of instrument values corresponding to low levels of parental education. This is consistent with the above discussion and the findings in kedagni2020discordant in that for low levels of parental education the exclusion restriction may fail. The LATEs for the pairs of instrument values that are not screened out are positive and significant. \paragraph{Notation:} We introduce some standard notation, following sun2021ivvalidity. All random elements are defined on a probability space $(\Omega,\mathcal{A},\mathbb{P})$. For all $m\in\mathbb{N}$, $\mathcal{B}_{\mathbb{R}^m}$ is the Borel $\sigma$-algebra on $\mathbb{R}^m$. We denote by $\mathcal{P}$ the set of probability measures such that if the data $\{(Y_i,D_i,Z_i)\}_{i=1}^n$ are i.i.d.\ and distributed according to some probability measure $Q\in\mathcal{P}$, then $Q(G)=\mathbb{P}((Y_i,D_i,Z_i)\in G)$ for all measurable sets $G$. For every $Q\in\mathcal{P}$ and every measurable function $v$, with some abuse of notation, we define $Q(v)=\int v\, \mathrm{d}Q$. The symbol $\leadsto$ denotes weak convergence in a metric space in the Hoffmann--J\o rgensen sense. For every set $B$, let $1_B$ denote the indicator function for $B$. Finally, to simplify the exposition of the theoretical results, we adopt the convention folland2013real, that \begin{equation} 0\cdot \infty =0. \end{equation} \section{Identification with Pairwise Valid Instruments} \subsection{Weakening Instrument Validity to Pairwise Validity} Consider a setting with an outcome variable $Y\in \mathbb{R}$, a treatment $D\in \mathcal{D}$, and an instrument (vector) $Z\in \mathcal{Z}$. In the main text, we focus on the leading case where the treatment is binary with $\mathcal{D}=\left\{0,1\right\}$. The extensions to multivalued ordered and unordered treatments can be found in the Appendix. The instrument is discrete with $\mathcal{Z}=\left\{z_{1},\ldots,z_K\right\} $, and can be ordered or unordered. Let $Y_{dz}\in\mathbb{R}$ for $(d,z)\in\mathcal{D}\times \mathcal{Z}$ denote the potential outcomes and let $D_{z}$ for $z\in \mathcal{Z}$ denote the potential treatments. The following assumption generalizes the standard LATE assumptions with binary instruments to multivalued instruments. \begin{assumption} IV validity for LATEs with binary treatments and multivalued instruments: \begin{enumerate}[label=(\roman*)] • Exclusion: For each $d\in\{0,1\}$, $Y_{dz_{1}}=\cdots=Y_{dz_{K}}$ almost surely (a.s.). • Random Assignment: $Z$ is jointly independent of $\left( Y_{0z_{1}},\ldots,Y_{0z_K}, Y_{1z_{1} },\ldots,Y_{1z_K}\right)$ and \\$\left(D_{z_{1}},\ldots,D_{z_{K}}\right)$. \label{ass.IV random assignment} \item Monotonicity: For all $ k\in\{1,\ldots, K-1\}$, $D_{z_{k+1}}\geq D_{z_k}$ a.s. \label{ass.IV validity binary D monotonicity} \end{enumerate} \end{assumption} Assumption \ref{ass.IV validity binary D} is similar to the LATE assumptions in, for example, \citet{imbens1994identification}, \citet{angrist1995two}, \citet{frolich2007nonparametric}, \citet{kitagawa2015test}, and \citet{sun2021ivvalidity}. It imposes the IV validity assumptions with respect to all possible values of the instrument $z\in \mathcal{Z}$. This assumption has a lot of identifying power: It identifies LATEs with respect to every pair of IV values $(z_{k},z_{k+1})$ with $\mathbb{P}(D_{z_{k+1}}>D_{z_{k}})>0$. However, Assumption \ref{ass.IV validity binary D} can be restrictive in applications. We therefore introduce the notion of \emph{pairwise instrument validity}, which weakens the conditions in Assumption \ref{ass.IV validity binary D}. Define the set of all possible pairs of values of $Z$ as \begin{align*} \mathscr{Z}=\left\{ \left( z_{1},z_{2}\right) ,\ldots,\left(z_{1},z_{K}\right),\left( z_{2},z_{3}\right) ,\ldots,\left(z_{2},z_{K}\right) ,\ldots,(z_{K-1},z_K),(z_{2},z_1),\ldots,\left( z_{K},z_{K-1}\right) \right\}. \end{align*} The number of the elements in $\mathscr{Z}$ is $K\cdot\left( K-1\right) $. We use $\mathcal{Z}_{\left( k,k^{\prime }\right) }$ to denote a pair $\left( z_{k},z_{k^{\prime}}\right) \in \mathscr{Z}$. Note that we include both $(z_k,z_{k'})$ and $(z_{k'},z_k)$ in $\mathscr{Z}$ so that we do not restrict the direction of the monotonicity (Assumption \ref{ass.IV validity binary D}\ref{ass.IV validity binary D monotonicity}) a priori. \begin{definition} \label{def.partial validity pairwise binary D} The instrument $Z$ is \textbf{pairwise valid} for the treatment $D\in\{0,1\}$ if there is a set $\mathscr{Z}_M=\{(z_{k_1},z_{k_1^{\prime}}),\ldots,(z_{k_M},z_{k_M^{\prime}})\}\subseteq\mathscr{Z}$ such that the following conditions hold for every $(z,z')\in\mathscr{Z}_M$: \begin{enumerate}[label=(\roman*)] \item Exclusion: For each $d\in\{0,1\}$, $Y_{dz}=Y_{dz^{\prime}}$ a.s.\label{def.pairwise exclusion} \item Random Assignment: $Z$ is jointly independent of $(Y_{0z},Y_{0z^{\prime}},Y_{1z},Y_{1z^{\prime}},D_{z},D_{z^{\prime}}) $.\footnote{This condition can be further weakened: The conditional distribution of $(Y_{0z},Y_{0z^{\prime}},Y_{1z},Y_{1z^{\prime}},D_{z},D_{z^{\prime}}) $ given $Z=z$ or $Z=z'$ is the same as the unconditional distribution.} \label{def.pairwise random assignment} \item Monotonicity: $D_{z^{\prime}}\geq D_{z}$ a.s. \label{def.pairwise monotonicity} \end{enumerate} The set $\mathscr{Z}_M$ is called a \textbf{validity pair set} of $Z$.\footnote{We use $\mathscr{Z}_M$ to denote an arbitrary validity pair set throughout the paper. To simplify the notation, we therefore only index $\mathscr{Z}$ by $M$ and not by the full index set $\{(k_1,k_1'),\dots,(k_M,k_M')\}$.} The union of all validity pair sets is the largest validity pair set, denoted by $\mathscr{Z}_{\bar{M}}$. A pair of instrument values $(z,z')$ is called a \textbf{valid pair} if $(z,z')\in\mathscr{Z}_{\bar{M}}$. A pair $(z,z')$ is called an \textbf{invalid pair} if $(z,z')\notin\mathscr{Z}_{\bar{M}}$. \end{definition} Definition \ref{def.partial validity pairwise binary D} separates the instrument value pairs into two groups: Valid pairs for which all the LATE assumptions hold and invalid pairs for which the LATE assumptions fail due to failures of exclusion, independence, monotonicity, or combinations thereof. We show in Appendix \ref{sec.no information} that absent additional restrictions on how Definition \ref{def.partial validity pairwise binary D} can be violated, the sharp identified set for the LATEs for the invalid pairs is the entire real line; that is, there is no information about the LATEs for the invalid pairs in the data. Pairwise validity does not require researchers to specify which LATE assumptions are violated and to which extent for the invalid pairs. The motivation for this is that it can be difficult to determine why exactly instrument pairs are invalid in applications. While pairwise validity does not require researchers to impose additional assumptions for the invalid pairs, it allows for the possibility that valid pairs restrict which LATE assumptions are violated for invalid pairs.\footnote{For example, suppose that $\mathcal{Z}= \{z_1,z_2,z_3\}$, where the pairs $(z_{1},z_3)$ and $(z_{2},z_3)$ are valid. This configuration implies that exclusion (Definition \ref{def.partial validity pairwise binary D}\ref{def.pairwise exclusion}) cannot be violated for the pair $(z_{1},z_2)$.} Definition \ref{def.partial validity pairwise binary D} complements the existing approaches for relaxing the LATE assumptions. These approaches typically impose more structure on which LATE assumptions are violated and how exactly the LATE assumptions are violated. They provide partial identification results and methods for performing sensitivity analyses \citep[e.g.,][]{huber2014sensitivity,noack2021sensitivity,kedagni2023identifying,cui2024robust} or focus on target parameters that are identified under weaker assumptions \citep[e.g.,][]{de2017tolerating,frandsen2023judging}. To illustrate Definition \ref{def.partial validity pairwise binary D}, consider a simple example where $Z\in\mathcal{Z}=\{z_1,z_2,z_3\}$. If $Z$ is fully valid as in Assumption \ref{ass.IV validity binary D} such that $D_{z_3}\ge D_{z_2}\ge D_{z_1}$ a.s., then $\mathscr{Z}_{\bar{M}}=\{(z_1,z_2),(z_1,z_3),(z_2,z_3)\}$. The orange solid lines in Figure \ref{fig:pairwise IV}(a) indicate that two instrument values, $\{z_k,z_{k'}\}$, form a valid pair: Either $(z_k,z_{k'})$ or $(z_{k'},z_k)$ satisfies the conditions in Definition \ref{def.partial validity pairwise binary D}. The full validity Assumption \ref{ass.IV validity binary D} requires that every pair of instrument values forms a valid pair. Definition \ref{def.partial validity pairwise binary D} relaxes Assumption \ref{ass.IV validity binary D} as it does not require every pair to form a valid pair. For example, it could be that only $(z_1,z_3)$ satisfies the conditions in Definition \ref{def.partial validity pairwise binary D}. The teal dashed lines in Figure \ref{fig:pairwise IV}(b) indicate that $\{z_1,z_2\}$ and $\{z_2,z_3\}$ do not form valid pairs. In this case, the instrument $Z$ is pairwise but not fully valid. \begin{figure} [H] \caption{Full IV Validity vs.\ Pairwise IV Validity} \label{fig:pairwise IV} \centering \begin{subfigure}[b]{0.4\textwidth} \centering \begin{tikzpicture} \draw [orange, very thick, -] (0,0) -- (1,1.732); \draw [orange, very thick, -] (2,0) -- (1,1.732); \draw [orange, very thick, -] (0,0) -- (2,0); \node[circle, draw=gray!70, fill=gray!10, very thick, minimum size=7mm] at (0,0) {$z_2$}; \node[circle, draw=gray!70, fill=gray!10, very thick, minimum size=7mm] at (1,1.732) {$z_1$}; \node[circle, draw=gray!70, fill=gray!10, very thick, minimum size=7mm] at (2,0) {$z_3$}; \end{tikzpicture} \subcaption{Fully Valid Instrument $Z$} \end{subfigure} \begin{subfigure}[b]{0.4\textwidth} \centering \begin{tikzpicture} \draw [teal, very thick, dashed,-] (0,0) -- (1,1.732); \draw [orange, very thick, -] (2,0) -- (1,1.732); \draw [teal, very thick, dashed, -] (0,0) -- (2,0); \node[circle, draw=gray!70, fill=gray!10, very thick, minimum size=7mm] at (0,0) {$z_2$}; \node[circle, draw=gray!70, fill=gray!10, very thick, minimum size=7mm] at (1,1.732) {$z_1$}; \node[circle, draw=gray!70, fill=gray!10, very thick, minimum size=7mm] at (2,0) {$z_3$}; \end{tikzpicture} \subcaption{Pairwise Valid Instrument $Z$} \end{subfigure} \end{figure} In applications where the instrument $Z$ is randomly assigned (e.g., in experiments with imperfect compliance), joint independence (Assumption \ref{ass.IV validity binary D}\ref{ass.IV random assignment}) holds by design. In such applications, Definition \ref{def.partial validity pairwise binary D} captures violations of exclusion and monotonicity. Such violations are easy to interpret. The pairwise exclusion assumption in Definition \ref{def.partial validity pairwise binary D}\ref{def.pairwise exclusion} requires that $Y_{dz}$, viewed as a function of $z$, is constant over some regions of $\mathcal{Z}$ and varies over others. This nests, for example, Condition E.3 in \citet{kedagni2020discordant}, which requires that $Y_{dt}\le Y_{dt'}$ for all $t\le t'$ and $Y_{dt}=Y_{dz}$ for all $t\ge z$. The pairwise monotonicity assumption (Definition \ref{def.partial validity pairwise binary D}\ref{def.pairwise monotonicity}) requires that $D_z$, viewed as a function of $z$, is monotonic over some regions of $\mathcal{Z}$ and non-monotonic over others. We discuss the relationship to existing relaxations of LATE monotonicity in more detail in Section \ref{sec.partial monotnicity}. In many quasi-experimental applications, the instrument $Z$ is not randomly assigned, and joint independence (Assumption \ref{ass.IV validity binary D}\ref{ass.IV random assignment}) may fail. A leading and practically relevant case where joint independence fails but pairwise independence (Definition \ref{def.partial validity pairwise binary D}\ref{def.pairwise random assignment}) holds is when there are multiple instruments, and some of them are not independent of all potential variables.\footnote{It is possible that joint independence fails but pairwise independence holds even if there is only one original instrument. To illustrate, let $D$ indicate college enrollment, and let $Z\in \{1,2,3\}$ measure distance to the closest college \citep[e.g.,][]{kane1993labor}, where $Z=1$ indicates \texttt{close}, $Z=2$ indicates \texttt{far}, and $Z=3$ indicates \texttt{very far}. Consider the selection mechanism $D_z=1\left\{B_0+B_11\{z=1\}+f(z)1\{z>1\}\le 0\right\}$, where $B_0$ and $B_1$ are random coefficients, $f(2)<f(3)$, and $B_0\perp\!\!\!\perp Z$. The coefficient $B_1$ captures the taste for living \texttt{close} to college relative to living farther away. If $B_1$ is correlated with actual distance $Z$, then $Z\not\!\perp\!\!\!\perp D_1$ and $Z\perp\!\!\!\perp (D_2,D_3)$.} To illustrate, consider the following example based on \citet[][Section II.C]{mogstad2021causal} and the empirical application in \citet{carneiro2011estimating}. Let $D$ be an indicator for college attendance. There are two binary instruments, $Z=(Z_1,Z_2)$, where $Z_1$ is an indicator for college proximity \citep[e.g.,][]{card1993geographic,kane1993labor} and $Z_2$ is an indicator for tuition subsidy.\footnote{We swap the order of the instruments relative to \citet[][Section II.C]{mogstad2021causal} for the purpose of illustration.} Individuals decide whether to attend college based on the following selection mechanism, \begin{equation} D_z=1\left\{B_0+B_1z_1+z_2\ge 0\right\}, \end{equation} where $B_0$ and $B_1$ are random coefficients. For simplicity, we assume that $Z\perp\!\!\!\perp B_0$. The coefficient $B_1$ measures the ``taste'' for proximity relative to tuition subsidy. If the taste for proximity, $B_1$, is correlated with actual proximity, $Z_1$, for example, due to spatial sorting, then the pair $(D_{(0,0)},D_{(0,1)})$ is independent of $Z$ but the pair $(D_{(1,0)},D_{(1,1)})$ is not.\footnote{See, for example, \citet{card1993geographic}, \citet{carneiro2011estimating}, \citet{slichter2014testing}, and \citet{kitagawa2015test} for discussions of the validity of the college proximity instrument.} \begin{remark}[Weakening Definition \ref{def.partial validity pairwise binary D} with Multiple Instruments] In Appendix \ref{sec.selectively valid instruments}, we introduce a weaker notion of pairwise validity (Definition \ref{def.partial validity pairwise binary D}) for settings where $Z$ contains multiple instruments: $Z=(Z_1,\ldots,Z_L)^T$, where $Z_l$ is a scalar instrument for $l\in\{1,\ldots,L\}$. \end{remark} \subsection{Relationship to Other Variants and Relaxations of Monotonicity} \label{sec.partial monotnicity} Here we discuss the connection between our pairwise monotonicity assumption (Definition \ref{def.partial validity pairwise binary D}\ref{def.pairwise monotonicity}) and three recently proposed variants and relaxations of the LATE monotonicity assumption. First, \citet{mogstad2021causal} propose a partial monotonicity (PM) condition for settings with multiple instruments, which is a special case of Condition \ref{def.pairwise monotonicity} in Definition \ref{def.partial validity pairwise binary D}; see also \citet{goff2020vector} for vector monotonicity assumption.\footnote{\citet{mogstad2021causal} motivate the PM condition by showing that full monotonicity imposes strong restrictions on the heterogeneity in individual choice behavior and is therefore likely violated in many applications.} For example, suppose that $Z=(Z_1,Z_2)\in\mathbb{R}^2$ and each element of $Z$ is binary so that $\mathcal{Z}=\{(0,0),(0,1),(1,0),(1,1)\}$. Suppose that Assumption PM of \citet{mogstad2021causal} holds with $D_{(0,0)}\ge D_{(0,1)}$, $D_{(0,0)}\ge D_{(1,0)}$, $D_{(1,1)}\ge D_{(0,1)}$, and $D_{(1,1)}\ge D_{(1,0)}$ a.s. (the sex composition instrument in \citet{angrist1998children} discussed in \citet{mogstad2021causal}), and that Conditions \ref{def.pairwise exclusion} and \ref{def.pairwise random assignment} of Definition \ref{def.partial validity pairwise binary D} hold. Then a validity pair set is $$\{((0,1),(0,0)),((1,0),(0,0)),((0,1),(1,1)),((1,0),(1,1))\}.$$ Second, \citet[][Section IV]{frandsen2023judging} study the interpretation of 2SLS under relaxations of monotonicity and exclusion. The relaxation of monotonicity, referred to as average monotonicity, requires $D_z$ to be positively correlated with the instrument propensity. Average monotonicity is fundamentally different from pairwise monotonicity (Definition \ref{def.partial validity pairwise binary D}\ref{def.pairwise monotonicity}). Pairwise montonicity operates at the level of pairs of instrument values, whereas average monotonicity implies restrictions across all instrument values. Also, \citet{frandsen2023judging} show that average monotonicity can be used to identify averages of treatment effects. By contrast, pairwise validity identifies the LATE for $(z,z')\in\mathscr{Z}_{\bar{M}}$, but does not identify the LATE for $(z,z')\notin \mathscr{Z}_{\bar{M}}$ (Corollary \ref{cor:no_information} in Appendix \ref{sec.no information}). Finally, \citet{noack2021sensitivity} considers a continuous relaxation of monotonicity when $Z$ is binary, parameterized by the fraction of defiers. We do not consider continuous relaxations and make no assumptions on the degree of violation. Combining VSIV estimation with continuous relaxations as in \citet{noack2021sensitivity} is an interesting direction for future research, as we discuss in Section \ref{sec:conclusion}. \subsection{Identification under Pairwise Validity} The following lemma establishes identification under pairwise validity. \begin{lemma}\label{lemma.pairwise beta binary D} Suppose that the instrument $Z$ is pairwise valid according to Definition \ref{def.partial validity pairwise binary D} with a known validity pair set $\mathscr{Z}_M=\{(z_{k_1},z_{k_1^{\prime}}),\ldots,(z_{k_M},z_{k_M^{\prime}})\}$.\footnote{Note that mathematically we do not need to impose a first-stage assumption here due to the convention \eqref{eq.0timesinfinity}.} Then we can define a random variable $Y_d(z_{k_m},z_{k_m^{\prime}})=Y_{dz_{k_m}}=Y_{dz_{k'_m}}$ a.s. for each $d\in\{0,1\}$ and every $(z_{k_m},z_{k_m^{\prime}})\in\mathscr{Z}_M$, and the following quantity can be identified for every $(z_{k_m},z_{k_m^{\prime}})\in\mathscr{Z}_M$: \begin{align}\label{eq.beta binary D} \beta _{k_{m}^{\prime},k_{m}}&\equiv E\left[Y_{1}(z_{k_m},z_{k_m^{\prime}})-Y_{0}(z_{k_m},z_{k_m^{\prime}})\big|D_{z_{k_m'}}>D_{z_{k_m}}\right]\notag \\ &=\frac{E\left[ Y|Z=z_{k_{m}^{\prime}}\right] -E\left[ Y|Z=z_{k_{m}}\right] }{E\left[ D|Z=z_{k_{m}^{\prime}}\right] -E\left[ D|Z=z_{k_{m}}\right] }. \end{align} \end{lemma} Lemma \ref{lemma.pairwise beta binary D} is a direct extension of Theorem 1 of \citet{imbens1994identification} for the case where $Z$ is pairwise valid. We follow \citet{imbens1994identification} and refer to $\beta_{k_{m}^{\prime},k_m}$ as a LATE. Lemma \ref{lemma.pairwise beta binary D} shows that if a validity pair set $\mathscr{Z}_{{M}}$ is known, we can identify every $\beta_{k_{m}^{\prime},k_m}$ with $(z_{k_m},z_{k_m^{\prime}})\in\mathscr{Z}_M$.\footnote{Note that if $(z_{k_m},z_{k_m^{\prime}})\in\mathscr{Z}_M$ with $D_{z_{k_m}}=D_{z_{k'_m}}$ a.s., then $\beta_{k_{m}^{\prime},k_m}=0$ by \eqref{eq.0timesinfinity}. Moreover, if $(z_{k_m},z_{k_m^{\prime}})\in\mathscr{Z}_{{M}}$ and $(z_{k'_m},z_{k_m})\in\mathscr{Z}_{{M}}$, then by Definition \ref{def.partial validity pairwise binary D}, $D_{z_{k_m}}=D_{z_{k'_m}}$ a.s.} In practice, however, $\mathscr{Z}_M$ is usually unknown. In this paper, we use testable implications of IV validity to estimate a pseudo-validity pair set $\mathscr{Z}_{0}$ containing $\mathscr{Z}_M$, and show how to use this estimated set to reduce the asymptotic bias in LATE estimation. We focus on the vector of LATEs $\{\beta_{k_{m}^{\prime},k_m}\}$ as our object of interest. Traditional IV estimators estimate weighted averages of LATEs \citep[e.g.,][]{imbens1994identification} and, thus, are strictly less informative (we can always compute linear IV estimands based on the LATEs). Moreover, VSIV estimation estimates LATEs that do not enter such weighted averages (Theorem 2 of \citet{imbens1994identification}). To illustrate, suppose $\mathcal{Z}=\{z_1,z_2,z_3\}$ and $\mathscr{Z}_{\bar{M}}=\{(z_1,z_2),(z_1,z_3),(z_2,z_3)\}$. The traditional IV estimator estimates a weighted average of $\beta_{2,1}$ and $\beta_{3,2}$, whereas our method estimates $(\beta_{2,1},\beta_{3,1},\beta_{3,2})^T$. Importantly, VSIV estimation allows for assigning researcher-specified weights to $\{\beta_{k_{m}^{\prime},k_m}\}$. In the traditional IV estimation, the weights are determined by the estimation procedure. \citet{mogstad2021causal} show that the weights assigned to the LATEs by 2SLS could be negative under partial monotonicity. Negative weights are not an issue for VSIV estimation because the weights can be chosen by researchers, instead of being determined by the estimation procedure. For example, we may define the weighted average as \begin{align}\label{eq.weighted average} \beta_w=\frac{p_{12}}{p_{123}}\beta_{2,1}+\frac{p_{13}}{p_{123}}\beta_{3,1}+\frac{p_{23}}{p_{123}}\beta_{3,2}, \end{align} where $p_{ij}=\mathbb{P}(Z\in\{z_i,z_j\})$ and $p_{123}=\mathbb{P}(Z\in\{z_1,z_2\})+\mathbb{P}(Z\in\{z_2,z_3\})+\mathbb{P}(Z\in\{z_1,z_3\})$. The asymptotic properties of the estimated weighted averages of LATEs follow straightforwardly from the asymptotic theory in Section \ref{sec: validity set iv estimation}. See Corollary \ref{corollary.weighted average weak convergence}. \begin{remark}[Extrapolation] The focus of VSIV estimation is on estimating the LATE parameters $\beta _{k^{\prime},k}$ for all pairs of instrument values $(z_{k},z_{k'})$ satisfying the testable restrictions of IV validity. This is because absent additional restrictions, there is no information in the data about the LATE for the invalid pairs (see Appendix \ref{sec.no information}). The local and DGP-dependent nature of LATE parameters has motivated the development of a variety of methods for assessing and restoring external validity \citep[e.g.,][]{heckman2003simple,angrist2013extrapolate,brinch2017beyond,mogstad2018using,wuthrich2020comparison,kowalski2023how}. The use of invalid instrument pairs will result in these approaches being biased and inconsistent. VSIV estimation constitutes a natural complement to the existing approaches to external validity. For settings where researchers are interested in externally valid effects, we recommend a two-step procedure: (i) Use VSIV estimation to eliminate invalid pairs. (ii) Apply a suitable approach to external validity based on the estimated validity pair set. In step (i), we recommend also reporting the VSIV LATE estimates because they summarize the available information about pairwise LATEs and are important inputs for approaches to external validity. \end{remark} \section{Validity Set IV Estimation} \label{sec: validity set iv estimation} \subsection{Overview} The goal of VSIV estimation is to exclude invalid instrument pairs. Specifically, we seek to exclude $(z_k,z_{k'})\notin{\mathscr{Z}}_{\bar{M}}$ from $\mathscr{Z}$, since if $(z_k,z_{k'})\notin{\mathscr{Z}}_{\bar{M}}$, then $\beta_{k',k}$ in \eqref{eq.beta binary D} is not identified absent additional restrictions (Appendix \ref{sec.no information}). Suppose that there is a set $\mathscr{Z}_0\subseteq\mathscr{Z}$ that satisfies the testable implications in \citet{kitagawa2015test}, \citet{mourifie2016testing}, \citet{kedagni2020generalized}, and \citet{sun2021ivvalidity}. Then we construct an estimator $\widehat{\mathscr{Z}_0}$ for $\mathscr{Z}_0$ and construct IV estimators based on $(z_{k},z_{k^{\prime}})\in \widehat{\mathscr{Z}_0}$. We refer to these estimators as \emph{VSIV estimators}. In the following, we assume that we have access to an estimator $\widehat{\mathscr{Z}_0}$, which is consistent for $\mathscr{Z}_{0}$ in the sense that $\mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_{0})\to 1$. We describe the testable implications and the construction of the proposed estimator satisfying $\mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_{0})\to 1$ in detail in Section \ref{sec.estimation validity set}. In Section \ref{sec.VSIV under consistency}, we study VSIV estimation under the assumption that $\mathscr{Z}_0=\mathscr{Z}_{\bar{M}}$ so that $\mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_{\bar{M}})\to 1$. In this case, the proposed VSIV estimators are asymptotically unbiased and normal under standard weak regularity conditions. Since ${\mathscr{Z}_0}$ is constructed based on the necessary (but not necessarily sufficient) conditions for the pairwise IV validity, ${\mathscr{Z}_0}$ could be larger than ${\mathscr{Z}}_{\bar{M}}$. (There exist no sufficient testable conditions for IV validity in general \citep{kitagawa2015test}.) In Section \ref{sec.bias reduction binary D}, we show that even if ${\mathscr{Z}_0}$ is larger than ${\mathscr{Z}}_{\bar{M}}$, VSIV estimators yield asymptotic bias reductions relative to standard LATE estimators that do not exploit testable implications to remove invalid instrument value pairs. Note that if $\mathscr{Z}_0=\varnothing$, VSIV estimation is trivial asymptotically since $\mathbb{P}(\widehat{\mathscr{Z}_0}=\varnothing)\to 1$. All the VSIV estimators converge to $0$ by the convention in \eqref{eq.0timesinfinity}. In this case, we do not report any IV estimates in practice. \subsection{VSIV Estimation under Consistent Estimation of Validity Pair Set} \label{sec.VSIV under consistency} Suppose that $\mathscr{Z}_0=\mathscr{Z}_{\bar{M}}$ so that $\mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_{\bar{M}})\to 1$, and we use $\widehat{\mathscr{Z}_{0}}$ to construct VSIV estimators for the LATEs. We impose the following standard regularity conditions. Let $g$ be a prespecified function that maps the value of $Z$ to $\mathbb{R}$. For example, we can simply set $g(z)=z$ for all $z$ if $Z$ is a scalar instrument.\footnote{The choice of $g$ may affect the efficiency of the VSIV estimators. We leave the formal analysis of the optimal choice of $g$ for future study.} \begin{assumption}\label{ass.iid data binary D} $\{(Y_i,D_i,Z_i)\}_{i=1}^{n}$ is an i.i.d.\ sample from a population such that all relevant moments exist. \end{assumption} \begin{assumption}\label{ass.first stage binary D} For every $\mathcal{Z}_{(k,k')}\in\mathscr{Z}_{\bar{M}}$, \begin{align}\label{eq.first stage} E[g(Z_i)D_i|Z_i\in\mathcal{Z}_{(k,k^{\prime})}]-E[D_i|Z_i\in\mathcal{Z}_{(k,k^{\prime})}]\cdot E[g(Z_i)|Z_i\in\mathcal{Z}_{(k,k^{\prime})}]\neq0. \end{align} \end{assumption} Assumption \ref{ass.iid data binary D} assumes an i.i.d.\ data set and requires the existence of the relevant moments. Assumption \ref{ass.first stage binary D} imposes a first-stage condition for every $\mathcal{Z}_{(k,k')}\in\mathscr{Z}_{\bar{M}}$. Note that \eqref{eq.first stage} may not hold for $\mathcal{Z}_{(k,k')}\notin\mathscr{Z}_{\bar{M}}$. This creates additional technical difficulties when establishing the asymptotic normality of the VSIV estimators, which we discuss below. Assumption \ref{ass.first stage binary D} also implies that if $\mathcal{Z}_{(k,k')}\in\mathscr{Z}_{\bar{M}}$, then $\mathcal{Z}_{(k',k)}\notin\mathscr{Z}_{\bar{M}}$. Otherwise, by Definition \ref{def.partial validity pairwise binary D}, $D_{z_k}=D_{z_{k'}}$ a.s., and \eqref{eq.first stage} does not hold. For every scalar random sample $\{\xi_i\}_{i=1}^n$ and every $\mathcal{A}\in\mathscr{Z}$, we define \begin{align*} \mathcal{E}_{n}\left( \xi_{i},\mathcal{A}\right) =\frac{\frac{1}{n}\sum _{i=1}^{n}\xi_{i}1\left\{ Z_{i}\in\mathcal{A}\right\}}{\frac{1}{n}\sum_{i=1}^{n} 1\left\{ Z_{i}\in\mathcal{A}\right\} } \text{ and }\mathcal{E}\left( \xi_{i},\mathcal{A}\right) =\frac{E\left[ \xi_{i}1\left\{ Z_{i}\in\mathcal{A}\right\} \right]}{E\left[ 1\left\{ Z_{i}\in\mathcal{A}\right\} \right]}. \end{align*} We define the VSIV estimators using regression-based IV estimators, following \citet{imbens1994identification}. For every $\mathcal{Z}_{( k,k^{\prime}) }\in{\mathscr{Z}}$, we run the IV regression \begin{align}\label{eq.VSIV estimation pairwise binary D} Y_i 1\left\{ Z_{i}\in\mathcal{Z}_{(k,k^{\prime})}\right\} =&\,\gamma_{(k,k^{\prime})}^01\left\{Z_{i}\in\mathcal{Z}_{(k,k^{\prime})}\right\}+\gamma_{(k,k^{\prime})}^1D_i1\left\{Z_{i}\in\mathcal{Z}_{(k,k^{\prime})}\right\}+\epsilon_{i}1\left\{Z_{i}\in\mathcal{Z}_{(k,k^{\prime})}\right\}, \end{align} using $g(Z_i)1\{Z_{i}\in\mathcal{Z}_{(k,k^{\prime})}\}$ as the instrument for $D_i1\{Z_{i}\in\mathcal{Z}_{(k,k^{\prime})}\}$. Given the estimated validity set $\widehat{\mathscr{Z}_0}$, we set the VSIV estimator for each $\mathcal{Z}_{(k,k^{\prime})}$ as \begin{align}\label{eq.VSIV estimator pairwise binary D} \widehat{\beta}_{( k,k^{\prime}) }^1=1\left\{\mathcal{Z}_{(k,k^{\prime})}\in\widehat{\mathscr{Z}_0}\right\}\cdot\frac{\mathcal{E}_{n}\left( g\left( Z_{i}\right) Y_{i},\mathcal{Z}_{(k,k^{\prime})}\right) -\mathcal{E}_{n}\left( g\left( Z_{i}\right) ,\mathcal{Z}_{(k,k^{\prime} )}\right) \mathcal{E}_{n}\left( Y_{i},\mathcal{Z}_{(k,k^{\prime})}\right) }{\mathcal{E}_{n}\left( g\left( Z_{i}\right) D_{i},\mathcal{Z} _{(k,k^{\prime})}\right) -\mathcal{E}_{n}\left( g\left( Z_{i}\right) ,\mathcal{Z}_{(k,k^{\prime})}\right) \mathcal{E}_{n}\left( D_{i} ,\mathcal{Z}_{(k,k^{\prime})}\right) }, \end{align} which is the IV estimator of $\gamma_{(k,k^{\prime})}^1$ in \eqref{eq.VSIV estimation pairwise binary D} multiplied by $1\{\mathcal{Z}_{(k,k^{\prime})}\in\widehat{\mathscr{Z}_0}\}$. For every $\mathcal{Z}_{(k,k')}\in\widehat{\mathscr{Z}}_0$, this IV estimation is equivalent to a conventional IV regression in the subsample of $\{(Y_i,D_i,Z_i)\}_{i=1}^n$ with $Z_i\in\mathcal{Z}_{(k,k')}$. Note that $\widehat{\beta}_{( k,k^{\prime}) }^1=0$ if $\mathcal{Z}_{(k,k^{\prime})}\notin\widehat{\mathscr{Z}_0}$. We discuss this convention further below. Define the vector of VSIV estimators as \[ \widehat{\beta}_{1}=\left( \widehat{\beta}_{\left( 1,2\right) }^{1},\ldots ,\widehat{\beta}_{\left( 1,K\right) }^{1},\ldots,\widehat{\beta}_{\left( K,1\right) }^{1},\ldots,\widehat{\beta}_{\left( K,K-1\right) }^{1}\right)^T. \] We also define \begin{align}\label{eq.VSIV true beta binary D} \beta_{( k,k^{\prime}) }^{1}=1\left\{\mathcal{Z}_{(k,k^{\prime})}\in\mathscr{Z}_{\bar{M}}\right\}\cdot\frac{\mathcal{E}\left( g\left( Z_{i}\right) Y_{i},\mathcal{Z}_{(k,k^{\prime})}\right) -\mathcal{E}\left( g\left( Z_{i}\right) ,\mathcal{Z}_{(k,k^{\prime})}\right) \mathcal{E} \left( Y_{i},\mathcal{Z}_{(k,k^{\prime})}\right) }{\mathcal{E}\left( g\left( Z_{i}\right) D_{i},\mathcal{Z}_{(k,k^{\prime})}\right) -\mathcal{E}\left( g\left( Z_{i}\right) ,\mathcal{Z}_{(k,k^{\prime} )}\right) \mathcal{E}\left( D_{i},\mathcal{Z}_{(k,k^{\prime})}\right) } \end{align} and \begin{align}\label{eq.VSIV true beta1 binary D} \beta_{1}=\left( \beta_{\left( 1,2\right) }^{1},\ldots,\beta_{\left( 1,K\right) }^{1},\ldots,\beta_{\left( K,1\right) }^{1},\ldots ,\beta_{\left( K,K-1\right) }^{1}\right)^T . \end{align} As we show formally in Theorem \ref{thm.IV estimator asymptotics pairwise binary D} below, $\beta_{( k,k^{\prime}) }^{1}=\beta_{k^{\prime},k}$ as defined in \eqref{eq.beta binary D} for every $(z_k,z_{k^{\prime}})\in\mathscr{Z}_{\bar{M}}$. If $\mathcal{Z}_{(k,k^{\prime})}\notin {\mathscr{Z}_{\bar{M}}}$, we set ${\beta}_{(k,k^{\prime})}^1=0$ by \eqref{eq.VSIV true beta binary D} and \eqref{eq.0timesinfinity}. Similarly, if $\mathcal{Z}_{(k,k^{\prime})}\notin \widehat{\mathscr{Z}_0}$, $\widehat{\beta}_{(k,k^{\prime})}^1=0$ by \eqref{eq.VSIV estimator pairwise binary D} and \eqref{eq.0timesinfinity}. Letting them be equal to $0$ facilitates the description of the theoretical results in Theorem \ref{thm.IV estimator asymptotics pairwise binary D}, and this will not affect the estimation of the weighted average of LATEs.\footnote{It is equivalent to not including them into the averages.} In practice, we recommend leaving $\widehat{\beta}_{(k,k^{\prime})}^1$ to be blank if $1\{\mathcal{Z}_{(k,k^{\prime})}\in\widehat{\mathscr{Z}_0}\}=0$, as in the application in Section \ref{sec.application}. We interpret the LATE corresponding to $\mathcal{Z}_{(k,k^{\prime})}\in {\mathscr{Z}_{\bar{M}}}$ in the usual way as the average treatment effects for compliers in the corresponding subgroup. We do not report the estimates for LATEs corresponding to invalid pairs since they are not identified and there is no information about them in the data absent additional restrictions (Appendix \ref{sec.no information}). The next theorem establishes the asymptotic distribution of the VSIV estimator $\widehat{\beta}_{1}$, obtained based on the estimator of the instrument validity pair set $\widehat{\mathscr{Z}_0}$. \begin{theorem}\label{thm.IV estimator asymptotics pairwise binary D} Suppose that the instrument $Z$ is pairwise valid for the treatment $D$ according to Definition \ref{def.partial validity pairwise binary D} with the largest validity pair set $\mathscr{Z}_{\bar{M}}=\{(z_{k_1},z_{k_1^{\prime}}),\ldots,(z_{k_{\bar{M}}},z_{k_{\bar{M}}^{\prime}})\}$, that the estimator $\widehat{\mathscr{Z}_0}$ satisfies $\mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_{\bar{M}})\to 1$, and that Assumptions \ref{ass.iid data binary D} and \ref{ass.first stage binary D} hold. Then \begin{align}\label{eq.beta asymptotic distribution} \sqrt{n}( \widehat{\beta}_{1}-\beta_1 ) \overset{d}\to N\left( 0,\Sigma \right), \end{align} where $\Sigma$ is defined in \eqref{eq.weak convergence pairwise2} in the Appendix. In addition, $\beta_{( k,k^{\prime}) }^{1}=\beta_{k^{\prime},k}$ as defined in \eqref{eq.beta binary D} for every $(z_k,z_{k^{\prime}})\in\mathscr{Z}_{\bar{M}}$. \end{theorem} Theorem \ref{thm.IV estimator asymptotics pairwise binary D} establishes the joint asymptotic normality of the VSIV estimator of the LATEs. Establishing the asymptotic distribution in \eqref{eq.beta asymptotic distribution} requires a careful treatment of the case where the first-stage Assumption \ref{ass.first stage binary D} does not hold for some pairs of instrument values $\mathcal{Z}_{(k,k')}$ that are not in the largest validity pair set $\mathscr{Z}_{\bar{M}}$, that is, $\mathcal{Z}_{(k,k')}\notin\mathscr{Z}_{\bar{M}}$. Specifically, we show that in this case, $\mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_{\bar{M}})\to 1$ implies that, if $\mathcal{Z}_{(k,k')}\notin\mathscr{Z}_{\bar{M}}$, then for every $\rho>0$, $n^{\rho}1\{\mathcal{Z}_{(k,k')}\in\widehat{\mathscr{Z}_0}\}=o_p(1)$. This guarantees the convergence in \eqref{eq.beta asymptotic distribution} even when \eqref{eq.first stage} does not hold for $\mathcal{Z}_{(k,k')}\notin\mathscr{Z}_{\bar{M}}$. The asymptotic covariance matrix $\Sigma$ defined in the Appendix can be consistently estimated under standard conditions. Importantly, the estimation of the instrument validity pair set does not affect the asymptotic covariance matrix such that standard inference methods can be applied. Under Theorem \ref{thm.IV estimator asymptotics pairwise binary D}, it is straightforward to obtain a consistent estimator of a weighted average of all ${\beta}_{(k,k')}^1$. Let $\widehat{W}$ be some estimated weights with \begin{align*} \widehat{W}=(\widehat{w}_{(1,2)},\ldots,\widehat{w}_{(1,K)},\ldots,\widehat{w}_{(K,1)},\ldots,\widehat{w}_{(K,K-1)})^T \end{align*} such that $\widehat{W}\overset{p}\rightarrow W$ for some $W$ with \begin{align*} {W}=({w}_{(1,2)},\ldots,{w}_{(1,K)},\ldots,{w}_{(K,1)},\ldots,{w}_{(K,K-1)})^T. \end{align*} Theorem \ref{thm.IV estimator asymptotics pairwise binary D} implies that $\widehat{W}^T\widehat{\beta}_1\overset{p}\rightarrow W^T{\beta}_1$. To establish the asymptotic distribution of the estimated weighted average of all $\beta^1_{(k,k')}$, we assume that $(\widehat{W}^T,\widehat{\beta}_1^T)^T$ is asymptotically normal. This is a weak condition in practice. It holds, for example, if the weights are defined as in the weighted average in \eqref{eq.weighted average}. The following corollary summarizes the result. \begin{corollary}\label{corollary.weighted average weak convergence} Suppose $\sqrt{n}\{(\widehat{W}^T,\widehat{\beta}_1^T)^T-({W}^T,{\beta}_1^T)^T\}\overset{d}\to N(0,\Sigma_W)$ for some matrix $\Sigma_W$. Then it follows that \begin{align}\label{eq.beta asymptotic distribution weighted average} \sqrt{n}(\widehat{W}^T \widehat{\beta}_{1}-{W}^T\beta_1 ) \overset{d}\to N\left( 0,(\beta_1^T,W^T)\Sigma_{W} (\beta_1^T,W^T)^T \right). \end{align} \end{corollary} The LATE $\beta_{k',k}$ is not identified if $\mathcal{Z}_{(k,k')}\notin\mathscr{Z}_{\bar{M}}$. Let $\beta_{1S}=(\beta_{(\kappa_{1},\kappa_{1}^{\prime})}^{1},\ldots,\beta _{(\kappa_{S},\kappa_{S}^{\prime})}^{1})^{T}$ for some $S>0$. In our context, it is interesting to test hypotheses about $\beta^1_{(k,k')}$ with $\mathcal{Z}_{(k,k')}\in\mathscr{Z}_{\bar{M}}$ ($\beta^1_{(k,k')}=\beta_{k',k}$ by Theorem \ref{thm.IV estimator asymptotics pairwise binary D}): \begin{align} \mathrm{H}_{0}:\mathcal{Z}_{( \kappa_{1},\kappa_{1}^{\prime}) } \in\mathscr{Z}_{\bar{M}},\ldots,\mathcal{Z}_{( \kappa_{S},\kappa_{S}^{\prime }) }\in\mathscr{Z}_{\bar{M}},~R\left( \beta_{1S}\right) =0, \end{align} where $R$ is a (possibly nonlinear) smooth $r$-dimensional function. Let $R^{\prime }(\beta_{S})$ be the $r\times S$ matrix of the continuous first derivative functions of $R$ at an arbitrary value $\beta_{S}$, that is, $R^{\prime}(\beta_{S})=\partial R\left( \beta _{S}\right) /\partial\beta_{S}^{T}$. Let $\mathcal{I}_{S}$ be a $S\times \left( K-1\right) K$ matrix such that \[ \mathcal{I}_{S}\beta=(\beta_{(\kappa_{1},\kappa_{1}^{\prime})},\ldots ,\beta_{(\kappa_{S},\kappa_{S}^{\prime})})^{T} \] for every $\beta=(\beta_{\left( 1,2\right) },\ldots,\beta_{\left( 1,K\right) },\ldots,\beta_{\left( K,1\right) },\ldots,\beta_{\left( K,K-1\right) })^{T}$. Theorem \ref{thm.IV estimator asymptotics pairwise binary D} implies that \[ \sqrt{n}( \widehat{\beta}_{1S}-\beta_{1S}) =\sqrt {n}\mathcal{I}_{S}( \widehat{\beta}_{1}-\beta_{1}) \overset {d}{\rightarrow}N\left( 0,\Sigma_{S}\right) , \] where $\Sigma_{S}=\mathcal{I}_{S}\Sigma\mathcal{I}_{S}^{T}$, so that by the delta method, we obtain \[ \sqrt{n}\left\{ R( \widehat{\beta}_{1S}) -R\left( \beta _{1S}\right) \right\} \overset{d}{\rightarrow}N\left( 0,R^{\prime }\left( \beta_{1S}\right) \Sigma_{S}R^{\prime}\left( \beta_{1S} \right) ^{T}\right) . \] We construct the test statistics as \[ TS_{1n}=\prod_{s=1}^{S}1\left\{ \mathcal{Z}_{\left( \kappa_{s},\kappa _{s}^{\prime}\right) }\in\widehat{\mathscr{Z}_{0}}\right\} \] and \begin{align}\label{eq.TS2_1} TS_{2n}=&\,\sqrt{n}R( \widehat{\beta}_{1S}) ^{T}\left\{ R^{\prime}( \widehat{\beta}_{1S}) \mathcal{I}_{S} \widehat{\Sigma}\mathcal{I}_{S}^{T}R^{\prime}( \widehat{\beta}_{1S} ) ^{T}\right\} ^{-1}\sqrt{n}R( \widehat{\beta}_{1S}) , \end{align} where $\widehat{\Sigma}$ is a consistent estimator of $\Sigma$, which can be constructed based on the formula in \eqref{eq.weak convergence pairwise2}. Suppose that Assumptions \ref{ass.iid data binary D} and \ref{ass.first stage binary D} hold and $\mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_{\bar{M}})\to 1$. If $\mathrm{H}_{0}$ is true and $R'(\beta_{1S})$ is of full row rank, then it follows from standard arguments that $TS_{2n}\overset{d}{\rightarrow}\chi^2_{r}$, where $\chi^2_{r}$ denotes the chi-square distribution with $r$ degrees of freedom. The decision rule of the test is to reject $\mathrm{H}_{0}$ if $TS_{1n}=0$ or $TS_{2n}>c_{r}(\alpha)$, where $c_{r}(\alpha)$ satisfies $\mathbb{P}(\chi^2_{r}>c_{r}(\alpha))=\alpha$ for some predetermined $\alpha\in(0,1)$. The following proposition establishes the formal properties of the proposed test. \begin{proposition}\label{prop.test} Suppose that Assumptions \ref{ass.iid data binary D} and \ref{ass.first stage binary D} hold and $\mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_{\bar{M}})\to 1$. \begin{enumerate}[label=(\roman*)] \item If $\mathrm{H}_{0}$ is true, $\mathbb{P}\left( \left\{ TS_{1n}=0\right\} \cup\left\{ TS_{2n} >c_{r}(\alpha)\right\} \right) \rightarrow\alpha$. \item If $\mathrm{H}_{0}$ is false, $\mathbb{P}\left( \left\{ TS_{1n}=0\right\} \cup\left\{ TS_{2n}>c_{r}(\alpha)\right\} \right) \rightarrow1$. \end{enumerate} \end{proposition} \subsection{Asymptotic Bias Reduction under VSIV Estimation}\label{sec.bias reduction binary D} In Section \ref{sec.VSIV under consistency}, we show that if the estimator of the largest validity pair set is consistent, $ \mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_{\bar{M}})\to 1$, the VSIV estimators are consistent and asymptotically normal under weak conditions. However, since ${\mathscr{Z}_0}$ is constructed based on necessary (but not necessarily sufficient) conditions for IV validity, we have $ \mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_0)\to 1$ in general, where the \emph{pseudo-validity pair set} $\mathscr{Z}_0$ could be larger than $\mathscr{Z}_{\bar{M}}$. In this case, VSIV estimators may not be asymptotically unbiased. Consider an arbitrary presumed validity pair set $\mathscr{Z}_P$, which could incorporate prior information. If no prior information is available, we set $\mathscr{Z}_P=\mathscr{Z}$. Here we show even if ${\mathscr{Z}_0}$ is larger than $\mathscr{Z}_{\bar{M}}$, VSIV estimators based on $\widehat{\mathscr{Z}'_0}=\widehat{\mathscr{Z}_0}\cap \mathscr{Z}_P$ have weakly lower asymptotic biases than standard LATE estimators based on $\mathscr{Z}_P$. Intuitively, VSIV estimators use the information in the data about IV validity to reduce the asymptotic bias. Since our target parameter is the vector $\beta_1$, a natural definition of the asymptotic bias is as follows. \begin{definition}\label{def.bias} The asymptotic bias of an arbitrary estimator $\tilde{\beta}_1$ for the true value $\beta_1$ defined in \eqref{eq.VSIV true beta1 binary D} is defined as $\mathrm{plim}_{n\rightarrow\infty}\Vert \tilde{\beta}_1 -\beta_1\Vert_2$, where $\Vert\cdot\Vert_2$ is the $\ell^2$-norm on Euclidean spaces. \end{definition} The next assumption extends Assumption \ref{ass.first stage binary D} to $\mathscr{Z}_0$. \begin{assumption}\label{ass.first stage binary D bias} For every $\mathcal{Z}_{(k,k')}\in\mathscr{Z}_{0}$, \begin{align}\label{eq.first stage bias} E[g(Z_i)D_i|Z_i\in\mathcal{Z}_{(k,k^{\prime})}]-E[D_i|Z_i\in\mathcal{Z}_{(k,k^{\prime})}]\cdot E[g(Z_i)|Z_i\in\mathcal{Z}_{(k,k^{\prime})}]\neq0. \end{align} \end{assumption} The following theorem shows that the VSIV estimators based on $\widehat{\mathscr{Z}'_0}$ exhibit a smaller asymptotic bias than standard LATE estimators based on $\mathscr{Z}_P$. \begin{theorem}\label{thm.bias reduction binary D} Suppose that Assumptions \ref{ass.iid data binary D} and \ref{ass.first stage binary D bias} hold and that $\mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_0)\to 1$ with $\mathscr{Z}_0\supseteq\mathscr{Z}_{\bar{M}}$. For every presumed validity pair set $\mathscr{Z}_P$, the asymptotic bias of $\widehat{\beta}_1$ is reduced by using $\widehat{\mathscr{Z}'_0}$ in the estimation \eqref{eq.VSIV estimator pairwise binary D} compared to the asymptotic bias from using $\mathscr{Z}_P$. \end{theorem} As shown in Proposition \ref{prop.consistent G hat pairwise Z1 binary D} below, the pseudo-validity pair set $\mathscr{Z}_0$ can be estimated consistently by $\widehat{\mathscr{Z}_0}$ under mild conditions. Compared to constructing standard IV estimators based on $\mathscr{Z}_P$, Theorem \ref{thm.bias reduction binary D} shows that the asymptotic bias, $\mathrm{plim}_{n\rightarrow\infty}\Vert \widehat{\beta}_1-\beta_1\Vert_2$, can be reduced by using VSIV estimators based on $\widehat{\mathscr{Z}'_0}=\widehat{\mathscr{Z}_0}\cap\mathscr{Z}_P$. The arguments used for establishing the asymptotic normality of the VSIV estimators in Section \ref{sec.VSIV under consistency} do not rely on the consistent estimation of $\mathscr{Z}_{\bar{M}}$ ($\mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_{\bar{M}})\to 1$). If $\mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_{0})\to 1$ with $\mathscr{Z}_{\bar{M}}\subsetneq\mathscr{Z}_{0}$, the VSIV estimators are asymptotically normal, centered at $\beta_1$ defined with $\mathscr{Z}_0$ instead of $\mathscr{Z}_{\bar{M}}$. However, note that $\beta_1$ can only be interpreted as a vector of LATEs under consistent estimation ($\mathbb{P}(\widehat{\mathscr{Z}_0}=\mathscr{Z}_{\bar{M}})\to 1$). \begin{example}[Asymptotic Bias Reduction under VSIV Estimation] Consider a simple example\\ where $\mathcal{Z}=\{1,2,3,4\}$ as in our application and suppose that $\mathscr{Z}_{\bar{M}}=\{(1,2)\}$. In this case, by \eqref{eq.VSIV true beta binary D} and \eqref{eq.0timesinfinity}, \begin{align*} \beta_{1}=\left( \beta_{\left( 1,2\right) }^{1},\ldots,\beta_{\left( 1,4\right) }^{1},\ldots,\beta_{\left( 4,1\right) }^{1},\ldots ,\beta_{\left( 4,3\right) }^{1}\right)^T= \left( \beta_{\left( 1,2\right) }^{1},0,\ldots,0\right)^T. \end{align*} Suppose that, by mistake, we assume $Z$ is valid according to Assumption \ref{ass.IV validity binary D} and use \begin{align*} \mathscr{Z}_P=\{(1,2),(1,3),(1,4),(2,3),(2,4),(3,4)\} \end{align*} as an estimator for $\mathscr{Z}_{\bar{M}}$. Then by \eqref{eq.VSIV estimator pairwise binary D} and \eqref{eq.0timesinfinity}, \begin{align}\label{eq.beta1_hat example1 binary D} \widehat{\beta}_{1}=\left( \widehat{\beta}_{\left( 1,2\right) }^{1},\widehat{\beta}_{\left( 1,3\right) }^{1}, \widehat{\beta}_{\left( 1,4\right) }^{1},\widehat{\beta}_{\left( 2,3\right) }^{1},\widehat{\beta}_{\left( 2,4\right) }^{1},\widehat{\beta}_{\left( 3,4\right) }^{1},0,0,0,0,0,0\right)^T, \end{align} where $\widehat{\beta}_{(1,3)}^1$, $\widehat{\beta}_{(1,4)}^1$, $\widehat{\beta}_{(2,3)}^1$, $\widehat{\beta}_{(2,4)}^1$, and $\widehat{\beta}_{(3,4)}^1$ may not converge to $0$ in probability. However, by definition ${\beta}_{(1,3)}^1=0$, ${\beta}_{(1,4)}^1=0$, ${\beta}_{(2,3)}^1=0$, ${\beta}_{(2,4)}^1=0$, and ${\beta}_{(3,4)}^1=0$. Thus, $\Vert \widehat{\beta}_1-\beta_1\Vert_2$ may not converge to $0$ in probability. The approach proposed in this paper helps reduce this asymptotic bias. We exploit the information in the data about IV validity to obtain the estimator $\widehat{\mathscr{Z}_0}$. Even if $\widehat{\mathscr{Z}_0}$ converges to a set larger than $\mathscr{Z}_{\bar{M}}$ (because we use the necessary but not sufficient conditions for IV validity), VSIV always reduces the asymptotic bias. Suppose that $\mathscr{Z}_0=\{(1,2),(3,4)\}$, which is larger than $\mathscr{Z}_{\bar{M}}$ but smaller than $\mathscr{Z}_P$. In this case, the VSIV estimator $\widehat{\beta}_1$ constructed by using $\widehat{\mathscr{Z}_0}\cap\mathscr{Z}_P$ converges in probability to \begin{align}\label{eq.beta1_hat example2 binary D} {\beta}'_{1}=\left( {\beta}_{\left( 1,2\right) }^{1},0, 0,0,0,{\beta}^{1\prime}_{\left( 3,4\right) },0,0,0,0,0,0\right)^T, \end{align} where ${\beta}^{1\prime}_{\left( 3,4\right) }$ is the probability limit of $\widehat{\beta}^1_{\left( 3,4\right) }$. Then, clearly, VSIV reduces the probability limit of $\Vert \widehat{\beta}_1-\beta_1\Vert_2 $. \end{example} \subsection{Partially Valid Instruments and Connection to Existing Results} \label{sec.partially valid instruments binary D} Suppose we estimate the following canonical IV regression model, \begin{align}\label{eq.model equation binary D} Y_i=\alpha_{0}+\alpha_{1}D_i+\epsilon_{i}, \end{align} using $g(Z_i)$ as the instrument for $D_i$. When the instrument $Z$ is fully valid, the traditional IV estimator of $\alpha_1$ is \begin{align}\label{eq.alpha_hat} \widehat{\alpha}_{1}=\frac{n\sum_{i=1}^{n}g\left( Z_{i}\right) Y_{i}-\sum _{i=1}^{n}g\left( Z_{i}\right) \sum_{i=1}^{n}Y_{i}}{n\sum_{i=1}^{n}g\left( Z_{i}\right) D_{i}-\sum_{i=1}^{n}g\left( Z_{i}\right) \sum_{i=1}^{n}D_{i}}. \end{align} The asymptotic properties of $\widehat{\alpha}_1$ can be found in \citet[p.~471]{imbens1994identification} and \citet[p.~436]{angrist1995two}. To connect VSIV estimation to canonical IV regression with fully valid instruments, consider the following special case of pairwise IV validity. \begin{definition}\label{def.partially validy instrument binary D} Suppose that the instrument $Z$ is pairwise valid for the treatment $D$ with the largest validity pair set $\mathscr{Z}_{\bar{M}}$. If there is a validity pair set $$\mathscr{Z}_{{M}}=\{(z_{k_1},z_{k_2}),(z_{k_2},z_{k_3}),\ldots,(z_{k_{M-1}},z_{k_M})\}$$ for some $M>0$, then the instrument $Z$ is called a \textbf{partially valid instrument} for the treatment $D$. The set $\mathcal{Z}_M=\{z_{k_1},\ldots,z_{k_M}\}$ is called a \textbf{validity value set} of $Z$. \end{definition} \begin{assumption}\label{ass.first stage binary D partial} The validity value set $\mathcal{Z}_M$ satisfies that \begin{align} E[g(Z_i)D_i|Z_i\in\mathcal{Z}_{M}]-E[D_i|Z_i\in\mathcal{Z}_{M}]\cdot E[g(Z_i)|Z_i\in\mathcal{Z}_{M}]\neq0. \end{align} \end{assumption} Suppose that $Z$ is partially valid for the treatment $D$ with a validity value set $\mathcal{Z}_M$ and that there is a consistent estimator $\widehat{\mathcal{Z}_{0}}$ of $\mathcal{Z}_M$. We then construct a VSIV estimator for $\alpha_1$ in \eqref{eq.model equation binary D} by estimating the model \begin{align}\label{eq.VSIV estimation binary D} Y_i1\left\{ Z_{i}\in\widehat{\mathcal{Z}_{0}}\right\} =\gamma_{0}1\left\{ Z_{i}\in\widehat{\mathcal{Z}_{0}}\right\}+\gamma_{1}D_i1\left\{ Z_{i}\in\widehat{\mathcal{Z}_{0}}\right\} +\epsilon_{i}1\left\{ Z_{i}\in\widehat{\mathcal{Z}_{0}}\right\}, \end{align} using $g(Z_i)1\{ Z_{i}\in\widehat{\mathcal{Z}_{0}}\}$ as the instrument for $D_i1\{ Z_{i}\in\widehat{\mathcal{Z}_{0}}\}$. We obtain the VSIV estimator for $\alpha_1$ in \eqref{eq.model equation binary D} by \begin{align}\label{eq.VSIV estimator binary D} \widehat{\theta}_{1}=\frac{n_z\sum_{i=1}^{n}g\left( Z_{i}\right) Y_{i}1\left\{ Z_{i}\in\widehat{\mathcal{Z}_{0}}\right\} -\sum_{i=1}^{n}g\left( Z_{i}\right) 1\left\{ Z_{i}\in\widehat{\mathcal{Z}_{0}}\right\} \sum _{i=1}^{n}Y_{i}1\left\{ Z_{i}\in\widehat{\mathcal{Z}_{0}}\right\} } {n_z\sum_{i=1}^{n}g\left( Z_{i}\right) D_{i}1\left\{ Z_{i}\in\widehat {\mathcal{Z}_{0}}\right\} -\sum_{i=1}^{n}g\left( Z_{i}\right) 1\left\{ Z_{i}\in\widehat{\mathcal{Z}_{0}}\right\} \sum_{i=1}^{n}D_{i}1\left\{ Z_{i}\in\widehat{\mathcal{Z}_{0}}\right\} }, \end{align} where $n_z=\sum_{i=1}^n1\{Z_{i}\in\widehat{\mathcal{Z}_{0}}\}$. We can see that $\widehat{\theta}_1$ is a generalized version of $\widehat{\alpha}_1$ in \eqref{eq.alpha_hat}, because when the instrument is fully valid, we can just let $\widehat{\mathcal{Z}_{0}}=\mathcal{Z}$ and then $\widehat{\theta}_1=\widehat{\alpha}_1$. \begin{theorem}\label{thm.IV estimator asymptotics binary D} Suppose that the instrument $Z$ is partially valid for the treatment $D$ according to Definition \ref{def.partially validy instrument binary D} with a validity value set $\mathcal{Z}_M=\{z_{k_1},\dots,z_{k_M} \}$, and that the estimator $\widehat{\mathcal{Z}_0}$ for $\mathcal{Z}_M$ satisfies $\mathbb{P}(\widehat{\mathcal{Z}_{0}}=\mathcal{Z}_{M})\rightarrow 1$. Under Assumptions \ref{ass.iid data binary D} and \ref{ass.first stage binary D partial}, it follows that $\widehat{\theta}_{1}\overset{p}\rightarrow\theta_1 $, where \begin{align*} \theta_{1}=\frac{E\left[ g\left( Z_{i}\right) Y_{i}| Z_{i}\in\mathcal{Z}_{M} \right] -E\left[ Y_{i}| Z_{i}\in\mathcal{Z}_{M} \right] E\left[ g\left( Z_{i}\right)| Z_{i}\in\mathcal{Z}_{M} \right] }{E\left[ g\left( Z_{i}\right) D_{i}| Z_{i}\in\mathcal{Z}_{M} \right] -E\left[ D_{i}| Z_{i}\in\mathcal{Z}_{M} \right] E\left[ g\left( Z_{i}\right)| Z_{i}\in\mathcal{Z}_{M} \right] }. \end{align*} Also, $\sqrt{n}( \widehat{\theta}_{1}-\theta_1 ) \overset{d}\to N\left( 0,\Sigma_1 \right) $, where $\Sigma_1$ is provided in \eqref{eq.asymptotic beta1} in the Appendix. In addition, the quantity $\theta_{1}$ can be interpreted as the weighted average of $\{ \beta _{k_{2},k_{1}},\ldots ,\beta_{k_{M},k_{M-1}}\} $ defined as in \eqref{eq.beta binary D}. Specifically, $\theta _{1}=\sum_{m=1}^{M-1}\mu _{m}\beta _{k_{m+1},k_{m}}$ with \begin{align*} \mu _{m}= \frac{\left[ p\left( z_{k_{m+1}}\right) -p\left( z_{k_{m}}\right) \right] \sum_{l=m}^{M-1}\mathbb{P}\left( Z_{i}=z_{k_{l+1}}|Z_i\in\mathcal{Z}_M\right) \left\{ g\left( z_{k_{l+1}}\right) -E\left[ g\left( Z_{i}\right) |Z_i\in\mathcal{Z}_M \right] \right\} }{\sum_{l=1}^{M}\mathbb{P}\left( Z_{i}=z_{k_{l}}|Z_i\in\mathcal{Z}_M\right) p\left( z_{k_{l}}\right) \left\{ g\left( z_{k_{l}}\right) -E\left[ g\left( Z_{i}\right)|Z_i\in\mathcal{Z}_M \right] \right\} }, \end{align*} $p\left( z_{k}\right) =E\left[ D_{i}|Z_{i}=z_{k}\right] $, and $\sum_{m=1}^{M-1}\mu _{m}=1$. \end{theorem} Theorem \ref{thm.IV estimator asymptotics binary D} is an extension of Theorem 2 of \citet{imbens1994identification} to the case where the instrument is partially but not fully valid. To establish a connection to existing results, Theorem \ref{thm.IV estimator asymptotics binary D} assumes consistent estimation of the validity value set such that $\mathbb{P}(\widehat{\mathcal{Z}_{0}}=\mathcal{Z}_{M})\rightarrow 1$. If $\widehat{\mathcal{Z}_{0}}$ converges to a larger set than $\mathcal{Z}_{M}$, the properties of VSIV estimation follow from the results in Section \ref{sec.bias reduction binary D} because partially valid instruments are a special case of pairwise valid instruments. \section{Definition and Estimation of $\mathscr{Z}_{0}$} \label{sec.estimation validity set} Here we discuss the definition and the estimation of $\mathscr{Z}_0$ based on the testable implications in \citet{kitagawa2015test}, \citet{mourifie2016testing}, \citet{kedagni2020generalized}, and \citet{sun2021ivvalidity} for pairwise IV validity. We show that under weak assumptions, the proposed estimator $\widehat{\mathscr{Z}_{0}}$ is consistent for the pseudo-validity pair set $\mathscr{Z}_{0}$ in the sense that $\mathbb{P}(\widehat{\mathscr{Z}_{0}}=\mathscr{Z}_{0})\rightarrow 1$. As a consequence, when $\mathscr{Z}_{0}=\mathscr{Z}_{\bar{M}}$, the largest validity pair set can be estimated consistently. Specifically, we first construct two sets, $\mathscr{Z}_1$ and $\mathscr{Z}_2$, of pairs of instrument values such that $\mathscr{Z}_1$ satisfies the testable implications in \citet{kitagawa2015test}, \citet{mourifie2016testing}, and \citet{sun2021ivvalidity}, and $\mathscr{Z}_2$ satisfies the testable implications in \citet{kedagni2020generalized}. We then construct $\mathscr{Z}_0$ as the intersection of these two sets, $\mathscr{Z}_0=\mathscr{Z}_1\cap \mathscr{Z}_2$ (see Appendices \ref{sec.estimation Z_0 ordered} and \ref{sec.estimation Z_0 unordered}). Lemma \ref{lemma.testable implications for Z2 weaker} shows that $\mathscr{Z}_1\subseteq\mathscr{Z}_2$ when $D$ is binary. For multivalued $D$, we are not aware of such a result. \begin{lemma}\label{lemma.testable implications for Z2 weaker} If $D\in\{0,1\}$, then $\mathscr{Z}_1\subseteq\mathscr{Z}_2$, and the testable restrictions defining $\mathscr{Z}_1$ are sharp.\footnote{The statement in Lemma \ref{lemma.testable implications for Z2 weaker} holds for ordered or unordered $D\in\mathcal{D}=\{d_1,d_2\}$.} \end{lemma} Lemma \ref{lemma.testable implications for Z2 weaker} shows that when $D\in\{0,1\}$, the testable implications of \citet{kedagni2020generalized} are implied by those of \citet{kitagawa2015test}, \citet{mourifie2016testing}, and \citet{sun2021ivvalidity}. This result may be of independent interest. Moreover, Lemma \ref{lemma.testable implications for Z2 weaker} establishes the sharpness of the testable implications used to define $\mathscr{Z}_1$. This follows from the arguments in Proposition 1.1(i) of \citet{kitagawa2015test}. Lemma \ref{lemma.testable implications for Z2 weaker} implies that $\mathscr{Z}_0=\mathscr{Z}_1\cap \mathscr{Z}_2=\mathscr{Z}_1$ when $D$ is binary. Therefore, we define $\mathscr{Z}_0$ based on the testable implications proposed in \citet{kitagawa2015test}, \citet{mourifie2016testing}, and \citet{sun2021ivvalidity}. These testable implications were originally proposed for full IV validity. In the following, we extend them to Definition \ref{def.partial validity pairwise binary D}. To describe the testable restrictions, we use the notation of \citet{sun2021ivvalidity}. Define conditional probabilities \begin{equation*} P_z\left( B,C\right) =\mathbb{P}\left( Y\in B,D\in C|Z=z\right) \end{equation*} for all Borel sets $B,C\in\mathcal{B}_{\mathbb{R}}$ and all $z\in\mathcal{Z}$. With $\mathscr{Z}_{\bar{M}}=\{(z_{k_1},z_{k_1^{\prime}}),\ldots,(z_{k_{\bar{M}}},z_{k_{\bar{M}}^{\prime}})\}$, for every $m\in\{1,\ldots,\bar{M}\}$, it follows that \begin{align}\label{eq.testable implication binary D} P_{z_{k_m}}\left( B,\{1\}\right) \leq P_{z_{k_m^{\prime}}}\left( B,\{1\}\right) \text{ and } P_{z_{k_m}}\left( B,\{0\}\right) \geq P_{z_{k_m^{\prime}}}\left( B,\{0\}\right) \end{align} for all $B\in\mathcal{B}_{\mathbb{R}}$. By definition, for all $B,C\in \mathcal{B}_{\mathbb{R}}$, \begin{equation*} \mathbb{P}\left( Y\in B,D\in C|Z=z\right) =\frac{\mathbb{P}\left( Y\in B,D\in C,Z=z\right) }{\mathbb{P}\left( Z=z\right) }. \end{equation*} Define the function spaces \begin{align}\label{def.function spaces binary D} & \mathcal{G}_P=\left\{ \left( 1_{\mathbb{R}\times \mathbb{R}\times \left\{ z_{k}\right\} },1_{\mathbb{R}\times \mathbb{R}\times \left\{ z_{k^{\prime }}\right\} }\right) : k,k^{\prime }\in\{1,\ldots, K\}, k\neq k'\right\} , \notag \\ & \mathcal{H}=\left\{ \left( -1\right) ^{d}\cdot 1_{B\times \left\{ d\right\} \times \mathbb{R}}:B\text{ is a closed interval in }\mathbb{R} ,d\in \{0,1\}\right\} , \text{ and }\notag \\ & \bar{\mathcal{H}}=\left\{ \left( -1\right) ^{d}\cdot 1_{B\times \left\{ d\right\} \times \mathbb{R}}:B\text{ is a closed, open, or half-closed interval in }\mathbb{R},d\in \left\{ 0,1\right\} \right\}. \end{align} Similarly to \citet{sun2021ivvalidity}, by Lemma B.7 in \citet{kitagawa2015test}, we use all closed intervals $ B\subseteq \mathbb{R}$ to construct ${\mathcal{H}}$ instead of all Borel sets. Suppose we have access to an i.i.d.\ sample $\{\left( Y_{i},D_{i},Z_{i}\right) \}_{i=1}^{n}$ distributed according to some probability distribution $P$ in $\mathcal{P}$, that is, $P(G)=\mathbb{P}((Y_{i},D_{i},Z_{i})\in G)$ for all measurable $G$. The closure of $\mathcal{H}$ in $L^{2}(P)$ is equal to $\bar{\mathcal{H}}$ by Lemma C.1 of \citet{sun2021ivvalidity}. For every $\left( h,g\right) \in {\bar{ \mathcal{H}}\times \mathcal{G}_P}$ with $g=(g_{1},g_{2})$, we define \begin{equation*} \phi \left( h,g\right) =\frac{P\left( h\cdot g_{2}\right) }{P\left( g_{2}\right) }-\frac{P\left( h\cdot g_{1}\right) }{P\left( g_{1}\right) } \end{equation*} and \begin{align}\label{eq.stat variance multi} \sigma^2(h,g)=\Lambda(P) \cdot \left\{\frac{ P\left( h^2\cdot g_{2}\right) }{P^{2}\left(g_{2}\right) } -\frac{ P^2\left( h\cdot g_{2}\right) }{P^{3}\left( g_{2}\right) } +\frac{ P\left( h^2\cdot g_{1}\right) }{P^{2}\left( g_{1}\right) } -\frac{ P^2\left( h\cdot g_{1}\right) }{P^{3}\left( g_{1}\right) }\right\}, \end{align} where $\Lambda(P)=\prod_{k=1}^{K}P(1_{\mathbb{R}\times\mathbb{R}\times\{z_k\}})$ and $P^m(g_j)=[P(g_j)]^m$ for $m\in\mathbb{N}$ and $j\in\{1,2\}$. We denote the sample analog of $\phi $ as \begin{equation*} \widehat{\phi}\left( h,g\right) =\frac{\widehat{P}(h\cdot g_{2})}{\widehat{P}(g_{2})}- \frac{\widehat{P}(h\cdot g_{1})}{\widehat{P}(g_{1})}, \end{equation*} where $\widehat{P}$ is the empirical probability measure corresponding to $P$ so that for every measurable function $v$ (by abuse of notation), \begin{equation}\label{eq.empirical P} \widehat{P}\left( v\right) =\frac{1}{n}\sum_{i=1}^{n}v\left( Y_{i},D_{i},Z_{i}\right). \end{equation} For every $\left( h,g\right) \in \bar{\mathcal{H}}\times\mathcal{ G}_P$ with $g=\left( g_{1},g_{2}\right) $, define the sample analog of $\sigma^2(h,g)$ as \begin{equation*} \widehat{\sigma}^{2}\left( h,g\right) =\frac{T_{n}}{n}\cdot \left\{ \frac{\widehat{P} \left( h^{2}\cdot g_{2}\right) }{\widehat{P}^{2}\left( g_{2}\right) }-\frac{\widehat{ P}^{2}\left( h\cdot g_{2}\right) }{\widehat{P}^{3}\left( g_{2}\right) }+\frac{ \widehat{P}\left( h^{2}\cdot g_{1}\right) }{\widehat{P}^{2}\left( g_{1}\right) }- \frac{\widehat{P}^{2}\left( h\cdot g_{1}\right) }{\widehat{P}^{3}\left( g_{1}\right) }\right\} , \end{equation*} where $T_{n}=n\cdot \prod_{k=1}^{K}\widehat{P}(1_{\mathbb{R}\times \mathbb{R} \times \{z_{k}\}})$. By \eqref{eq.0timesinfinity}, $\widehat{\sigma}^{2}$ is well defined. By similar arguments as in the proof of Lemma 3.1 in \citet{sun2021ivvalidity}, ${\sigma}^{2}$ and $\widehat{\sigma}^{2}$ are uniformly bounded in $(h,g)$. The following lemma reformulates the testable restrictions in \eqref{eq.testable implication binary D} in terms of $\phi$. Below, we use this reformulation to define $\mathscr{Z}_0$ and the corresponding estimator $\widehat{\mathscr{Z}_0}$. \begin{lemma}\label{lemma.superset of Z pairwise binary D} Suppose that the instrument $Z$ is pairwise valid for the treatment $D$ with the largest validity pair set $\mathscr{Z}_{\bar{M}}=\{(z_{k_1},z_{k_1^{\prime}}),\ldots,(z_{k_{\bar{M}}},z_{k_{\bar{M}}^{\prime}})\}$. For every $m\in\{1,\ldots,\bar{M}\}$, $\sup_{h\in {\mathcal{H}}}\phi \left( h,g\right) =0$ with $g=( 1_{\mathbb{R}\times \mathbb{R}\times \{ z_{k_m}\} },1_{\mathbb{R}\times \mathbb{R}\times \{ z_{k_{m}^{\prime}}\} })$. \end{lemma} Lemma \ref{lemma.superset of Z pairwise binary D} reformulates the necessary conditions based on \citet{kitagawa2015test}, \citet{mourifie2016testing}, and \citet{sun2021ivvalidity} for the validity pair set $\mathscr{Z}_{\bar{M}}$. Define \begin{equation}\label{eq.G0 pair binary D} \mathcal{G}_{0}=\left\{ g\in \mathcal{G}_P:\sup_{h\in {\mathcal{H}}}\phi\left( h,g\right) =0\right\} \text{ and } \widehat{\mathcal{G}_{0}}=\left\{ g\in \mathcal{G}_P:\sqrt{T_n}\left\vert \sup_{h\in {\mathcal{H}} }\frac{\widehat{\phi} \left( h,g\right)}{\xi_{0}\vee \widehat{\sigma}(h,g)} \right\vert \leq \tau _{n}^{g}\right\}, \end{equation} where $1/\min_{g\in\mathcal{G}_P}\tau_{n}^g\to0$ in probability and $\max_{g\in\mathcal{G}_P}\tau_{n}^g/\sqrt{n}\to0$ in probability as $n\to\infty$, and $\xi_{0}$ is a small positive number.\footnote{ In practice, we use $\xi_0=10^{-100}$.} Here, we allow $\tau_n^g$ to be different for every $g$. This added flexibility helps improve the finite sample performance of VSIV estimation. We discuss our explicit choice of $\tau_n^g$ in more detail in Section \ref{sec.simulation}. The set $\mathcal{G}_0$ is different from the contact sets defined in \citet{Beare2015improved}, \citet{Beare2017improved}, and \citet{sun2021ivvalidity} in different contexts because of the presence of the map $\sup$. We refer to \citet{linton2010improved} and \citet{lee2018testing} for further discussions of contact set estimation. Define ${\mathscr{Z}_0}$ as the collection of all $(z,z')$ associated with some $g\in{\mathcal{G}_0}$: \begin{align}\label{eq.Z1 pair binary D} \mathscr{Z}_0=\left\{ (z_{k},z_{k^{\prime}})\in\mathscr{Z}: g=( 1_{\mathbb{R}\times \mathbb{R}\times \{ z_{k}\} },1_{\mathbb{R}\times \mathbb{R}\times \{ z_{k^{\prime}}\} })\in\mathcal{G}_0\right\}. \end{align} Note that $\mathscr{Z}_0$ is the set of all pairs that satisfy the testable implications in \eqref{eq.testable implication binary D}. For example, if $K=4$ and ${\mathcal{G}_0}=\{( 1_{\mathbb{R}\times \mathbb{R}\times \left\{ z_{1}\right\} },1_{\mathbb{R}\times \mathbb{R}\times \left\{ z_{2}\right\} }), ( 1_{\mathbb{R}\times \mathbb{R}\times \left\{ z_{3}\right\} },1_{\mathbb{R}\times \mathbb{R}\times \left\{ z_{4}\right\} })\}$, then ${\mathscr{Z}_0}=\{(z_1,z_2),(z_3,z_4)\}$. We use $\widehat{\mathcal{G}_0}$ to construct the estimator of $\mathscr{Z}_0$, denoted by $\widehat{\mathscr{Z}_0}$, which is defined as the set of all $(z,z^{\prime})$ associated with some $g\in\widehat{\mathcal{G}_0}$: \begin{align}\label{eq.Z1_hat pair binary D} \widehat{\mathscr{Z}_0}=\left\{ (z_{k},z_{k^{\prime}})\in\mathscr{Z}: g=( 1_{\mathbb{R}\times \mathbb{R}\times \{ z_{k}\} },1_{\mathbb{R}\times \mathbb{R}\times \{ z_{k^{\prime}}\} })\in\widehat{\mathcal{G}_0}\right\}. \end{align} Note that \eqref{eq.Z1_hat pair binary D} is the sample analog of \eqref{eq.Z1 pair binary D}. The following proposition establishes consistency of $\widehat{\mathscr{Z}_0}$. \begin{proposition} \label{prop.consistent G hat pairwise Z1 binary D} Under Assumption \ref{ass.iid data binary D}, $\mathbb{P}(\widehat{\mathcal{G}_0}=\mathcal{G}_0)\rightarrow 1$, and thus $\mathbb{P}(\widehat{\mathscr{Z}_{0}}=\mathscr{Z}_{0})\rightarrow 1$. \end{proposition} Proposition \ref{prop.consistent G hat pairwise Z1 binary D} is related to the contact set estimation in \citet{sun2021ivvalidity}. Since, by definition, $\mathcal{G}_0\subseteq\mathcal{G}_P$ and $\mathcal{G}_P$ is a finite set, we can use techniques similar to those in \citet{sun2021ivvalidity} to obtain the stronger result in Proposition \ref{prop.consistent G hat pairwise Z1 binary D}, that is, $\mathbb{P}(\widehat{\mathcal{G}}_0=\mathcal{G}_0)\rightarrow 1$. \begin{remark}[Uniqueness of $\mathscr{Z}_{\bar{M}}$ and $\mathscr{Z}_0$] By definition, $\mathscr{Z}_{\bar{M}}$ is the collection of all valid pairs of values of $Z$, while $\mathscr{Z}_0$ is the collection of all pairs that satisfy the testable restrictions in Lemma \ref{lemma.superset of Z pairwise binary D}. Thus, both $\mathscr{Z}_{\bar{M}}$ and $\mathscr{Z}_0$ are unique, and clearly $\mathscr{Z}_{\bar{M}}\subseteq\mathscr{Z}_0$. \end{remark} \begin{remark}[Local Violations of the Testable Restrictions] The theory for VSIV estimation relies on selection consistency, $\mathbb{P}(\widehat{\mathscr{Z}_{0}}=\mathscr{Z}_{0})\rightarrow 1$. A potential concern with this approach is that selection consistency is theoretically only possible if the violations of the testable restrictions are well-separated from zero. However, VSIV estimation remains useful even in the presence of local (to-zero) violations. It will always yield asymptotic bias reductions from removing pairs of instrument values for which the violations are well-separated from zero. An important feature of VSIV estimation is that it proceeds pairwise, so that the presence of pairs corresponding to local violations does not impact the selection performance for pairs corresponding to well-separated violations. We explore the performance of our method when we change the magnitude of the violations in Appendix \ref{sec.additional simulations}. \end{remark} \section{Simulations and Application}\label{sec.simulation} In this section, we evaluate the finite sample performance of VSIV estimation in Monte Carlo simulations designed based on an empirical application. First, we discuss the empirical context. Second, we present the results from the simulation study and determine the choice of the tuning parameter $\tau_n^g$. Finally, we present the results from the empirical application. \subsection{Empirical Context} We revisit the analysis of the causal effect of college education on earnings. We use the dataset analyzed by \citet{heckman2001four} and \citet{kedagni2020discordant}, consisting of data on 1,230 white males from the National Longitudinal Survey of Youth of 1979 (NLSY). The outcome of interest ($Y$) is the log wage and the treatment ($D$) is a dummy variable for college enrollment. \citet{kedagni2020discordant} use the maximum parental education ($E$) as an instrument for college enrollment. Since the overall sample size is relatively small, we consider a coarsened version of this instrument: ${Z}=1\{E<12\}+2\cdot1\{E=12\}+3\cdot1\{12<E<16\}+4\cdot1\{E\ge 16\}$. The parental education instrument likely violates the exclusion restriction due to its potential positive effect on earnings, especially at lower levels of parental education. Therefore, \citet{kedagni2020discordant} consider a relaxation of IV validity, building on \citet{manski2000monotone,manski2009more}. As discussed in Section \ref{sec.pairwise valid instrument binary D}, their relaxation allows the exclusion restriction to be violated for IV values below a cutoff, which they estimate to be $E=11$ (Table 2, Column (3)). The results in \citet{kedagni2020discordant} suggest that the parental education instrument is partially invalid, especially for pairs of instrument values corresponding to lower levels of parental education, thus providing an ideal setting for illustrating the usefulness of VSIV estimation. We emphasize that there are two major differences between our analysis here and the one in \citet{kedagni2020discordant}. First, we consider a different relaxation of IV validity. Unlike \citet{kedagni2020discordant}, we do not impose any assumptions on how IV validity fails for invalid pairs. Second, \citet{kedagni2020discordant} focus on partial identification of the average treatment effect. By contrast, we consider the estimation of LATEs for all pairs of instrument values that are not screened out based on the testable implications. \subsection{Simulation Evidence} \subsubsection{Data Generating Processes} Here we describe the data generating processes (DGPs) that we use in the simulations. All DGPs are calibrated to the joint empirical distribution of $(D,Z)$ in the application. Denote by $\widehat{\mathbb{P}}$ the empirical probability measure of ${\mathbb{P}}$. In the empirical application, we have $\widehat{\mathbb{P}}(Z=1)=0.1317$, $\widehat{\mathbb{P}}(Z=2)=0.4716$, $\widehat{\mathbb{P}}(Z=3)=0.1495$, $\widehat{\mathbb{P}}(Z=4)=0.2472$, $\widehat{\mathbb{P}}(D=1|Z=1)=0.1420$, $\widehat{\mathbb{P}}(D=1|Z=2)=0.3086$, $\widehat{\mathbb{P}}(D=1|Z=3)=0.5054$, and $\widehat{\mathbb{P}}(D=1|Z=4)=0.7796$. Given the prior information that $Z$ may be positively correlated with $D$, we assume that $\mathscr{Z}_P=\{(1,2),(1,3),(1,4),(2,3),(2,4),(3,4)\}$.\footnote{This is consistent with the fact that the empirical proportions $\widehat{\mathbb{P}}(D=1|Z=z)$ are increasing in $z$.} We consider five data generating processes (DGPs (0)--(4)), where Definition \ref{def.partial validity pairwise binary D} holds for each pair in $\mathscr{Z}_P$ under DGP (0) and is violated for every pair under DGPs (1)--(4). These DGPs are constructed following those in \citet{kitagawa2015test} and \citet{sun2021ivvalidity}. For all DGPs, we specify $U\sim\mathrm{Unif}(0,1)$, $V\sim\mathrm{Unif}(0,1)$, $W\sim\mathrm{Unif}(0,1)$, $Z=1\{U \le 0.1317\}+2\times\{0.1317<U \le 0.6033\}+3\times\{0.6033<U\le 0.7528\}+4\times 1\{U>0.7528\}$, $D_1=1\{V\le 0.1420\}$, $D_2=1\{V\le 0.3086\}$, $D_3=1\{V\le 0.5054\}$, $D_4=1\{V\le 0.7796\}$, and $D=\sum_{z=1}^4 1\{Z=z\}\times D_z$. We let $N_Z\sim N(0,1)$, and let $N_{10a}(\sigma)\sim N(-1,\sigma^2)$, $N_{10b}(\sigma)\sim N(-0.5,\sigma^2)$, $N_{10c}(\sigma)\sim N(0,\sigma^2)$, $N_{10d}(\sigma)\sim N(0.5,\sigma^2)$, $N_{10e}(\sigma)\sim N(1,\sigma^2)$, and $N_{C}(\sigma)=1\{W\le 0.15\}\times N_{10a}(\sigma)+1\{0.15<W\le 0.35\}\times N_{10b}(\sigma)+1\{0.35<W\le 0.65\}\times N_{10c}(\sigma)+1\{0.65<W\le 0.85\}\times N_{10d}(\sigma)+1\{W>0.85\}\times N_{10e}(\sigma)$ for $\sigma>0$. The DGPs (0)--(4) are specified as follows. \begin{enumerate}[start=0,label=(\arabic*):] \item $N_{dz}=N_Z$ for each $d\in\{0,1\}$ and all $z\in\{1,2,3,4\}$, $Y=\sum_{z=1}^{4}1\{Z=z\}\times(\sum_{d=0}^{1} 1\{D=d\}\times N_{dz})$ \item For $(z_1,z_2)\in\mathscr{Z}_P$, $N_{1z_1}\sim N(\mu_{(z_1,z_2)},1)$, $N_{1z_2}\sim N(0,1)$, $N_{0z_1}\sim N(0,1)$, $N_{0z_2}\sim N(\mu_{(z_1,z_2)},1)$ with $\mu_{(1,2)}=-0.9$, $\mu_{(1,3)}=-1.1$, $\mu_{(1,4)}=-1.3$, $\mu_{(2,3)}=-0.9$, $\mu_{(2,4)}=-1.1$, $\mu_{(3,4)}=-0.9$, $N_{dz}\sim N(0,1)$ for $d\in\{0,1\}$ and $z\in\{1,2,3,4\}\setminus \{z_1,z_2\}$, $Y=\sum_{z=1}^{4}1\{Z=z\}\times(\sum_{d=0}^{1} 1\{D=d\}\times N_{dz})$ \item For $(z_1,z_2)\in\mathscr{Z}_P$, $N_{1z_1}\sim N(0,1)$, $N_{1z_2}\sim N(0,\sigma_{(z_1,z_2)}^2)$, $N_{0z_1}\sim N(0,\sigma_{(z_1,z_2)}^2)$, $N_{0z_2}\sim N(0,1)$ with $\sigma_{(1,2)}=3$, $\sigma_{(1,3)}=5$, $\sigma_{(1,4)}=7$, $\sigma_{(2,3)}=3$, $\sigma_{(2,4)}=5$, $\sigma_{(3,4)}=3$, $N_{dz}\sim N(0,1)$ for $d\in\{0,1\}$ and $z\in\{1,2,3,4\}\setminus \{z_1,z_2\}$, $Y=\sum_{z=1}^{4}1\{Z=z\}\times(\sum_{d=0}^{1} 1\{D=d\}\times N_{dz})$ \item For $(z_1,z_2)\in\mathscr{Z}_P$, $N_{1z_1}\sim N(0,1)$, $N_{1z_2}\sim N(0,\sigma_{(z_1,z_2)}^2)$, $N_{0z_1}\sim N(0,\sigma_{(z_1,z_2)}^2)$, $N_{0z_2}\sim N(0,1)$ with $\sigma_{(1,2)}=0.5$, $\sigma_{(1,3)}=0.45$, $\sigma_{(1,4)}=0.4$, $\sigma_{(2,3)}=0.5$, $\sigma_{(2,4)}=0.45$, $\sigma_{(3,4)}=0.5$, $N_{dz}\sim N(0,1)$ for $d\in\{0,1\}$ and $z\in\{1,2,3,4\}\setminus \{z_1,z_2\}$, $Y=\sum_{z=1}^{4}1\{Z=z\}\times(\sum_{d=0}^{1} 1\{D=d\}\times N_{dz})$ \item For $(z_1,z_2)\in\mathscr{Z}_P$, $N_{1z_1}\sim N_C(\sigma_{(z_1,z_2)})$, $N_{1z_2}\sim N(0,1)$, $N_{0z_1}\sim N(0,1)$, $N_{0z_2}\sim N_C(\sigma_{(z_1,z_2)})$ with $\sigma_{(1,2)}=0.08$, $\sigma_{(1,3)}=0.04$, $\sigma_{(1,4)}=0.02$, $\sigma_{(2,3)}=0.08$, $\sigma_{(2,4)}=0.04$, $\sigma_{(3,4)}=0.08$, $N_{dz}\sim N(0,1)$ for $d\in\{0,1\}$ and $z\in\{1,2,3,4\}\setminus \{z_1,z_2\}$, $Y=\sum_{z=1}^{4}1\{Z=z\}\times(\sum_{d=0}^{1} 1\{D=d\}\times N_{dz})$ \end{enumerate} The random variables $U$, $V$, $W$, and those generated from $N(\mu,\sigma^2)$ for some $\mu$ and $\sigma$ are mutually independent. \subsubsection{Choice of $\tau_n^g$} \label{sec.validity set estimation and choice of tau} As shown in \eqref{eq.G0 pair binary D}, the tuning parameter $\tau_n^g$ should satisfy that $1/\min_{g\in\mathcal{G}_P}\tau_{n}^g\to0$ in probability and $\max_{g\in\mathcal{G}_P}\tau_{n}^g/\sqrt{n}\to0$ in probability as $n\to\infty$. In our simulations, for every $g=( 1_{\mathbb{R}\times \mathbb{R}\times \{ z_{k}\} },1_{\mathbb{R}\times \mathbb{R}\times \{ z_{k^{\prime}}\} })$, we set $\tau_n^g$ as \begin{align} \tau_n^g=c\cdot \frac{(n\widehat{\mathbb{P}}(Z\in\{z_k,z_{k'}\}))^{1/5}}{\vert \widehat{\mathbb{P}}(D=1|Z=z_k)-\widehat{\mathbb{P}}(D=1|Z=z_{k'}) \vert^{1/5}}\label{eq:tuning_parameter} \end{align} for some constant $c>0$.\footnote{If ${\mathbb{P}}(D=1|Z=z_k)-{\mathbb{P}}(D=1|Z=z_{k'})=0$ for some pair $(z_k,z_{k'})$, we still have that $\tau_n^g/\sqrt{n}\to0$ in probability as $n\to\infty$ for every $g\in\mathcal{G}_P$, and $\tau_n^g$ satisfies the two conditions above in this case.} The numerator of \eqref{eq:tuning_parameter} captures the relevant sample size, and the denominator captures the idea that if the difference is large, violations are harder to detect. This (heuristic) adjustment in the denominator does not affect the asymptotic properties but works well in simulations. \subsubsection{Simulation Results} We consider two different sample sizes: $n=1230$ as in the empirical application and $n=2460$ to explore how the properties of our methods improve as the sample size increases. We report results for $\tau_n^g$ in \eqref{eq:tuning_parameter} with $c\in \{0.1,0.2,\dots,1\}$. For each simulation, we use 1,000 Monte Carlo iterations. To calculate the supremum in $\sqrt{T_n}\vert \sup_{h\in\mathcal{H}}{\widehat{\phi} \left( h,g\right)}/({\xi_{0}\vee \widehat{\sigma}(h,g)}) \vert $ for every $g$, we use the approach employed by \citet{kitagawa2015test} and \citet{sun2021ivvalidity}. Specifically, we compute the supremum based on the closed intervals $[a,b]$ with the realizations of $\{Y_i\}$ as endpoints, that is, intervals $[a,b]$ where $a,b\in\{Y_i\}$ and $a\le b$. Tables \ref{tab:DGP0}--\ref{tab:DGP4} show the empirical probabilities with which each element of $\mathscr{Z}_P$ is selected to be in $\widehat{\mathscr{Z}_0}$ in the simulations. The results show that choosing $c$ is subject to a trade-off between the ability of our method to screen out invalid pairs and its ability to include valid pairs. Given the nature of the method, screening out invalid pairs is particularly important since LATE estimators based on these pairs are inconsistent. Our method with $c=0.6$ detects invalid pairs almost perfectly while selecting most of the valid pairs with high probability. With $c=0.6$, as $n$ increases from $1230$ to $2460$, the selection rates for valid pairs are increasing and those for invalid pairs are decreasing to $0$. Overall, the simulation results show that the proposed method performs well in identifying the validity pair set in finite samples. \begin{table}[h!] \centering \caption{Validity Pair Set Estimation for DGP (0)} \scalebox{0.9}{ \begin{tabular}{ c c c c c c c c } \hline \hline $n$ & $c $ & (1, 2) & {(1, 3)} & {(1, 4)} & {{(2, 3)}} & {(2, 4)} & {{(3, 4)}} \\ \hline \multirow{10}{*}{1230} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.002 & 0.000 & 0.000 & 0.001 \\ & 0.5 & 0.000 & 0.053 & 0.198 & 0.003 & 0.140 & 0.217 \\ & 0.6 & 0.003 & 0.585 & 0.726 & 0.264 & 0.733 & 0.821 \\ & 0.7 & 0.138 & 0.901 & 0.953 & 0.736 & 0.967 & 0.985 \\ & 0.8 & 0.542 & 0.983 & 0.996 & 0.956 & 0.998 & 1.000 \\ & 0.9 & 0.865 & 0.997 & 1.000 & 0.996 & 1.000 & 1.000 \\ & 1 & 0.968 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 \\ \hline \multirow{10}{*}{2460} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.006 \\ & 0.5 & 0.000 & 0.182 & 0.370 & 0.018 & 0.290 & 0.568 \\ & 0.6 & 0.018 & 0.802 & 0.898 & 0.482 & 0.908 & 0.956 \\ & 0.7 & 0.344 & 0.977 & 0.995 & 0.901 & 0.998 & 0.997 \\ & 0.8 & 0.821 & 0.998 & 1.000 & 0.989 & 1.000 & 1.000 \\ & 0.9 & 0.968 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 \\ & 1 & 0.997 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 \\ \hline \hline \end{tabular} } \label{tab:DGP0} \end{table} \begin{table}[h!] \centering \caption{Validity Pair Set Estimation for DGP (1)} \scalebox{0.9}{ \begin{tabular}{ c c c c c c c c } \hline \hline $n$ & $c $ & (1, 2) & {(1, 3)} & {(1, 4)} & {{(2, 3)}} & {(2, 4)} & {{(3, 4)}} \\ \hline \multirow{10}{*}{1230} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.002 & 0.011 & 0.000 & 0.001 & 0.017 \\ & 0.7 & 0.000 & 0.021 & 0.104 & 0.020 & 0.028 & 0.108 \\ & 0.8 & 0.000 & 0.086 & 0.298 & 0.126 & 0.148 & 0.311 \\ & 0.9 & 0.010 & 0.199 & 0.538 & 0.390 & 0.357 & 0.583 \\ & 1 & 0.051 & 0.375 & 0.743 & 0.689 & 0.592 & 0.787 \\ \hline \multirow{10}{*}{2460} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.000 & 0.004 & 0.000 & 0.000 & 0.001 \\ & 0.7 & 0.000 & 0.003 & 0.058 & 0.003 & 0.010 & 0.019 \\ & 0.8 & 0.001 & 0.010 & 0.219 & 0.037 & 0.046 & 0.093 \\ & 0.9 & 0.001 & 0.037 & 0.405 & 0.182 & 0.164 & 0.276 \\ & 1 & 0.005 & 0.118 & 0.648 & 0.509 & 0.352 & 0.533 \\ \hline \hline \end{tabular} } \label{tab:DGP1} \end{table} \begin{table}[h!] \centering \caption{Validity Pair Set Estimation for DGP (2)} \scalebox{0.9}{ \begin{tabular}{ c c c c c c c c } \hline \hline $n$ & $c $ & (1, 2) & {(1, 3)} & {(1, 4)} & {{(2, 3)}} & {(2, 4)} & {{(3, 4)}}\\ \hline \multirow{10}{*}{1230} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.000 & 0.001 & 0.000 & 0.000 & 0.004 \\ & 0.7 & 0.000 & 0.000 & 0.039 & 0.000 & 0.010 & 0.027 \\ & 0.8 & 0.000 & 0.003 & 0.188 & 0.024 & 0.063 & 0.134 \\ & 0.9 & 0.000 & 0.012 & 0.416 & 0.192 & 0.189 & 0.342 \\ & 1 & 0.008 & 0.037 & 0.630 & 0.479 & 0.418 & 0.564 \\ \hline \multirow{10}{*}{2460} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.000 & 0.001 & 0.000 & 0.000 & 0.000 \\ & 0.7 & 0.000 & 0.000 & 0.030 & 0.000 & 0.002 & 0.004 \\ & 0.8 & 0.000 & 0.000 & 0.185 & 0.007 & 0.021 & 0.050 \\ & 0.9 & 0.000 & 0.000 & 0.436 & 0.041 & 0.093 & 0.135 \\ & 1 & 0.000 & 0.000 & 0.652 & 0.185 & 0.241 & 0.282 \\ \hline \hline \end{tabular} } \label{tab:DGP2} \end{table} \begin{table}[h!] \centering \caption{Validity Pair Set Estimation for DGP (3)} \scalebox{0.9}{ \begin{tabular}{ c c c c c c c c } \hline \hline $n$ & $c $ & (1, 2) & {(1, 3)} & {(1, 4)} & {{(2, 3)}} & {(2, 4)} & {{(3, 4)}} \\ \hline \multirow{10}{*}{1230} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.002 \\ & 0.7 & 0.000 & 0.005 & 0.022 & 0.002 & 0.000 & 0.050 \\ & 0.8 & 0.000 & 0.033 & 0.141 & 0.064 & 0.004 & 0.282 \\ & 0.9 & 0.000 & 0.164 & 0.337 & 0.270 & 0.035 & 0.668 \\ & 1 & 0.003 & 0.436 & 0.630 & 0.556 & 0.182 & 0.922 \\ \hline \multirow{10}{*}{2460} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.7 & 0.000 & 0.000 & 0.001 & 0.000 & 0.000 & 0.001 \\ & 0.8 & 0.000 & 0.000 & 0.006 & 0.002 & 0.000 & 0.013 \\ & 0.9 & 0.000 & 0.003 & 0.050 & 0.055 & 0.000 & 0.193 \\ & 1 & 0.000 & 0.058 & 0.214 & 0.237 & 0.006 & 0.593 \\ \hline \hline \end{tabular} } \label{tab:DGP3} \end{table} \begin{table}[h!] \centering \caption{Validity Pair Set Estimation for DGP (4)} \scalebox{0.9}{ \begin{tabular}{ c c c c c c c c } \hline \hline $n$ & $c $ & (1, 2) & {(1, 3)} & {(1, 4)} & {{(2, 3)}} & {(2, 4)} & {{(3, 4)}} \\ \hline \multirow{10}{*}{1230} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.030 \\ & 0.7 & 0.000 & 0.001 & 0.001 & 0.036 & 0.002 & 0.302 \\ & 0.8 & 0.002 & 0.023 & 0.051 & 0.269 & 0.054 & 0.700 \\ & 0.9 & 0.043 & 0.173 & 0.239 & 0.606 & 0.241 & 0.914 \\ & 1 & 0.180 & 0.433 & 0.530 & 0.843 & 0.508 & 0.983 \\ \hline \multirow{10}{*}{2460} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.016 \\ & 0.7 & 0.000 & 0.000 & 0.000 & 0.031 & 0.000 & 0.220 \\ & 0.8 & 0.000 & 0.002 & 0.015 & 0.261 & 0.021 & 0.613 \\ & 0.9 & 0.010 & 0.030 & 0.136 & 0.649 & 0.150 & 0.865 \\ & 1 & 0.114 & 0.168 & 0.381 & 0.883 & 0.395 & 0.962 \\ \hline \hline \end{tabular} } \label{tab:DGP4} \end{table} Table \ref{tab:DGP0Inference} shows the coverage rates of the confidence intervals constructed based on the asymptotic distribution in Theorem \ref{thm.IV estimator asymptotics pairwise binary D} for DGP (0) under which all pairs are valid. We construct the confidence interval as follows: If a pair is in the estimated validity pair set, then the confidence interval is constructed based on the asymptotic distribution; if the pair is not in the estimated validity pair set, then the confidence interval is $\mathbb{R}$, since in this case the LATE is not identified. The results in Table \ref{tab:DGP0Inference} show that the coverage rates are close to $1-\alpha$, where $\alpha=0.05$, for most pairs when $c=0.6$. Since the sample size for each pair is relatively small in our simulations for both choices of $n$, the coverage rates can be somewhat higher than $0.95$ because of the way we construct the confidence intervals if the pair is not in the estimated validity pair set. In Appendix \ref{sec.simulation balanced}, we provide additional simulation results for a variant of DGP (0) with a more balanced design. The results in Tables \ref{tab:DGP0Balanced} and \ref{tab:DGP0InferenceBalanced} show that in this case for $c=0.6$, the selection rates for all pairs are high and converging to one, and the coverage rates are converging to $95%$. In Tables \ref{tab:RMSEDGP1}--\ref{tab:RMSEDGP4}, we explore the asymptotic bias reduction property of VSIV estimation by presenting the Root Mean Square Errors (RMSEs) of $\sqrt{n}(\widehat{\beta}_{(k,k')}^1-{\beta}_{(k,k')}^1)$ for every pair in $\mathscr{Z}_P$ under DGPs (1)--(4) under which all instrument pairs are invalid. As shown in Theorem \ref{thm.IV estimator asymptotics pairwise binary D}, $\sqrt{n}(\widehat{\beta}_{(k,k')}^1-{\beta}_{(k,k')}^1)\to0$ in probability if $(z_k,z_{k'})$ is invalid. Our simulation results show that overall, the RMSEs are getting closer to $0$ as $n$ increases for $c=0.6$. For larger $c$, the RMSEs for some invalid pairs may not be small enough in finite samples. Also note that for larger $c$, the RMSEs may not be decreasing in the sample size due to the rescaling by $\sqrt{n}$. According to these simulation results, we suggest using $c=0.6$ in applications. The results for other values of $c$ may also be presented for consideration. \begin{table}[h!] \centering \caption{Coverage Rates of the Confidence Intervals for DGP (0)} \scalebox{0.9}{ \begin{tabular}{ c c c c c c c c } \hline \hline $n$ & $c $ & (1, 2) & {(1, 3)} & {(1, 4)} & {{(2, 3)}} & {(2, 4)} & {{(3, 4)}} \\ \hline \multirow{10}{*}{1230} & 0.1 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 \\ & 0.2 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 \\ & 0.3 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 \\ & 0.4 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 \\ & 0.5 & 1.000 & 0.999 & 0.994 & 0.999 & 0.997 & 0.991 \\ & 0.6 & 1.000 & 0.984 & 0.964 & 0.995 & 0.968 & 0.955 \\ & 0.7 & 0.997 & 0.974 & 0.960 & 0.979 & 0.953 & 0.949 \\ & 0.8 & 0.985 & 0.969 & 0.959 & 0.971 & 0.952 & 0.948 \\ & 0.9 & 0.978 & 0.969 & 0.959 & 0.970 & 0.952 & 0.948 \\ & 1 & 0.972 & 0.969 & 0.959 & 0.970 & 0.952 & 0.948 \\ \hline \multirow{10}{*}{2460} & 0.1 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 \\ & 0.2 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 \\ & 0.3 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 \\ & 0.4 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 & 1.000 \\ & 0.5 & 1.000 & 0.995 & 0.979 & 1.000 & 0.983 & 0.979 \\ & 0.6 & 1.000 & 0.966 & 0.964 & 0.988 & 0.940 & 0.966 \\ & 0.7 & 0.986 & 0.953 & 0.954 & 0.967 & 0.934 & 0.965 \\ & 0.8 & 0.961 & 0.952 & 0.954 & 0.964 & 0.934 & 0.965 \\ & 0.9 & 0.954 & 0.952 & 0.954 & 0.963 & 0.934 & 0.965 \\ & 1 & 0.953 & 0.952 & 0.954 & 0.963 & 0.934 & 0.965 \\ \hline \hline \end{tabular} } \label{tab:DGP0Inference} \end{table} \begin{table}[h!] \centering \caption{RMSEs for DGP (1)} \scalebox{0.9}{ \begin{tabular}{ c c c c c c c c } \hline \hline $n$ & $c $ & (1, 2) & {(1, 3)} & {(1, 4)} & {{(2, 3)}} & {(2, 4)} & {{(3, 4)}} \\ \hline \multirow{10}{*}{1230} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.940 & 0.332 & 0.000 & 0.204 & 3.067 \\ & 0.7 & 0.000 & 2.751 & 1.843 & 3.642 & 0.877 & 8.122 \\ & 0.8 & 0.000 & 7.177 & 3.444 & 11.820 & 2.392 & 15.625 \\ & 0.9 & 9.110 & 12.447 & 4.930 & 23.042 & 4.150 & 23.877 \\ & 1 & 21.949 & 19.690 & 6.233 & 32.068 & 6.184 & 29.984 \\ \hline \multirow{10}{*}{2460} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.000 & 0.421 & 0.000 & 0.000 & 0.681 \\ & 0.7 & 0.000 & 1.996 & 1.431 & 1.498 & 0.734 & 4.616 \\ & 0.8 & 2.945 & 3.824 & 3.467 & 7.503 & 1.616 & 10.312 \\ & 0.9 & 2.945 & 7.404 & 5.354 & 17.387 & 3.253 & 19.648 \\ & 1 & 7.884 & 14.607 & 7.466 & 32.231 & 5.749 & 30.738 \\ \hline \hline \end{tabular} } \label{tab:RMSEDGP1} \end{table} \begin{table}[h!] \centering \caption{RMSEs for DGP (2)} \scalebox{0.9}{ \begin{tabular}{ c c c c c c c c } \hline \hline $n$ & $c $ & (1, 2) & {(1, 3)} & {(1, 4)} & {{(2, 3)}} & {(2, 4)} & {{(3, 4)}} \\ \hline \multirow{10}{*}{1230} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.000 & 0.492 & 0.000 & 0.000 & 2.559 \\ & 0.7 & 0.000 & 0.000 & 6.044 & 0.000 & 2.065 & 4.397 \\ & 0.8 & 0.000 & 1.514 & 13.078 & 4.721 & 5.465 & 9.519 \\ & 0.9 & 0.000 & 3.599 & 20.644 & 14.404 & 9.314 & 14.755 \\ & 1 & 3.964 & 6.682 & 26.301 & 23.569 & 14.271 & 20.772 \\ \hline \multirow{10}{*}{2460} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.000 & 0.159 & 0.000 & 0.000 & 0.000 \\ & 0.7 & 0.000 & 0.000 & 6.373 & 0.000 & 0.989 & 2.103 \\ & 0.8 & 0.000 & 0.000 & 15.145 & 2.702 & 3.455 & 6.309 \\ & 0.9 & 0.000 & 0.000 & 22.974 & 5.737 & 7.433 & 10.970 \\ & 1 & 0.000 & 0.000 & 28.708 & 13.680 & 10.664 & 16.494 \\ \hline \hline \end{tabular} } \label{tab:RMSEDGP2} \end{table} \begin{table}[h!] \centering \caption{RMSEs for DGP (3)} \scalebox{0.9}{ \begin{tabular}{ c c c c c c c c } \hline \hline $n$ & $c $ & (1, 2) & {(1, 3)} & {(1, 4)} & {{(2, 3)}} & {(2, 4)} & {{(3, 4)}} \\ \hline \multirow{10}{*}{1230} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.333 \\ & 0.7 & 0.000 & 0.307 & 0.366 & 0.384 & 0.000 & 1.571 \\ & 0.8 & 0.000 & 0.928 & 0.921 & 3.482 & 0.149 & 4.226 \\ & 0.9 & 0.000 & 2.612 & 1.429 & 7.000 & 0.460 & 6.845 \\ & 1 & 0.340 & 4.145 & 2.119 & 10.064 & 1.239 & 8.501 \\ \hline \multirow{10}{*}{2460} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.7 & 0.000 & 0.000 & 0.045 & 0.000 & 0.000 & 0.022 \\ & 0.8 & 0.000 & 0.000 & 0.262 & 0.175 & 0.000 & 0.631 \\ & 0.9 & 0.000 & 0.194 & 0.668 & 2.536 & 0.000 & 3.089 \\ & 1 & 0.000 & 1.160 & 1.327 & 5.498 & 0.288 & 6.092 \\ \hline \hline \end{tabular} } \label{tab:RMSEDGP3} \end{table} \begin{table}[h!] \centering \caption{RMSEs for DGP (4)} \scalebox{0.9}{ \begin{tabular}{ c c c c c c c c } \hline \hline $n$ & $c $ & (1, 2) & {(1, 3)} & {(1, 4)} & {{(2, 3)}} & {(2, 4)} & {{(3, 4)}} \\ \hline \multirow{10}{*}{1230} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 1.976 \\ & 0.7 & 0.000 & 0.024 & 0.035 & 2.166 & 0.304 & 5.750 \\ & 0.8 & 1.301 & 1.212 & 1.118 & 7.129 & 1.119 & 9.226 \\ & 0.9 & 4.496 & 3.836 & 2.472 & 10.658 & 2.444 & 10.591 \\ & 1 & 8.876 & 5.881 & 3.687 & 12.397 & 3.532 & 11.038 \\ \hline \multirow{10}{*}{2460} & 0.1 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.2 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.3 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.4 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.5 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 \\ & 0.6 & 0.000 & 0.000 & 0.000 & 0.000 & 0.000 & 1.407 \\ & 0.7 & 0.000 & 0.000 & 0.000 & 1.994 & 0.000 & 5.094 \\ & 0.8 & 0.000 & 0.398 & 0.786 & 6.534 & 0.755 & 8.282 \\ & 0.9 & 1.803 & 1.366 & 1.715 & 10.912 & 1.790 & 9.792 \\ & 1 & 5.893 & 3.564 & 3.012 & 12.921 & 2.941 & 10.315 \\ \hline \hline \end{tabular} } \label{tab:RMSEDGP4} \end{table} \newpage \subsection{Empirical Results}\label{sec.application} In this section, we apply VSIV to estimate the returns of college education. We choose the tuning parameter $\tau_n^g$ using formula \eqref{eq:tuning_parameter} in Section \ref{sec.validity set estimation and choice of tau}. The tuning parameter $\tau_n^g$ depends on the user-specified constant $c$. Table \ref{tab:Application_FullSample} presents the estimated validity pair set for a grid of values for $c$. The simulation results in the previous section suggest choosing $c=0.6$. For this choice of $c$, only the pairs $(2,4)$ and $(3,4)$ are selected, while the other pairs are screened out, so that the estimated validity pair set is $\widehat{\mathscr{Z}}_0=\{(2,4),(3,4)\}$.\footnote{In Appendix \ref{sec.validity set estimation using tests}, we present the results of a standard IV validity test.} This result is consistent with the results in \citet{kedagni2020discordant} in that for small values of $Z$, the exclusion condition may fail. The resulting VSIV estimates reported in the last row of Table \ref{tab:Application_FullSample} are $\widehat{\beta}^1_{(2,4)}=0.542$ and $\widehat{\beta}^1_{(3,4)}=0.652$. The corresponding $95%$ confidence intervals based on heteroskedasticity-robust standard errors are $[0.38493,0.69954]$ and $[0.28062,1.0233]$, respectively. \begin{table}[h!] \centering \caption{Validity Pair Set Estimation in Application} \scalebox{1}{ \begin{tabular}{ c c c c c c c } \hline \hline {$c$}& (1, 2) & {{(1, 3)}} & {{(1, 4)}} & (2, 3) & (2, 4) & {{(3, 4)}} \\ \hline 0.1 &\ding{55}&\ding{55}&\ding{55}&\ding{55}&\ding{55}&\ding{55}\\ 0.2 &\ding{55}&\ding{55}&\ding{55}&\ding{55}&\ding{55}&\ding{55}\\ 0.3 &\ding{55}&\ding{55}&\ding{55}&\ding{55}&\ding{55}&\ding{55}\\ 0.4 &\ding{55}&\ding{55}&\ding{55}&\ding{55}&\ding{55}&\ding{55}\\ 0.5 &\ding{55}&\ding{55}&\ding{55}&\ding{55}&\ding{55}&\ding{55}\\ \textbf{\textcolor{orange}{0.6}} &\ding{55}&\ding{55}&\ding{55}&\ding{55}&\ding{51}&\ding{51}\\ 0.7 &\ding{55}&\ding{51}&\ding{55}&\ding{51}&\ding{51}&\ding{51}\\ 0.8 &\ding{55}&\ding{51}&\ding{51}&\ding{51}&\ding{51}&\ding{51}\\ 0.9 &\ding{55}&\ding{51}&\ding{51}&\ding{51}&\ding{51}&\ding{51}\\ 1&\ding{51}&\ding{51}&\ding{51}&\ding{51}&\ding{51}&\ding{51}\\ \hline $\widehat{\beta}^1_{(k,k')}$& -- & -- & -- & -- &0.542 & 0.652\\ \hline \hline \end{tabular} } \label{tab:Application_FullSample} \end{table} \section{Conclusion}\label{sec:conclusion} We propose an approach for estimating LATEs when the instruments are partially invalid. We focus on settings where the instruments are pairwise valid. Under pairwise validity, there are two types of instrument value pairs: Pairs for which the LATE assumptions hold and pairs for which the LATE assumptions fail due to violations of exclusion, independence, monotonicity, or combinations thereof. Pairwise validity is a natural starting point and useful in applications in which it is difficult to determine which LATE assumptions fail and how. However, in settings where researchers have information about how exactly the LATE assumptions fail, it would be interesting to consider intermediate cases that incorporate such information. Under restrictions on which LATE assumptions fail and how, it is often possible to derive nontrivial bounds on LATEs \citep[e.g.,][]{huber2014sensitivity,noack2021sensitivity,kedagni2023identifying,cui2024robust}. Throughout this paper, we focus on average treatment effects for the compliers. However, under pairwise validity, both marginal potential outcome distributions are identified for each $(z,z')\in \mathscr{Z}_{\bar{M}}$. Therefore, another interesting direction for future research is to extend VSIV to allow for the estimation of distributional treatment effects for the compliers, such as local quantile treatment effects abadie2002instrumental,frolich2013unconditional,melly2017local. \section*{Acknowledgements} We are grateful to the Editor (Elie Tamer), Associate Editor, two anonymous referees, Zheng Fang, Martin Huber, Toru Kitagawa, Julian Martinez-Iriarte, D\'esir\'e K\'edagni, Ismael Mourifi\'e, Xiaoxia Shi, and all seminar and conference participants for their insightful suggestions and comments. We thank D\'esir\'e K\'edagni for sharing the data for the empirical application with us. Sun acknowledges funding by the National Natural Science Foundation of China [Grant Number 72103004]. W\"uthrich is also affiliated with CESifo. The usual disclaimer applies. \putbib
bibunit