EconBase
← Back to paper

Instrumental variables with unordered treatments: Theory and evidence from returns to fields of study

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

111,333 characters · 0 sections · 39 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Instrumental variables with unordered treatments: Theory and evidence from returns to fields of study

\global\long\global\long\def\d#1{\mathbbm{1}_{[#1]}}

abstract\begin{singlespace} We revisit the identification argument of kirkeboen_field_2016 who showed how one may combine instruments for multiple unordered treatments with information about individuals\textquoteright ranking of these treatments to achieve identification while allowing for both observed and unobserved heterogeneity in treatment effects. We show that the key assumptions underlying their identification argument have testable implications. We also provide a new characterization of the bias that may arise if these assumptions are violated. Taken together, these results allow researchers not only to test the underlying assumptions, but also to argue whether the bias from violation of these assumptions are likely to be economically meaningful. Guided and motivated by these results, we estimate and compare the earnings payoffs to post-secondary fields of study in Norway and Denmark. In each country, we apply the identification argument of kirkeboen_field_2016 to data on individuals' ranking of fields of study and field-specific instruments from discontinuities in the admission systems. We empirically examine whether and why the payoffs to fields of study differ across the two countries. We find strong cross-country correlation in the payoffs to fields of study, especially after removing fields with violations of the assumptions underlying the identification argument. \end{singlespace}

\thispagestyle{empty}

\setcounter{page}{1}

onehalfspace\@startsection {section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus.2ex} {\reset@font}{Introduction}

Instrumental variables (IV) estimation of treatment effects is challenging if there are multiple unordered treatments. Not only does identification require (at least) one instrument per alternative, but it is also necessary to deal with the issue that individuals who choose the same treatment may have different next-best treatments. One way to resolve this issue is to assume homogeneous treatment effects. If effects are heterogeneous across individuals (conditional on observable characteristics), then standard 2SLS does not identify the payoff to any individual or group of the population from choosing one treatment instead of another.\footnote{A number of studies in diverse fields report evidence of unobserved heterogeneity in causal effects (see, for example, the review article by mogstad2018identification).}

We revisit the identification argument of kirkeboen_field_2016 who showed how one may combine instruments for multiple unordered treatments with information about individuals\textquoteright ranking of these treatment to achieve identification while allowing for both observed and unobserved heterogeneity in treatment effects.\footnote{kirkeboen_field_2016 contributes to a larger literature on identification of treatment effects in unordered choice models. heckman2006understanding and heckman2010comparing discuss the challenges associated with the identification and interpretation of treatment effects in such models. See also the recent work by kamat2017identification and lee2020filtered.} We show that the key assumptions underlying their identification argument have testable implications. We also provide a new characterization of the bias that may arise if these assumptions are violated.\footnote{Throughout the paper, we use the term bias to describe the difference between two population quantities, namely the IV estimand and the parameter of interest, that is the positively weighted average of treatment effects for some complier group.} Taken together, these results allow researchers not only to test the underlying assumptions, but also to argue whether the bias from violation of these assumptions are likely to be economically meaningful. Guided and motivated by these results, we estimate and compare the earnings payoffs to post-secondary fields of study in Norway and Denmark.\footnote{There is a growing body of work on the payoffs to field of study or college major, reviewed in altonji2012heterogeneity,altonji2015analysis, and kirkeboen_field_2016. The latter study also reports IV estimates of the payoffs to fields of study from Norway. Thus, our empirical contribution is the new payoff estimates from Denmark, the examination of the IV assumptions, and the comparison of payoffs to fields of study between Norway and Denmark.} In each country, we apply the identification argument of kirkeboen_field_2016 to data on individuals' ranking of fields of study and field-specific instruments from discontinuities in the admission systems. We then empirically examine the extent to which and why the payoffs to fields of study differ across the two countries.

In Section 2, we begin by briefly reviewing IV in settings with multiple unordered treatments, laying the groundwork for our analysis. As in the analysis of binary treatments in imbens_identification_1994, we allow for heterogeneous effects and assume that each instrument is exogenous and satisfies a monotonicity condition. Our point of departure is the key result in kirkeboen_field_2016: IV can then be used to identify local average treatment effects (LATEs) of unordered treatments under the additional assumptions that the analyst observes individuals' next-best alternatives and an irrelevance condition on preferences.

The next two sections of the paper examine whether the additional assumptions of kirkeboen_field_2016 have testable implications and the bias that may arise if they are violated. To do so, it is useful to stratify the population into a set of instrument-dependent groups, sometimes referred to as principal strata. These groups are defined by the manner in which members of the population react to the instruments. In addition to the usual compliers, always takers, and never takers of imbens_identification_1994, there are two so-called defier groups (both of which are distinct from the usual defier group that exists if the usual monotonicity condition fails). The first is the next-best defiers. In the context of our empirical application, this group consists of individuals who would choose their preferred field if above the admission cutoff, but otherwise choose fields other than the stated next-best alternative. The others are the irrelevance-defiers. In our context, the irrelevance assumption means that if crossing the admission cutoff to a given field does not make an individual choose that field, it should not affect her choice of other fields either.

In Section 3, we use this stratification of the population to characterize the bias in the IV estimands that may arise in the presence of next-best defiers, or irrelevance defiers, or both. It is useful to observe that the bias due to each type of defier has a product structure: It depends on the number of defiers compared to compliers, multiplied by the difference between compliers and defiers in the average payoff to choosing one type of education compared to another. Thus, there will be zero bias if there either are no defiers or if the average payoff to choosing one type of education compared to another is the same for defiers and compliers. Furthermore, the bias becomes large only if there are many defiers relative to compliers and there are large differences in the payoff between compliers and defiers.

In Section 4, we show that the shares of next-best and irrelevance defiers can be bounded, but not point identified. We derive sharp bounds -- which are nontrivial -- and, thus, provides testable implications of the additional assumptions of kirkeboen_field_2016. We show that these results have implications for the recent work of nibbering_clustered_2022 who propose an algorithm which aggregate fields into clusters based on estimated first-stage coefficients. The motivation for their approach is to avoid bias from irrelevance and next-best defiers. We show that their approach requires point identification of the shares of next-best and irrelevance defiers, and that it may produce biased estimates even if effects are constant across individuals (in contrast to standard 2SLS).

The last three sections of the paper take the theoretical results discussed above to the data by comparing payoff estimates for two countries, Denmark and Norway. These are two geographically and culturally close open-economies with very comparable educational institutions as well as similar tax, welfare and social benefits systems. It seems therefore natural to expect that payoffs will -- at least to a degree -- be aligned, and that differences can potentially be understood in light of violations of the assumptions underlying approach outlined above and detailed below.

In Section 5, we present the institutional background and data sources in Denmark and Norway. This section highlights the common institutional framework and data sources, documents how educational classifications and outcomes are harmonized across countries, and discusses differences that may be consequential for the analysis and results. In Section 6, we present the empirical specification that generates the payoff estimates for the two countries, following closely kirkeboen_field_2016. Two challenges that must be met when comparing estimates from two different populations are the reference population and measurement error. Section 6 therefore also defines the population of compliers that we use to anchor the estimates, and presents an error-in-variables approach that addresses bias arising from measurement error when comparing the noisy payoff estimates.

In Section 7, we present the estimation results. We first turn to the first-stages and, building on the results from Section 4, document that in both countries the violations of irrelevance or next-best are non-trivial and appear to be of similar magnitude but of a different nature. In Norway, there is clear evidence of violations of next-best, but little if any sign of violations of irrelevance; in Denmark, the two types of violations seem to be equally frequent. The accompanying second-stages (with earnings measured eight years after application) show that, on average, the estimated annual payoff to completing a field-of-study instead of the next-best is about 2,200 USD in Denmark, while in Norway the payoff estimates are substantially larger and around 22,000 USD. The payoffs in Norway also exhibit a higher variance than in Denmark, and overall we strongly reject that the payoffs are the same. Despite these differences in levels and variation the payoffs significantly co-vary, and we estimate a correlation coefficient of 0.65 for our population of interest.

This correlation substantially increases when we exclude the estimates with the most violations of the irrelevance and next-best assumptions. However, violations of irrelevance and next-best do not appear to explain the lower level and variation of the payoffs in Denmark compared to Norway. Additional exploratory analyses show that these across country differences are mostly driven by heterogeneity in next-best fields, and can partly be explained by differences in selectivity (as measured by students' high school test scores).

\@startsection {section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus.2ex} {\reset@font}{IV with multiple unordered treatments}

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{Models and assumptions}

We assume individuals choose between three mutually exclusive and collectively exhaustive alternatives $d\in\{0,1,2\}$. To fix ideas we envision these as enrolling in three different fields of study. We suppress the individual index and abstract from control variables. We want to interpret IV estimates of the equation

equation[equation omitted — 100 chars of source]

where $y$ is an observed outcome such as earnings, and $d_{j}\equiv\d{d=j}$ is a treatment indicator. Without loss of generality we choose field 0 as reference field, so that $\beta_{1}^{IV}$ ($\beta_{2}^{IV}$) is the payoff from choosing field 1 (2) over field 0.

We suppose individuals are randomly assigned to one of three mutually exclusive and collectively exhaustive groups $Z\in\{0,1,2\}$ and let $z_{j}=\d{Z=j}$ be an indicator variable that equals 1 if an individual is assigned to group $j$ and 0 otherwise. The indicator $z_{j}$ can be thought of as an instrument shifting the costs or benefits of choosing field $j$. For each individual, this gives three potential field choices $d^{z}$ and nine potential outcomes $y^{d,z}$. We let $\mathbf{d}$ denote the column vector of treatment indicators and $\mathbf{z}$ the column vector of instruments. We define $d_{j}^{z}\equiv\d{d^{z}=j}$ to be an indicator variable that tells us whether an individual would choose field $j$ for a given value of $Z$.

As in the analysis of binary treatments in imbens_identification_1994, we allow for heterogeneous effects and assume that each instrument satisfies the following assumptions:

assumptionIV Assumptions \begin{enumerate}[label={(\alph*)}, ref={1(\alph*)}] • Exclusion: $y^{d,z}=y^{d}$ for all $d,z$ • Independence: $y^{d},d^{z}\perp Z$ for all $d,z$\emph{Rank:} $E[\mathbf{z}\mathbf{d}^{\top}]$ has full rank • \textbf{\emph{Monotonicity:}} $d_{k}^{k}\geq d_{k}^{k'}$ for each field assignment pair $k,k'$ \end{enumerate}

Given our notation and assumptions, we can link the observed and potential outcomes and choices as follows,

align[align omitted — 173 chars of source]

These equations represent a model with multiple unordered treatment that permits unrestricted unobserved heterogeneity in treatment effects. Extending the model (and our theoretical results) to more than three choice alternatives is straightforward.

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{Principal strata}

In Table (ref) we invoke assumptions (ref)--(ref) and characterize the principal strata, that is, the groups of individuals defined by how their potential field choices depend on the instrument. The table considers the (sub)population with field 0 as the stated next-best alternative. For brevity, it does not include always takers of field 1 (2) (those who chose field 1 (2) irrespective of instrument value) and never takers of field 1 (2) (those who choose field 2 and 0 (1 and 0) irrespective of instrument value).

As shown in the table, there are two types of compliers, $C_{1}$ and $C_{2}$. The $C_{1}$ ($C_{2}$) compliers are individuals who choose field 1 (2) when the instrument takes value 1 (2), and the reference field 0 when the instrument takes the value 0. In addition, there are four types of defiers, irrelevance and next-best defiers of instruments 1 and 2. Irrelevance defiers $ID_{1}$ ($ID_{2}$) are individuals who choose field 2 (1) when the instrument takes value 1 (2) while choosing field 0 if the instrument takes value 0. Next-best defiers $ND_{1}$ ($ND_{2}$) are individuals who choose field 2 (1) when the instrument takes value 0 while choosing field 1 (2) if the instrument takes value 1 (2).

table[table omitted — 2,105 chars of source]

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{Identification result}

kirkeboen_field_2016 suggest the following assumptions on the groups in Table (ref) to obtain identification:\footnote{kirkeboen_field_2016 are imprecise about whether auxiliary assumption (ref) is imposed on everyone or only those individuals whose treatment status depends on the instrument. However, this is immaterial for their results, as well as ours. The reason is that always takers and never takers drop out of the IV estimand because their treatment status does not change with the instrument.}

assumptionAuxiliary Assumptions \begin{enumerate}[label={(\alph*)}, ref={2(\alph*)}] • Irrelevance: $d_{k}^{k}-d_{k}^{0}=0\implies d_{k'}^{k}=d_{k'}^{0}$ for all pairs $k,k'$ • Next-best: We are able to condition on $d_{1}^{0}=d_{2}^{0}=0$ i.e. $d_{0}^{0}=1$. \end{enumerate}

The irrelevance condition assumes that if changing z from 0 to 1 (2) does not induce an individual to choose treatment 1 (2), then it does not make her choose treatment 2 (1) either. In our context, for example, this assumption means that if crossing the admission cutoff to field 1 does not make an individual choose field 1, it does not make her choose field 2 either. The next-best alternative condition is effectively assuming that individuals' stated next-best alternative is their actual next-best alternative. The following lemma is immediate from these two assumptions:

lemSuppose Assumptions (ref)--(ref) hold. Then $\beta_{1}^{IV},\beta_{2}^{IV}$ have a causal interpretation as positively weighted averages of treatment effects for compliers, and \begin{align*} \beta_{1}^{IV} & =\mathbb{E}[y^{1}-y^{0}\mid C_{1}]\\ \beta_{2}^{IV} & =\mathbb{E}[y^{2}-y^{0}\mid C_{2}] \end{align*}
proofFor a proof, see kirkeboen_field_2016.

The core of Lemma (ref) is that the IV estimand of $\beta_{1}$ ($\beta_{2}$) can be given an interpretation as a local average treatment effect (LATE) of an instrument-induced shift from field 0 to field 1 (2) for compliers when irrelevance and next-best defiers are assumed away.

\@startsection {section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus.2ex} {\reset@font}{Interpretation of IV estimands if auxiliary assumptions fail}

If Assumptions (ref)--(ref) do not hold, the IV estimand of $\beta_{1}$ ($\beta_{2}$) does not have a causal interpretation as a positively weighted average of treatment effects of choosing field 1 (2) over field 0. In the following, we characterize the bias that will occur in this case, and discuss in which situations the bias will be large and small.

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{Assuming only next-best}

The IV estimands of $\beta_{1}$ and $\beta_{2}$ can be decomposed into a LATE for compliers and a bias term using IV moment conditions. In particular, if only next-best holds, but not irrelevance, we get the following decomposition, as shown in Appendix (ref).

propSuppose Assumptions (ref)--(ref) and (ref) hold. Then $\beta_{1}^{IV},\beta_{2}^{IV}$ do not have a causal interpretation as positively weighted averages of treatment effects for compliers, \begin{align} \beta_{1}^{IV}=\underbrace{\mathbb{E}[y^{1}-y^{0}\mid C_{1}]}_{\substack{A} }\quad & +\quad\underbrace{\frac{P(ID_{1})P(ID_{2})}{W'}}_{\substack{\omega_{1}} }\times\enskip(\underbrace{\mathbb{E}[y^{1}-y^{0}\mid C_{1}]-\mathbb{E}[y^{1}-y^{0}\mid ID_{2}]}_{\substack{\Delta_{1}} })\\[1pt] & -\quad\underbrace{\frac{P(ID_{1})P(C_{2})}{W'}}_{\substack{\omega_{2}} }\enskip\times\enskip(\underbrace{\mathbb{E}[y^{2}-y^{0}\mid C_{2}]-\mathbb{E}[y^{2}-y^{0}\mid ID_{1}]}_{\substack{\Delta_{2}} })\nonumber \end{align} where $W'=P(C_{1})P(C_{2})-P(ID_{1})P(ID_{2})$ and the expression for $\beta_{2}^{IV}$ follows by symmetry. A is the complier LATE, $\omega_{1}$ and $\omega_{2}$ are defier group weights, and $\Delta_{1}$ and $\Delta_{2}$ are differences in the causal effects between compliers and irrelevance defiers.
proofSee appendix (ref).

Imposing the constant effects assumption implies that the differences in the causal effects between defier groups ($\Delta_{1}$, $\Delta_{2}$) go to zero. In this case, $\beta_{1}^{IV}$ ($\beta_{2}^{IV}$) would recover the causal effect, $\mathbb{E}[y^{1}-y^{0}]$ ($\mathbb{E}[y^{2}-y^{0}]$). Imposing irrelevance implies that the defier weights ($\omega_{1}$, $\omega_{2}$) go to zero. In this case, $\beta_{1}^{IV}$ ($\beta_{2}^{IV}$) would recover the complier LATE, $\mathbb{E}[y^{1}-y^{0}\mid C_{1}]$ ($\mathbb{E}[y^{2}-y^{0}\mid C_{2}]$).

A central question for empirical researchers is when the bias in Proposition (ref) is likely to be large. To answer this question, it is useful to observe that the two bias terms in equation (ref) are the products of a difference in causal effects and a defier weight consisting of the product of the propensities of irrelevance defiers divided by the difference between complier and defier propensity products.

Note that as long as $P(C_{1})P(C_{2})>2\times P(ID_{1})P(ID_{2})$ the weight $\omega_{1}$ is below 1. This will occur when there are many compliers relative to defiers. When the weight is below 1, the bias will always be smaller than the difference in causal effects. Due to the product structure ($\omega_{j}\times\Delta_{j}$) the bias due to violations of the irrelevance assumption will be very small when both $\omega_{j}$ and $\Delta_{j}$ are small. Conversely, in order for a large bias to occur, there needs to be both many defiers relative to compliers and a large difference in causal effects between the compliers and the irrelevance defiers.

We illustrate this with two examples. In both examples, we fix the LATE for compliers at \$1000. We focus on the first instrument, fixing the propensities of compliers and irrelevance defiers of instrument 2 to $P(ID_{2})=0.2$ and $P(C_{2})=0.8$, and, for simplicity, assume no always takers or never takers for any of the instruments, such that $P(C_{1})=1-P(ID_{1})$.

In Figure (ref) we show how the bias from the first term varies with the propensity of irrelevance defiers. We let the difference in causal effects between compliers and instrument 2-defiers be fixed at three different levels: 10%, 20% and 50% of the complier LATE. In Figure (ref) we show the bias from the first term when varying the difference in causal effects between compliers and defiers. We let the propensity of irrelevance defiers be fixed at three different levels: low (0.1), middle (0.2), and high (0.5). The key take away is that the bias will be small even when there is a sizable number of defiers and a nontrivial difference in causal effects between the compliers and the defiers.

figure[figure omitted — 1,053 chars of source]

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{Assuming only irrelevance}

If irrelevance holds, but next-best is not observed, we may decompose the IV estimand into a complier LATE and a bias term.

propSuppose Assumptions (ref)--(ref) and (ref) hold. Then $\beta_{1}^{IV},\beta_{2}^{IV}$ do not have a causal interpretation as positively weighted averages of treatment effects for compliers, \begin{align} \beta_{1}^{IV}=\underbrace{\mathbb{E}[y^{1}-y^{0}\mid C_{1}]}_{\substack{A} }\quad & +\quad\underbrace{\frac{P(ND_{1})P(C_{2})}{\hat{W}}}_{\substack{\omega_{3}} }\quad\times\enskip(\underbrace{\mathbb{E}[y^{1}-y^{0}\mid ND_{1}]-\mathbb{E}[y^{1}-y^{0}\mid C_{1}]}_{\substack{\Delta_{3}} })\\[4pt] & -\quad\underbrace{\frac{P(ND_{1})P(C_{2})}{\hat{W}}}_{\substack{\omega_{4}} }\quad\times\enskip(\underbrace{\mathbb{E}[y^{2}-y^{0}\mid ND_{1}]-\mathbb{E}[y^{2}-y^{0}\mid C_{2}]}_{\substack{\Delta_{4}} })\nonumber \\[1pt] & +\quad\underbrace{\frac{P(ND_{1})P(ND_{2})}{\hat{W}}}_{\substack{\omega_{5}} }\times\enskip(\underbrace{\mathbb{E}[y^{1}-y^{0}\mid ND_{1}]-\mathbb{E}[y^{1}-y^{0}\mid ND_{2}]}_{\substack{\Delta_{5}} })\nonumber \\[1pt] & -\quad\underbrace{\frac{P(ND_{1})P(ND_{2})}{\hat{W}}}_{\substack{\omega_{6}} }\times\enskip(\underbrace{\mathbb{E}[y^{2}-y^{0}\mid ND_{1}]-\mathbb{E}[y^{2}-y^{0}\mid ND_{2}]}_{\substack{\Delta_{6}} })\nonumber \end{align} where $\hat{W}=P(C_{1})P(C_{2})+P(C_{1})P(ND_{2})+P(ND_{1})P(C_{2})$ and the expression for $\beta_{2}^{IV}$ follows by symmetry. A is the complier LATE, $\omega_{3}$ through $\omega_{6}$ are defier group weights and $\Delta_{3}$ through $\Delta_{6}$ are differences in the causal effects between complier and defier groups.
proofSee appendix (ref).

Imposing the constant effects assumption implies that the differences in causal effects between defier groups ($\Delta_{3}$ through $\Delta_{6}$) go to zero. In this case, $\beta_{1}^{IV}$ ($\beta_{2}^{IV}$) would recover the causal effect, $\mathbb{E}[y^{1}-y^{0}]$ ($\mathbb{E}[y^{2}-y^{0}]$). Observing the next-best alternative implies that the defier weights ($\omega_{3}$ through $\omega_{6}$) go to zero. In this case, $\beta_{1}^{IV}$ ($\beta_{2}^{IV}$) would recover the complier LATE, $\mathbb{E}[y^{1}-y^{0}\mid C_{1}]$ ($\mathbb{E}[y^{2}-y^{0}\mid C_{2}]$).

As in equation ((ref)), the bias terms in equation ((ref)) are the products of a difference in causal effects and a weight consisting of the product of the propensities of defiers divided by the sum of complier and defier propensity products.

Note that the weight in the first and second terms of equation ((ref)) ($\omega_{3}$, $\omega_{4}$) are below 1, but that the weights for the two latter terms ($\omega_{5}$, $\omega_{6}$) can be above 1 if $P(ND_{1})P(ND_{2})>P(C_{1})P(C_{2})+P(C_{1})P(ND_{2})+P(ND_{1})P(C_{2})$. When the weight is below 1, the bias from the term will always be smaller than the difference in causal effects. Due to the product structure ($\omega_{j}\times\Delta_{j}$) the bias due to violations of the next-best assumption will be very small when both $\omega_{j}$ and $\Delta_{j}$ are small. Conversely, in order for a large bias to occur, we need both many defiers relative to compliers and a large difference in causal effects between the different groups.

We keep the same numerical example as in Section (ref) and focus on the term $\omega_{3}\times\Delta_{3}$. In Figure (ref) we show how the bias from this term varies with the propensity of next-best defiers. We let the difference in causal effects between compliers and defiers be fixed at three different levels: at 10%, 20% and 50% of the complier LATE. In Figure (ref) we show the bias when varying the difference in causal effects between compliers and defiers. We let the propensity of next-best defiers be fixed at three different levels: low (0.1), middle (0.2), and high (0.5). The key take away is as above that the bias will be small even when there is a sizable number of defiers and a nontrivial difference in causal effects between the compliers and the defiers.

figure[figure omitted — 1,066 chars of source]

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{Assuming neither irrelevance nor next-best}

If one neither makes the irrelevance assumption nor the next-best assumption, the IV estimand becomes the sum of the complier LATE, all bias terms from Propositions (ref) and (ref), as well as a third set of interacted bias terms.

propSuppose Assumptions (ref)--(ref) holds. Then $\beta_{1}^{IV},\beta_{2}^{IV}$ do not have a causal interpretation as positively weighted averages of treatment effects for compliers, \begin{align} \beta_{1}^{IV}=\underbrace{\mathbb{E}[y^{1}-y^{0}\mid C_{1}]}_{\substack{A} }\enskip & \enskip+\enskip\underbrace{\frac{P(ID_{1})P(ID_{2})}{\bar{W}}}_{\substack{\omega_{1}} }\enskip\enskip\times\enskip(\underbrace{\mathbb{E}[y^{1}-y^{0}\mid C_{1}]-\mathbb{E}[y^{1}-y^{0}\mid ID_{2}]}_{\substack{\Delta_{1}} })\\[1pt] & \enskip-\enskip\underbrace{\frac{P(ID_{1})P(C_{2})}{\bar{W}}}_{\substack{\omega_{2}} }\enskip\quad\times\enskip(\underbrace{\mathbb{E}[y^{2}-y^{0}\mid C_{2}]-\mathbb{E}[y^{2}-y^{0}\mid ID_{1}]}_{\substack{\Delta_{2}} })\nonumber \\[1pt] & \enskip+\enskip\underbrace{\frac{P(ND_{1})P(C_{2})}{\bar{W}}}_{\substack{\omega_{3}} }\quad\times\enskip(\underbrace{\mathbb{E}[y^{1}-y^{0}\mid ND_{1}]-\mathbb{E}[y^{1}-y^{0}\mid C_{1}]}_{\substack{\Delta_{3}} })\nonumber \\[1pt] & \enskip-\enskip\underbrace{\frac{P(ND_{1})P(C_{2})}{\bar{W}}}_{\substack{\omega_{4}} }\quad\times\enskip(\underbrace{\mathbb{E}[y^{2}-y^{0}\mid ND_{1}]-\mathbb{E}[y^{2}-y^{0}\mid C_{2}]}_{\substack{\Delta_{4}} })\nonumber \\[4pt] & \enskip+\enskip\underbrace{\frac{P(ND_{1})P(ND_{2})}{\bar{W}}}_{\substack{\omega_{5}} }\times\enskip(\underbrace{\mathbb{E}[y^{1}-y^{0}\mid ND_{1}]-\mathbb{E}[y^{1}-y^{0}\mid ND_{2}]}_{\substack{\Delta_{5}} })\nonumber \\[1pt] & \enskip-\enskip\underbrace{\frac{P(ND_{1})P(ND_{2})}{\bar{W}}}_{\substack{\omega_{6}} }\times\enskip(\underbrace{\mathbb{E}[y^{2}-y^{0}\mid ND_{1}]-\mathbb{E}[y^{2}-y^{0}\mid ND_{2}]}_{\substack{\Delta_{6}} })\nonumber \\[1pt] & \enskip-\enskip\underbrace{\frac{P(ND_{1})P(ID_{2})}{\bar{W}}}_{\substack{\omega_{7}} }\enskip\times\enskip(\underbrace{\mathbb{E}[y^{1}-y^{0}\mid C_{1}]-\mathbb{E}[y^{1}-y^{0}\mid ID_{2}]}_{\substack{\Delta_{7}} })\nonumber \\[1pt] & \enskip+\enskip\underbrace{\frac{P(ID_{1})P(ND_{2})}{\bar{W}}}_{\substack{\omega_{8}} }\enskip\times\enskip(\underbrace{\mathbb{E}[y^{1}-y^{0}\mid ND_{2}]-\mathbb{E}[y^{1}-y^{0}\mid C_{1}]}_{\substack{\Delta_{8}} })\nonumber \\[1pt] & \enskip-\enskip\underbrace{\frac{P(ID_{1})P(ND_{2})}{\bar{W}}}_{\substack{\omega_{9}} }\enskip\times\enskip(\underbrace{\mathbb{E}[y^{2}-y^{0}\mid ND_{2}]-\mathbb{E}[y^{2}-y^{0}\mid ID_{1}]}_{\substack{\Delta_{9}} })\nonumber \end{align} where \begin{align*} \bar{W}\enskip & =\enskip P(C_{1})P(C_{2})+P(C_{1})P(ND_{2})+P(ND_{1})P(C_{2})\\[1pt] & \enskip+P(ND_{1})P(ID_{2})+P(ID_{1})P(ND_{2})-P(ID_{1})P(ID_{2}) \end{align*} and the expression for $\beta_{2}^{IV}$ follows by symmetry.
proofSee appendix (ref).

A is the complier LATE, $\omega_{1}$ and $\omega_{2}$ are defier weights which also occur when observing the next-best alternative, $\omega_{3}$ through $\omega_{6}$ are defier weights which also occur under irrelevance and $\omega_{7}$ through $\omega_{9}$ are defier weights which occur only when neither assumption holds. $\Delta_{1}$, $\Delta_{2}$ and $\Delta_{7}$ are differences in the causal effects between irrelevance defiers and compliers, $\Delta_{3}$, $\Delta_{4}$ and $\Delta_{8}$ are differences in the causal effects between next-best defiers and compliers, while $\Delta_{5}$ and $\Delta_{6}$ are differences in causal effects between next-best defiers for the two different instruments, and $\Delta_{9}$ is the difference in the causal effects between next-best defiers of instrument 2 and irrelevance defiers of instrument 1.

Imposing the constant effects assumption implies that the differences in causal effects between defier groups ($\Delta_{1}$ through $\Delta_{9}$) go to zero. In this case, $\beta_{1}^{IV}$ ($\beta_{2}^{IV}$) would recover the causal effect, $\mathbb{E}[y^{1}-y^{0}]$ ($\mathbb{E}[y^{2}-y^{0}]$). Imposing the next-best assumption yields the result from Proposition (ref), as weights $\omega_{3}$ through $\omega_{9}$ go to zero. Imposing the irrelevance assumption yields Proposition (ref), as weights $\omega_{1}$, $\omega_{2}$ and $\omega_{7}$ through $\omega_{9}$ go to zero. Imposing both irrelevance and observing the next-best alternative make all defier weights ($\omega_{1}$ through $\omega_{9}$) go to zero. Then $\beta_{1}^{IV}$ ($\beta_{2}^{IV}$) would recover the complier LATE, $\mathbb{E}[y^{1}-y^{0}\mid C_{1}]$ ($\mathbb{E}[y^{2}-y^{0}\mid C_{2}]$).

Note that the bias in Proposition (ref) is the sum of all bias terms from Propositions (ref) and (ref), in addition to three new bias terms (except for a different denominator of the weights). These are terms following from interactions between irrelevance and next-best defiers, and rely on both types of defiers being present and having differences in causal effects between each other and with the complier group. As a result, the bias will be small unless there are relatively many of both types of defiers and the causal effects are materially different between these groups and the compliers.

\@startsection {section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus.2ex} {\reset@font}{Testable implications and aggregation}

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{How to test the auxiliary assumptions}

The first stage equations for the IV estimates of equation ((ref)) are given by:

align[align omitted — 206 chars of source]

We now examine if it is possible to devise a test of whether the auxiliary assumptions (ref)--(ref) hold empirically. To do so, it is useful to characterize the quantities that the first stage coefficients recover:

lemSuppose Assumptions (ref)--(ref) hold. Then \begin{align*} \alpha_{1}^{0} & =P(AT_{1})\equiv P(OT_{2})+P(ND_{2}) & \alpha_{2}^{0} & =P(AT_{2})\equiv P(OT_{1})+P(ND_{1})\\ \alpha_{1}^{1} & =P(C_{1})+P(ND_{1}) & \alpha_{2}^{2} & =P(C_{2})+P(ND_{2})\\ \alpha_{2}^{1} & =P(ID_{1})-P(ND_{1}) & \alpha_{1}^{2} & =P(ID_{2})-P(ND_{2}) \end{align*} where $AT_{1}$($AT_{2}$) are always-takers of field $1$ (2) when $z$ equals 0 or 1 (0 or 2), and $OT_{1}$ ($OT_{2}$) are global (for every value of the instrument) always takers of the other field 2 (1). See Appendix Table (ref) for formal definitions of these instrument-specific strata.
proofSee appendix (ref).

This result paves the way for the main result on the testability of the irrelevance and next best assumptions:

propSuppose Assumptions (ref)--(ref) hold. Then $P(ID_{1})$ and $P(ND_{1})$ are partially identified. \begin{align*} P(ND_{1}) & \in[\max\{0,-\alpha_{2}^{1}\},\qquad\qquad\quad\,\,\min\{\alpha_{1}^{1},\alpha_{2}^{0}\}]\\ P(ID_{1}) & \in[\max\{0,\enskip\,\alpha_{2}^{1}\},\max\{0,\alpha_{2}^{1}+\min\{\alpha_{1}^{1},\alpha_{2}^{0}\}\}] \end{align*} where results for $P(ID_{2})$ and $P(ND_{2})$ follow by symmetry.
proofSee appendix (ref).

The practical implication of Proposition (ref) is that we cannot point identify the defier propensities without further assumptions. Yet, the assumptions are testable as the bounds will generally be nontrivial. Furthermore, if either assumption (ref) or (ref) is known to hold, the other assumption can be tested separately and $P(ID_{1})$ or $P(ND_{1})$ is point identified.

corSuppose Assumptions (ref)--(ref) and (ref) hold. Then $P(ND_{1})=P(ND_{2})=0$ and we can test whether assumption (ref) (irrelevance) holds, as $\alpha_{2}^{1}=P(ID_{1})$ and $\alpha_{1}^{2}=P(ID_{2})$.
corSuppose Assumptions (ref)--(ref) and (ref) hold. Then $P(ID_{1})=P(ID_{2})=0$ and we can test whether assumption (ref) (next-best) holds, as $\alpha_{2}^{1}=-P(ND_{1})$ and $\alpha_{1}^{2}=-P(ND_{2})$.

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{How aggregation may cause violations of the exclusion restriction}

nibbering_clustered_2022 propose an algorithm which aggregates fields into clusters based on estimated first-stage coefficients. The motivation for their approach is to avoid bias from irrelevance and next-best defiers. Before discussing their approach, it is important to observe that the resulting IV estimates between such clusters will, at best, identify a positively weighted average of the causal effects of choosing one field versus a linear combination of the other fields, for example, the effects of choosing field 1 versus field 0 or 2. Hence, this approach involves moving the goalpost from clearly defined field contrasts that govern individuals' educational investments to clusters of different fields. In the discussion below, we accept at faith that such contrasts are parameters of interest.

\@startsection{subsubsection}{3}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {-1em} {\reset@font}{Bias From Exclusion Violation}

We continue to consider the situation with three fields, discussed above. The algorithm takes as a starting point all individuals with a certain reported next-best alternative (in our case taken to be 0), and test the hypothesis that the off-diagonal coefficients, $\alpha_{2}^{1}$ and $\alpha_{1}^{2}$, are zero. If this hypothesis is rejected, the sign of the coefficient is evaluated and the treatments are clustered according to the rules laid out in Table (ref). For example, if $\alpha_{2}^{1}$ is negative and $\alpha_{1}^{2}$ is either zero or positive, fields 0 and 2 become the control cluster and field 1 the treatment cluster. Conversely, if $\alpha_{2}^{1}$ is either zero or positive and $\alpha_{1}^{2}$ is negative, fields 0 and 1 become the control cluster and field 2 the treatment cluster.

table[table omitted — 3,505 chars of source]

After performing the clustering based algorithm, nibbering_clustered_2022 estimate cluster treatment effects: Let $\tilde{d}(d)=\d{d\in S_{1}}$ be the binary cluster treatment indicator and $\tilde{z}(Z)=\d{Z=d\in S_{1}}$ the cluster instrument indicator. The no clustering-scenario is equivalent to the field level. In the two other scenarios (control clustering or treatment clustering) we consider IV estimates of the equation \[ y=\tilde{\beta}_{0}+\tilde{\beta}_{1}\tilde{d}+\varepsilon \] where the first stage is \[ \tilde{d}=\pi_{0}+\pi_{1,0}\tilde{z}+\nu \] and $\pi_{1,0}$ is the first stage coefficient. Observed and potential outcomes and choices are linked as

align[align omitted — 185 chars of source]

where $\tilde{d}^{j}\equiv\d{\tilde{d}^{j}=1}$ denotes the cluster-level potential treatment and $\tilde{y}^{j}$ is the potential outcome in cluster $j$. In Appendix (ref) we show that this IV estimand does not, under Assumptions (ref)--(ref), have a causal interpretation as a positively weighted average of treatment effects for the cluster complier groups. This result is summarized in Proposition (ref).

propSuppose Assumptions (ref)--(ref) hold. \begin{enumerate}[label={(\alph*)}, ref={6(\alph*)}] • Under control clustering, $\tilde{\beta}_{1}^{IV}$ does not have a causal interpretation as a positively weighted average of treatment effects for the cluster complier group. If the clustering is $S_{1}=\{1\}$ and $S_{0}=\{2,0\}$, we have \begin{align*} \tilde{\beta}_{1,0}^{IV}\enskip & =\enskip\underbrace{\frac{P(C_{1}\cup ND_{2})}{\pi_{1,0}}\mathbb{E}[y^{1}-y^{0}\mid C_{1}\cup ND_{2}]+\frac{P(C_{2}\cup ND_{1})}{\pi_{1,0}}\mathbb{E}[y^{1}-y^{2}\mid C_{2}\cup ND_{1}]}_{\substack{A} }\\[4pt] & \qquad+\underbrace{\frac{P(ID_{1})}{\pi_{1,0}}}_{\substack{\tilde{\omega}_{1}} }\enskip\enskip\underbrace{\mathbb{E}[y^{2}-y^{0}\mid ID_{1}]}_{\substack{\tilde{\Delta}_{1}} }-\underbrace{\frac{P(ID_{2})}{\pi_{1,0}}}_{\substack{\tilde{\omega}_{2}} }\enskip\enskip\underbrace{\mathbb{E}[y^{2}-y^{0}\mid ID_{2}]}_{\substack{\tilde{\Delta}_{1}} } \end{align*} where $\pi_{1,0}=P(C_{1}\cup C_{2}\cup ND_{1}\cup ND_{2})$. A is a positively weighted average of cluster complier LATEs, $\tilde{\omega}_{1}$ and $\tilde{\omega}_{2}$ are defier group weights, and $\tilde{\Delta}_{1}$ and $\tilde{\Delta}_{2}$ are differences in potential outcomes for irrelevance defiers in cluster $S_{0}$, i.e. never takers of the clustered treatment. The result for the clustering $S_{1}=\{2\}$ and $S_{0}=\{1,0\}$ is symmetric. • Under treatment clustering, $\tilde{\beta}_{1}^{IV}$ does not have a causal interpretation as a positively weighted average of treatment effects for the cluster complier group. We have \begin{align*} \tilde{\beta}_{1,0}^{IV}\enskip & =\enskip\underbrace{\frac{P(C_{1}\cup ID_{2})}{\pi_{1,0}}\mathbb{E}[y^{1}-y^{0}\mid C_{1}\cup ID_{2}]+\frac{P(C_{2}\cup ID_{1})}{\pi_{1,0}}\mathbb{E}[y^{2}-y^{0}\mid C_{2}\cup ID_{1}]}_{\substack{A} }\\[4pt] & \qquad+\underbrace{\frac{P(ND_{1})}{\pi_{1,0}}}_{\substack{\tilde{\omega}_{3}} }\enskip\underbrace{\mathbb{E}[y^{1}-y^{2}\mid ND_{1}]}_{\substack{\tilde{\Delta}_{3}} }-\underbrace{\frac{P(ND_{2})}{\pi_{1,0}}}_{\substack{\tilde{\omega}_{4}} }\enskip\underbrace{\mathbb{E}[y^{1}-y^{2}\mid ND_{2}]}_{\substack{\tilde{\Delta}_{4}} } \end{align*} where $\pi_{1,0}=P(C_{1}\cup C_{2}\cup ID_{1}\cup ID_{2})$. A is a positively weighted average of cluster complier LATEs, $\tilde{\omega}_{3}$ and $\tilde{\omega}_{4}$ are defier group weights, and $\tilde{\Delta}_{3}$ and $\tilde{\Delta}_{4}$ are differences in potential outcomes for irrelevance defiers in cluster $S_{1}$, i.e. always takers of the clustered treatment. \end{enumerate}
proofSee Appendix (ref).

Imposing the irrelevance assumption under control clustering implies that the defier weights ($\tilde{\omega}_{1},\tilde{\omega}_{2}$) go to zero. In this case, $\tilde{\beta}_{1,0}^{IV}$ recovers a positively weighted average of the causal effect of choosing field 1 over 0 for compliers of instrument 1 and next-best defiers of instrument 2, and of choosing field 1 over 2 for compliers of instrument 2 and next-best defiers of instrument 1, weighted by the number of compliers and defiers. Under control clustering, this is the new parameter of interest.

Imposing the next-best assumption under treatment clustering implies that the defier weights ($\tilde{\omega}_{3},\tilde{\omega}_{4}$) go to zero. In this case, $\tilde{\beta}_{1,0}^{IV}$ recovers a positively weighted average of the causal effect of choosing field 1 over 0 for compliers of instrument 1 and irrelevance defiers of instrument 2, and of choosing field 2 over 0 for compliers of instrument 2 and irrelevance defiers of instrument 1, weighted by the number of compliers and defiers. Under treatment clustering, this is the new parameter of interest.

If neither irrelevance nor next-best assumptions hold, the IV estimand does not have a causal interpretation as a positively weighted average of treatment effects for the cluster complier group. The bias terms reflect that individuals may in response to changes in the cluster instrument be switching across fields in the treatment cluster and/or across fields in the control cluster. Such switches will generally involve changes in potential outcomes, yet no change in the cluster treatment status. Thus, the exclusion restriction at the cluster level will be violated. The reason for this bias is that the algorithm equates the sign of the off-diagonal coefficients with the presence and absence of irrelevance and nex-best defiers. As shown in Lemma (ref), this is wrong. The off-diagonal coefficients tell us only if there are more or less next-best defiers than irrelevance defiers. One cannot in general use the sign of $\alpha_{2}^{1}$ ($\alpha_{1}^{2}$) to show that there are no irrelevance defiers of instrument 1 (2) if $\alpha_{2}^{1}<0$ ($\alpha_{1}^{2}<0$) and no next-best defiers of instrument 1 (2) if $\alpha_{2}^{1}>0$ ($\alpha_{1}^{2}>0$).

It is also important to observe that the constant effects assumption is not sufficient for $\tilde{\beta}_{1,0}^{IV}$ to recover a positively weighted average of treatment effects between clusters 0 and 1 and obtain a causal interpretation. This result is summarized in Proposition (ref).\footnote{One exception to this negative result is the special case in which the number of defiers for each instrument happen to be equal, i.e. that $P(ID_{1})=P(ID_{2})$ under control clustering or $P(ND_{1})=P(ND_{2})$ under treatment clustering.}

propSuppose Assumptions (ref)--(ref) hold and we further assume constant treatment effects. \begin{enumerate}[label={(\alph*)}, ref={6(\alph*)}] • Under control clustering, $\tilde{\beta}_{1}^{IV}$ does not recover the causal effect. If the clustering is $S_{1}=\{1\}$ and $S_{0}=\{2,0\}$, we have \begin{align*} \tilde{\beta}_{1,0}^{IV}\enskip & =\enskip\underbrace{\frac{P(C_{1}\cup ND_{2})}{\pi_{1,0}}\mathbb{E}[y^{1}-y^{0}]+\frac{P(C_{2}\cup ND_{1})}{\pi_{1,0}}\mathbb{E}[y^{1}-y^{2}]}_{\substack{A} }\\[4pt] & \qquad+\underbrace{\frac{P(ID_{1})-P(ID_{2})}{\pi_{1,0}}}_{\substack{\dot{\omega}_{1}} }\enskip\enskip\underbrace{\mathbb{E}[y^{2}-y^{0}]}_{\substack{\dot{\Delta}_{1}} } \end{align*} where $\pi_{1,0}=P(C_{1}\cup C_{2}\cup ND_{1}\cup ND_{2})$. A is a positively weighted average of the causal effects of choosing field 1 over 0 and of choosing field 1 over 2, $\dot{\omega}_{1}$ is a difference between defier group weights, and $\dot{\Delta}_{1}$ is the difference in potential outcomes for irrelevance defiers in cluster $S_{0}$, i.e. never takers of the clustered treatment. The result for the clustering $S_{1}=\{2\}$ and $S_{0}=\{1,0\}$ is symmetric. • Under treatment clustering, $\tilde{\beta}_{1}^{IV}$ does not recover the causal effect. We have \begin{align*} \tilde{\beta}_{1,0}^{IV}\enskip & =\enskip\underbrace{\frac{P(C_{1}\cup ID_{2})}{\pi_{1,0}}\mathbb{E}[y^{1}-y^{0}]+\frac{P(C_{2}\cup ID_{1})}{\pi_{1,0}}\mathbb{E}[y^{2}-y^{0}]}_{\substack{A} }\\[4pt] & \qquad+\underbrace{\frac{P(ND_{1})-P(ND_{2})}{\pi_{1,0}}}_{\substack{\dot{\omega}_{2}} }\enskip\underbrace{\mathbb{E}[y^{1}-y^{2}]}_{\substack{\dot{\Delta}_{2}} } \end{align*} where $\pi_{1,0}=P(C_{1}\cup C_{2}\cup ID_{1}\cup ID_{2})$. A is a positively weighted average of the causal effects of choosing field 1 over 0 and of choosing field 2 over 0, $\dot{\omega}_{2}$ is a difference between defier group weights, and $\dot{\Delta}_{2}$ is the difference in potential outcomes for irrelevance defiers in cluster $S_{1}$, i.e. always takers of the clustered treatment. \end{enumerate}
proofThe constant effects assumption reduces all conditional expectations to unconditional expectations, i.e. $\mathbb{E}[y^{j}-y^{k}\mid G]=\mathbb{E}[y^{j}-y^{k}]$ for any group $G$ and any combination of fields $j,k$. The result is immediate.

In contrast, the approach of kirkeboen_field_2016 recovers the causal effect under the constant effects assumption. This shows that the clustering method relies on different, not weaker assumptions than kirkeboen_field_2016.

The following auxiliary exclusion restriction can be made to obtain identification under the clustering approach.

assumptionCluster Exclusion Assumptions \begin{enumerate}[label={(\alph*)}, ref={3(\alph*)}] • Control Cluster Exclusion: $\tilde{d}^{1}=\tilde{d}^{0}=0\implies\tilde{y}^{0,1}=\tilde{y}^{0,0}$ • Treatment Cluster Exclusion: $\tilde{d}^{1}=\tilde{d}^{0}=1\implies\tilde{y}^{1,1}=\tilde{y}^{1,0}$ \end{enumerate}

Assumptions (ref) and (ref) ensure that the bias from switchers within clusters (irrelevance defiers under control clustering and next-best defiers under treatment clustering) disappear, irrespective of the number of switchers. These assumptions are homogeneity restrictions on potential outcomes across different fields, and, thus, difficult to justify. Nevertheless, if one is willing to invoke Assumptions (ref) and (ref), one may obtain the following identification result:

propUnder control clustering, suppose Assumptions (ref)--(ref) and (ref) hold. $\tilde{\beta}_{1}^{IV}$ has a causal interpretation as the positively weighted average of treatment effects for cluster compliers. If the clustering is $S_{1}=\{1\}$ and $S_{0}=\{2,0\}$, we have \begin{align*} \tilde{\beta}_{1,0}^{IV}\enskip & =\enskip\frac{P(C_{1}\cup ND_{2})}{\pi_{1,0}}\mathbb{E}[y^{1}-y^{0}\mid C_{1}\cup ND_{2}]+\frac{P(C_{2}\cup ND_{1})}{\pi_{1,0}}\mathbb{E}[y^{1}-y^{2}\mid C_{2}\cup ND_{1}] \end{align*} where $\pi_{1,0}=P(C_{1}\cup C_{2}\cup ND_{1}\cup ND_{2})$. The result for clustering $S_{1}=\{2\}$ and $S_{0}=\{1,0\}$ is symmetric. Under treatment clustering, suppose Assumptions (ref)--(ref) and (ref) hold. $\tilde{\beta}_{1}^{IV}$ has a causal interpretation as a positively weighted average of treatment effects for cluster compliers, and \begin{align*} \tilde{\beta}_{1,0}^{IV}\enskip & =\enskip\frac{P(C_{1}\cup ID_{1})}{\pi_{1,0}}\mathbb{E}[y^{1}-y^{0}\mid C_{1}\cup ID_{1}]+\frac{P(C_{2}\cup ID_{1})}{\pi_{1,0}}\mathbb{E}[y^{2}-y^{0}\mid C_{2}\cup ID_{1}] \end{align*} where $\pi_{1,0}=P(C_{1}\cup C_{2}\cup ID_{1}\cup ID_{2})$.
proofAssumption (ref) ((ref)) eliminates the bias terms in the results from Proposition (ref) by letting $\tilde{\Delta}_{1},\tilde{\Delta}_{2}$ ($\tilde{\Delta}_{3},\tilde{\Delta}_{4}$) go to zero. The result is immediate.

\@startsection {section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus.2ex} {\reset@font}{Empirical analysis}

Guided and motivated by the formal results above, we now turn to the empirical analysis of the payoffs to field of study in Norway and Denmark.

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{Institutional settings}

The Danish and Norwegian post-secondary education systems are similar in many respects. Their post-secondary education sectors consist of public universities and a larger number of public and private university colleges. The vast majority of students attend a public institution, and even the private institutions are publicly funded and regulated. Universities all offer a wide selection of fields. By comparison, the university colleges rarely offer fields like Law, Medicine, Science, or Technology, but tend to offer professional degrees in fields like Engineering, Health, Business, and Teaching. Obtaining a post-secondary degree normally requires three to five years; there are no tuition fees; most students receive financial support (in the form of grants/loans) from the state.

The admission process is centralized in both countries. Applications are submitted to a central organization that handles the admission process to universities and university colleges. An applicant ranks programs (up to 15 in Norway and 8 in Denmark), each defined by a detailed field and an institution. The number of slots for each program is effectively determined by each country\textquoteright s ministry of education. For many programs, demand exceeds supply. Most slots in programs with excess demand are filled based on an application score derived from high school GPA. Offers are determined by the applicants' application score: the highest ranked applicant receives an offer for her preferred program; the second highest applicant receives an offer for her highest ranked program among the remaining programs; and so on. This is repeated until either slots run out, or applicants run out. This allocation mechanism corresponds to a so-called serial dictatorship, which is both Pareto efficient and strategy-proof svensson1999strategy and should therefore elicit the applicants\textquoteright true ranking of fields at the time of application.\footnote{A possible threat to strategy-proofness is the truncation of the application list (at 15 programs in Norway and 8 in Denmark) which might induce individuals to list a safe option as their last choice. However, this is likely unimportant in practice, as less than 0.1 percent of Norwegian applicants are offered their 15th choice, and less than 1 percent of Danish applicants list eight programs.} If students want to change field or institution, they usually need to participate in next year's admission process on equal terms with other applicants.\footnote{Most programs in Denmark also have a standby (waiting) list and the GPA threshold for the standby list is typically a little lower than the main threshold. On the application form, applicants can choose whether to apply for the standby list. Applicants admitted to the standby list are guaranteed a study place the following year, but they are not considered for any of the lower-ranked programs on their application. Appendix (ref) provides a more detailed discussion.}

For both countries, the exact thresholds are unpredictable at the time of application. They are not published until after the allocation process, and variation in thresholds over time is considerable. For programs with excess demand, the admission process implies that applicants scoring above a certain threshold are much more likely to receive an offer for a program they prefer compared to applicants with the same program preferences but marginally lower application score. This gives rise to credible instruments from discontinuities that effectively randomize applicants near admission cutoffs into different programs.

As explained in greater detail in kirkeboen_field_2016, the instruments are defined around local course rankings on students' application lists. These local rankings define the \textquotedblleft preferred\textquotedblright and \textquotedblleft next-best\textquotedblright alternatives. For example, consider two fields, A and B, with A having a higher admission cutoff than B. Consider students who rank A just above B and have an application score that is either just below or just above the admission cutoff to A. These students will have A as the preferred field and B as the next-best field, no matter if A and B are ranked at the top, in the middle or at the bottom of the list. In other words, what matters for the relevance of instrument and the definition of preferred and next-best is the local ranking at which individuals are shifted in or out of a program because the application score is slightly above or below the relevant admission cutoff.

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{Data and descriptive statistics}

For each country, we combine several sources of administrative data. For Norway, we use data for all applications to post-secondary education for the years 1998--2004. For Denmark, we use data for all applications to post-secondary education for the years 1994--2002. For both countries we retain the individuals\textquoteright first observed application and exclude those who had a post-secondary degree at the time of application. We link these applicants to the population register and other registers to obtain background information, information on completed field, and annual earnings. In our main analysis we use data on treatment (completed field) and outcome (annual earnings) eight years after application as in kirkeboen_field_2016, and restrict the sample to those who have completed a field within eight years from application. The measure of earnings includes wage income, income from self-employment, and transfers that replace such income like short-term sickness pay and paid parental leave (but excludes unemployment benefits). Earnings are deflated using the CPI with 2011 as base year and converted to 1,000s of US dollars using the average exchange rates for the years 2010--2016 (6.5 Norwegian and 5.9 Danish crowns per US dollar).

We aggregate detailed fields into nine broad fields of study. We essentially follow the same classification of fields as in kirkeboen_field_2016. The only difference is that Technology now covers the integrated and more vocational/professional short and long cycle degrees at university colleges and universities and consist mostly of computer science and engineering degrees. Science corresponds to more open-ended bachelor programs in different sciences such as physics, biology and mathematics, as well as agriculture, forestry and aquaculture. For the main analysis, we retain all applicants who applied to at least two broad fields, where the most preferred field has an admission cutoff. If an applicant applied to several programs within her preferred broad field we use the lowest program cutoff as the effective cutoff to the preferred field.

While the full sample of applicants is of comparable size for the two countries, the final estimation sample is smaller in Denmark than in Norway, primarily because of fewer fields with admission restrictions in Denmark. Figure (ref) shows the distributions of completed field among applicants in Norway and Denmark eight years after applying. While the distributions are similar, there are some notable differences. The share of applicants completing teaching is substantially higher in Denmark than in Norway, as is the share of applicants having completed a degree with Technology. On the flip side the shares in Science, Social Science and Humanities are larger in Norway than in Denmark.\footnote{Some of the cross-country differences may be due to differences in the classification of specific fields into the nine broad fields. For instance, one reason why the share having completed Teaching is large in Denmark is that all individuals having completed a bachelor\textquoteright s degree in social education are included in Teaching regardless of the specialization (e.g., kindergarten teacher, nursery teacher, nursery nurse, child and youth worker, support worker), while some of the specializations could alternatively be classified as Other Health if they were observed as separate educations.}

figure[figure omitted — 431 chars of source]
figure[figure omitted — 656 chars of source]

As an indicator of relative selectivity, we standardize high school GPA within country and show in sub-graph (a) of Figure (ref) the average standardized GPA by field and country. In both countries average GPA is relatively low in Teaching and Other Health. Average GPA is very high for Medicine in both countries, but Law, Social Science and Science are nearly equally selective as Medicine in Denmark.

Sub-graph (b) of Figure (ref) compares earnings by field across country. Average earnings levels, indicated by the red dotted lines, are very similar in the two countries. Earnings in Medicine and Social Science are higher in Denmark consistent with their higher selectivity, but the same is not observed for Law and Science. Earnings are higher in Norway for Technology. In both countries, earnings are particularly low for Humanities. However, it is important to note that earnings are measured eight years after application, which is very early in the career, especially for those choosing longer programs or programs characterized by a more difficult school-to-work transition. We examine the importance of this issue in a specification check that uses earnings measured later in the working life as the outcome variable.

In Appendix Figures (ref) and (ref) we present results similar to Figures (ref) and (ref), but not restricted to applicants. For Norway, data for Figures (ref) and (ref) consist of everybody born 1979--1983 such that we have application data for the years they are aged 19--21. Similarly, for Denmark the population sample consists of the cohorts born 1975--1981. Completed field and earnings are measured at age 28. The results for these broader populations are similar to the results for applicants in Figures (ref) and (ref).

\@startsection {section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus.2ex} {\reset@font}{How we estimate and compare payoffs}

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{2SLS specification }

The identification results in Sections (ref) and (ref) motivate and guide the specification of the empirical model. We consider the following system of equations separately for individuals with next-best field $l$ (in the local field ranking):

align[align omitted — 217 chars of source]

where ((ref)) is the second stage equation, and ((ref)) are the first-stage equations, one for each field. In these equations, $j$ denotes the completed field, $l$ denotes the stated next best alternative (in the local field ranking), and $k$ denotes the preferred field (in the local field ranking). The $l$ index is necessary in these equations, since we are now considering all possible preferred and next best fields (and not only focusing on field 0 as the stated next best alternative, as we did in the simple example in equations ((ref))-((ref))).

The instruments $z_{k}$ in ((ref)) are the predicted offers for field $k$, and $z_{k}$ is therefore equal to one if $k$ is the individual's preferred field and her application score exceeds the admission cutoff for field $k$ and zero otherwise. We therefore have as many binary instruments as treatments (one $z_{k}$ for each completed field dummy $d_{j}$), and for a given individual at most one of the instruments $z_{k}$ can equal 1 (namely the one of her preferred field in the local field ranking).

Our estimation approach exploits the fuzzy regression discontinuity design implicit in the admission process described above, where individuals with application scores above the cutoff are more likely to receive an offer for their preferred field. Although the identification in this setup is ultimately local, we use 2SLS because our sample sizes do not allow for local non-parametric estimation. While the model laid out above abstracted from any control variables, we now need to include certain covariates to ensure the exogeneity of our instruments.

First, all equations include controls for the running variable. While our baseline specification controls for the application score linearly on each side of the admission cutoff, kirkeboen_field_2016 reported results from several specification checks, all of which support our main findings. Second, we control for individuals' preferences by adding fixed effects for preferring field $k$ and having $l$ as the next-best field (in the local field ranking): $\lambda_{l}^{k}$ and $\eta_{jl}^{k}$. To gain precision, we estimate the system of equations ((ref))--((ref)) jointly for all completed and next-best fields, allowing for separate intercepts for preferred field and for next-best field by completed field (i.e. $\lambda_{l}^{k}=\mu^{k}+\theta_{j}$ and $\eta_{jl}^{k}=\tau_{j}^{k}+\sigma_{j}^{k}$). In a robustness check, kirkeboen_field_2016 show that their estimates are robust to allowing for separate intercepts for every interaction between preferred and next-best field. Finally, to reduce residual variance we also add controls for gender, cohort and age at application, which are pre-determined.

From the resulting 2SLS estimation of equations ((ref))--((ref)) across all next-best fields, we obtain a matrix of the payoffs to field $j$ compared to $k$ for those who prefer $j$ and have $k$ as next-best field. In our baseline specification of the fields, we have 9 completed fields ($j$), 9 possible preferred fields/instruments ($k$), 8 possible next-best fields ($l$).\footnote{In both countries, the number of applicants with Medicine as next-best is very small and these are therefore omitted in our analysis. Thus, there are 9 preferred fields but only 8 next-best fields.} Because preferred field can never be the same as the next-best alternative, we get 576 (and not 648) unique first stage coefficients, $\alpha_{jl}^{k}$. Because $\sum d_{j}=1$ for each applicant, creating a within-applicant correlation between different $d_{j}$, we allow the residuals $u_{jl}$ to be clustered within applicant.

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{Comparing payoff estimates}

We want to compare payoffs to field of study across two different populations:

equation[equation omitted — 139 chars of source]

where we have re-centered the payoffs relative to the average Norwegian payoffs for interpretational convenience: it allows us to interpret the intercept $a_{0}$ as the payoff difference between Denmark and Norway at the average Norwegian payoff. The interpretation of the slope $a_{1}$ -- which quantifies the average increase in the Danish payoffs for a one unit increase in the Norwegian payoffs -- is unaffected by the centering.

There are two considerations that we need to pay attention to when taking equation ((ref)) to the data: measurement error and across-population comparison. Unweighted estimation of ((ref)) would assume that the estimated returns are from populations of similar size. In practice, the return estimates in the two countries will have differently sized groups, where some estimates are based on many applicants shifted by the instrument (when there are many applicants with given preferred and next-best fields and the first stage is large), while others are based on few applicants shifted (when there are less applicants in the preferred/next-best field cell or the first stage is close to zero).

To take these unequal underlying population sizes into account we will weigh our regressions with a measure of the number of applicants that are shifted. For each payoff estimate $\beta_{jl}^{c}$ in country $c$ we calculate the net number of applicants that are shifted on that margin as follows \[ n_{jl}^{k,c}=|\alpha_{jl}^{k,c}|\cdot N_{kl}^{c}\cdot\bar{z}_{kl}^{c} \] where $\alpha_{jl}^{k}$ is the first-stage coefficient, $N_{kl}$ the number of applicants with preferred field $k$ and next-best field $l$, and $\bar{z}$ the share of these applicants above the cutoff. We then construct weights\footnote{It should be noted that in practice using these weights gives very similar results to using population weights $N_{kl}^{NO}+N_{kl}^{DK}$.}\textsuperscript{,}\footnote{When we study distributions of first-stage coefficients we will also use the weights $w_{jl}$.} \[ w_{jl}=\sum_{k}(n_{jl}^{k,NO}+n_{jl}^{k,DK}) \]

Measurement error concerns arise because rather than relating population payoffs as in ((ref)) we will be comparing two sets of noisily estimated population payoffs:

equation[equation omitted — 167 chars of source]

It is well know that measurement error in explanatory variables results in estimation bias. Assuming classical measurement error $\hat{\beta}_{jl}^{c}=\beta_{jl}^{c}+\epsilon_{jl}^{c}$ with $\epsilon_{jl}^{c}$ i.i.d. and $\sigma_{\epsilon,c}^{2}\equiv var(\epsilon_{jl}^{c})$, we can quantify the bias as follows\footnote{Classical measurement error in the dependent variable affects the precision but not the consistency of the regression estimates.} \[ \hat{a}_{1}=\frac{cov(\hat{\beta}_{jl}^{DK},\hat{\beta}_{jl}^{NO})}{var(\hat{\beta}_{jl}^{NO})}\rightarrow a_{1}\frac{var(\beta_{jl}^{NO})}{var(\beta_{jl}^{NO})+var(\epsilon_{jl}^{NO})}=a_{1}R_{NO} \] where the estimate of $a_{1}$ is attenuated by a factor $R_{NO}=1-\sigma_{\epsilon,NO}^{2}/\sigma_{\hat{\beta},NO}^{2}$ (with $\sigma_{\hat{\beta},NO}^{2}\equiv var(\hat{\beta}_{jl}^{NO})$). $R_{NO}$ quantifies the reliability of $\hat{\beta}_{jl}^{NO}$ and, provided we can estimate it, implies that we can adjust $\hat{a}_{1}$ by $1/\hat{R}_{NO}$ to recover an unbiased estimate of the true $a_{1}$.\footnote{We use the Stata command -eivreg- to perform the error-in-variable regression.} We construct an estimate of $R_{NO}$ by plugging in the variance of the payoff estimates as an estimate of $\sigma_{\hat{\beta},c}^{2}$, and using the average squared standard errors of the payoffs as an estimate of $\sigma_{\epsilon,c}^{2}$.\footnote{sullivan2001note shows that this approach is robust to measurement error heteroskedasticity.} Finally, we can use the so-called total reliability $R_{Total}=\sqrt{R_{NO}\cdot R_{DK}}$ to construct an estimate of the correlation of the payoffs across the two countries \[ \hat{\rho}=\rho(\hat{\beta}_{jl}^{DK},\hat{\beta}_{jl}^{NO})/\hat{R}_{Total}\rightarrow\rho\equiv\rho(\beta_{jl}^{DK},\beta_{jl}^{NO}) \]

Table (ref) reports the standard deviation of the estimated payoffs, the square root of their average standard errors squared, as well as the resulting estimated reliability ratios. The first two columns report the unweighted estimates. We see that the payoff estimates vary more in Norway than in Denmark and are on average also more noisily estimated. These unweighted estimates do however not map into a population. The analysis in this paper will therefore investigate weighted results and the next two columns report the weighted reliability estimates. For shifted applicants the variability in the estimates and the average standard error is reduced, especially for the Norwegian estimates. The estimated reliability of the Norwegian payoff estimates is 0.86 compared to 0.72 for the Danish ones. Reliability is therefore high for both countries.\footnote{Using country-specific weights gives slightly higher but very similar estimates, namely a reliability of 0.89 for Norway and 0.78 for Denmark.}

table[table omitted — 1,529 chars of source]

\@startsection {section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus.2ex} {\reset@font}{Payoffs to fields of study}

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{Examining the violation of next-best and irrelevance}

We start by examining whether we can statistically reject the irrelevance and/or next best assumptions. As shown above, these assumptions are rejected if any of the off-diagonal first-stage coefficients are significantly different from zero. Joint tests strongly reject this null-hypothesis for both countries (cf. Tables (ref) and (ref)). This implies that there are irrelevance-violators (detected by positive off-diagonal first-stage coefficients) and/or next-best-violators (detected by negative off-diagonal first-stage coefficients). We tend to detect such violation for most field of studies.

As a first indication of the relative importance of next-best vs. irrelevance violations we consider the signs of the off-diagonal coefficients that are individually significant (Tables (ref) and (ref)). For Norway this reveals that few if any of the positive coefficients are individually significant, especially after adjusting for multiple testing (using the Bonferroni correction). However, a large number of the negative off-diagonal coefficients are significant, also after adjusting for multiple testing. For Norway, we therefore mostly find evidence for violations of next-best. This stands in contrast to the results for the Danish data, which are consistent with violations of the irrelevance and next-best assumptions being approximately equally frequent.

figure[figure omitted — 726 chars of source]

With enough data any model can be rejected, no matter how minor the misspecification. We therefore gauge the empirical relevance of the violations of the irrelevance and next-best alternative conditions by quantifying the relative size of the associated applicant groups. The results are reported in Figure (ref), which shows the distribution of the relevant first-stage coefficients weighted with the number of applicants shifted and where the densities are rescaled so that they to sum to unity. The mass under each density -- reported in parenthesis in the Figure -- quantifies the relative size of the complier/defier group in question.

The left-panel of Figure (ref) shows that at least 64% of the shifted applicants in Norway are shifted at the expected (on-diagonal) margin. Of the remaining shifted applicants nearly 90% are shifted on margins with negative coefficients. For Norway we therefore continue to find evidence for violations of next-best but not irrelevance when we take the size of the shifted applicant groups into account. The results for Denmark in the right-panel of Figure (ref) show that a similar share of applicants is shifted at the diagonal. Off-diagonal the shifted applicants are however evenly distributed between positive and negative margins. This reinforces the earlier conclusion, suggesting that violations of the irrelevance and next-best assumptions are approximately equally frequent. It should be emphasized however that, depending on their sign, the (absolute values of the) off-diagonal first-stage coefficients give a lower bound on each type of violator, while the on-diagonal coefficients provide upper bounds on the compliers. We therefore conclude that in both countries the violations of irrelevance or next-best are quantitatively non-trivial, appear to be of similar magnitude, but of a different nature.

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{Comparison of payoffs}

Figure (ref) reports the reliability-corrected and weighted densities for the Norwegian and Danish estimates of the payoffs of completing a field-of-study instead of the next-best.\footnote{Appendix tables (ref) and (ref) report the payoff estimates.} In each country, the payoffs are measured in terms of annual earnings eight years after application. On average the annual payoff in Denmark is about 2,200 USD, while in Norway the returns are substantially larger at about 22,000 USD. In addition, there is also more variation in the payoffs in Norway compared to Denmark. A joint test of equality of the payoffs across countries gives a $\chi_{64}^{2}$ statistic of 441.7 with a corresponding p-value smaller than 0.0001. We therefore strongly reject that the payoffs are the same. In the following we investigate these differences in more detail.

figure[figure omitted — 185 chars of source]
figure[figure omitted — 212 chars of source]

Figure (ref) starts out with comparing the Norwegian and Danish payoff estimates directly. It plots the estimates in the two countries against each other, with the size of the marker being proportional to the size of the sum of the Norwegian and Danish shifted applicant groups and, in addition to the 45-degree line, the figure also shows the regression line from the following error-in-variables regression ((ref)) described in section (ref) above:

equation[equation omitted — 169 chars of source]

Changes in the intercept $a_{0}$ as we omit estimates with evidence of defiance shows whether the average Danish payoff become more aligned with the average Norwegian payoffs. We also report changes in the slope $a_{1}$ and the estimated correlation between the Danish and Norwegian payoffs $\rho$.

Figure (ref) reports the results of this exercise and shows that, consistent with the low average and lower spread of the Danish estimates in Figure (ref), the Danish estimates increase less than one-to-one with the Norwegian estimates with an estimated slope of 0.38 (s.e. 0.07), and are on average substantially lower (the estimated payoff difference is -19.9 with a s.e. of 2.1). However, even though their levels are different, we find that the payoffs exhibit a relatively strong positive correlation of 0.65 after adjusting for measurement error.

Above we found evidence of violations of the irrelevance and next-best assumptions for both two countries. Can these violations explain the observed differences in the estimated payoffs? We investigate this question by successively removing preferred-next-best combinations with a high share of detected defiers, thus reducing the share of defiers in the sample and see how this impacts the relationships between the Norwegian and Danish payoff estimates.

We first compute, for each completed and next-best field in both countries, the share of applicants that are shifted by off-diagonal instruments. This quantifies the net-flow of irrelevance and next-best defiers at that particular margin. We then progressively drop the estimates with the largest shares of net-defiance and estimate the weighted error-in-variables regression on the resulting sub-sample of Danish payoffs on Norwegian payoffs.

figure[figure omitted — 1,267 chars of source]

Sub-graph (a) shows the estimated intercepts and their confidence intervals. The x-axis in sub-graph (a) shows the maximum share of net-defiance allowed in the sample of estimates for the two countries. Reducing this share one percentage point at a time we re-estimate ((ref)). The first estimate is dropped at about 66 percent net-defiance. Then progressively more payoff estimates are excluded as we restrict the maximum share of defiers below 50 percent. We see that $a_{0}$ stays approximately constant close to -20 until we restrict the share of defiers to be below 50 percent. After this $a_{0}$ gradually increases, and reaches -17 when we restrict the sample to max 20 percent defiers.

Sub-graph (a) also shows the shares of the 64 payoff estimates and of the shifted applicants that are excluded. The pairs of completed/next-best fields that have the highest shares of net-defiers have relatively few applicants shifted on the diagonal. Restricting the maximum to 50 percent we exclude 9 percent of estimates and 2 percent of compliers. Restricting further has a stronger impact on estimates and shifted applicants retained, and when we ultimately restrict the sample to max 20 percent defiers only 4 out of 64 estimates and 21 percent of the shifted applicants are retained.

In sub-graph (b) we plot $a_{0}$ against the share of shifted applicants that are excluded. As a function of applicants excluded, $a_{0}$ rises about linearly. However, as can be seen from the confidence bands in sub-graph (a), the estimated intercepts for different samples are never significantly different.

In sub-graphs (c) and (d) we show similar results for the slope parameter $a_{1}$ from ((ref)). While $a_{1}$ increases somewhat in the beginning, it is mostly stable across the different samples. Finally, in sub-graphs (e) and (f) we show the reliability-adjusted weighted coefficient of correlation. This increases steadily with the share of compliers excluded, from 0.65 in the full sample to 1 when restricting to less than 20 percent net-defiers.

While we found above that the Norwegian and Danish payoff estimates are strongly correlated, this correlation substantially increases further when we exclude the estimates with more evidence of defiance of irrelevance and next-best. The intercept and slope from the regression ((ref)) are however relatively stable, suggesting that violations of irrelevance and next-best do not explain the lower level and variation of the payoffs in Denmark compared to Norway.

\@startsection{subsection}{2}{\z@} {-3.25ex\@plus -1ex \@minus -.2ex} {1.5ex \@plus .2ex} {\reset@font}{Other explanations for differences in payoffs across the countries}

table[table omitted — 4,665 chars of source]

To explore other explanations for the between-country differences in payoffs, we re-estimate ((ref)) while adjusting for completed and next-best field dummies, as well as differences in average selectivity and earnings (cf. Figure (ref)) across completed and next-best fields. Table (ref) reports the results. The first column reproduces the basic results reported above in Figure (ref) where we found that the payoff difference was about 20,000 USD, and that the payoffs in Denmark increased by less than one for each unit increase in Norway reflecting the smaller variance in the payoff distribution in Denmark.

We next investigate whether payoffs are more aligned across completed fields or across next-best fields. The next two columns of Table (ref) therefore adjust for completed field and next-best field dummies. Keeping completed field fixed we now obtain a slope estimate of 0.29 in column (2), while keeping next-best field fixed in column (3) increases the slope substantially to 0.70. This shows that differences between next-best fields contribute more to between-country differences in payoffs than differences between completed fields.

In a next step we investigate two potential explanation of such differences. First we verify whether differential selectivity plays a role by adjusting for across country differences in average GPA in both the completed field $j$ and the next-best field $l$. This changes the interpretation of the intercept which now corresponds to the across country payoff difference keeping the average GPA the same in the completed and next-best field. The estimates in column (4) show that this reduces the payoff gap with 25% from about -20,000 to -16,000 USD, while at the same time payoff become more evenly distributed as shown by the increase in the slope coefficient from 0.37 to 0.56. A similar exercise using average earnings in column (5) and (6) shows that this does not explain across country differences.

As noted earlier, looking at earnings eight years after application corresponds to relatively early career outcomes, especially for 5-year programs and studies that are not closely tied to a narrow set of occupations and which may therefore have longer and more complex school-to-work transitions. We therefore also compare the Danish payoff estimates 13 years after applying with the Norwegian estimates 13 years after applying. Extending the time horizon by 5 years has some impact on the estimated reliability of the Norwegian estimates which drops to 0.64, but the raw correlation between the $t=8$ and $t=13$ estimates is high (0.80).\footnote{We need to exclude three very imprecisely estimated payoffs with Law as next-best field from the $t=13$ analysis (see appendix Table (ref)) to recover a non-negative reliability estimate. } For Denmark the reliability is slightly higher at $t=13$ as is the raw correlation between the $t=8$ and $t=13$ estimates (0.86).

Figure (ref) shows the estimated distributions of these longer-run payoffs across the two countries. Compared to the early career payoffs, the Norwegian and Danish payoff distribution are now much more aligned in terms of location and scale. This can also be seen in column (7) of Table (ref). The average payoff gap between the two countries is now about 10,000 USD, and the slope coefficient has increased to 0.70.\footnote{Appendix Tables (ref) and (ref) report the estimates, and appendix Figure (ref) compares the payoff estimates and reports the error-in-variables regression line.} Column (8) shows that after adjusting for differential selectivity the payoff estimates are on average aligned and we cannot reject that the intercept equals zero and the slope equals one (the corresponding F-test gives a p-value of 0.67). However, the estimated (reliability corrected) correlation coefficient between the Norwegian and Danish payoff estimates barely moves when comparing $t=8$ vs. $t=13$ (0.66 vs 0.65).

figure[figure omitted — 225 chars of source]

To summarize, we find that payoff estimates are strongly correlated across countries but have initially different levels and dispersion. Violations of the irrelevance and next-best assumptions that underpin the empirical approach do weaken the correlation, but appear to have little consequence for the estimated level and variance differences. Over time, the level and variance difference converge across countries, but this does not affect the correlation of the payoffs. Additional exploratory analyses show that these across country differences are mostly driven by heterogeneity in next-best fields which can partly be explained by differences in selectivity.

\@startsection {section}{1}{\z@} {-3.5ex \@plus -1ex \@minus -.2ex} {2.3ex \@plus.2ex} {\reset@font}{Conclusion}

onehalfspaceWe revisited the identification argument of kirkeboen_field_2016 who showed how one may combine instruments for multiple unordered treatments with information about individuals\textquoteright ranking of these treatments to achieve identification while allowing for both observed and unobserved heterogeneity in treatment effects. We showed that the key assumptions underlying their identification argument have testable implications. We also provided a new characterization of the bias that may arise if these assumptions are violated. Taken together, these results allow researchers not only to test the underlying assumptions, but also to argue whether the bias from violation of these assumptions are likely to be economically meaningful.
onehalfspaceGuided and motivated by these results, we estimated and compared the earnings payoffs to post-secondary fields of study in Norway and Denmark. In each country, we applied and assessed the identification argument of kirkeboen_field_2016 to data on individuals' ranking of fields of study and field-specific instruments from discontinuities in the admission systems. We empirically examined whether and why the payoffs to fields of study differ across the two countries. We found strong cross-country correlation in the payoffs to fields of study, especially after removing fields with violations of the assumptions underlying the identification argument.

While our empirical findings are specific to the context of postsecondary education in the Nordic countries, there could be lessons from our work for other settings with unordered choices. Our study highlights key challenges and possible solutions to understanding what the causal effects of these choices are. Examples can be found in observational studies that use IV to study workers\textquoteright selection of occupation, students' choice of education, firms\textquoteright decision on location, or families\textquoteright choice of where to live. Another example is the frequent use of IV to analyze encouragement designs in experiments where treatments are made available but take up is not universal duflo2007using.