Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Do Test Scores Help Teachers Give Better Track Advice to Students? A Principal Stratification Analysis
abstractEvery year, over one million EU students choose a secondary school track based on teacher recommendations, yet little evidence shows this yields optimal assignments. Using Dutch data, we examine whether access to standardized test scores improves recommendation quality. We develop a Principal-Stratification metric in a quasi-randomized setting, conduct a welfare analysis that flexibly weights short- and long-term losses, and assess principal fairness by examining whether test-score access affects equity across protected attributes. Results are robust to replacing the Exclusion Restriction assumption underlying our main identification strategy with alternative assumptions. Allowing recommendation upgrades when test scores exceed expectations increases successful placement in more demanding tracks by at least 6%, while misplacing 7% of weaker students. Only unrealistically high weights on short-term losses would justify banning such upgrades. Test-score access also yields fairer recommendations for immigrant and low-SES students. Our methodology and findings contribute to the literature on algorithm-assisted human decisions.
JEL-Code: I2
Keywords: principal stratification; secondary school track recommendations.
\thispagestyle{empty}
spacing{1.5}
\setcounter{page}{1}
\section{Introduction}
Every year, more than one million EU students, upon completing primary education, choose a secondary school track based on recommendations received from their teachers. In some institutional settings—such as the Dutch school system examined in this paper—only under exceptional circumstances students may begin secondary education in a more demanding track than the one recommended to them. Despite the potentially dramatic consequences that bad track recommendations may have, there is no clear evidence that assigning to teachers alone the role of providing these recommendations is the best solution and generates an optimal track assignment.\footnote{There is a wide literature on teachers' biases that may affect the quality of secondary school track recommendations. Recent examples are:
Driessen2008,
Burgess2013,
Gershenson2016,
Osikominu2021,
vanLeest2021,
Geven2021,
Carlana2022b,
Carlana2022c,
Carlana2022a,
Ferman2022,
Batruch2023
Bach2023,
Alesina2024,
Huizen2021, and
Falk2020.
}
Our goal in this paper is to establish, using a well-defined metric and a quasi-randomized experiment, whether teachers can “improve" the quality of their recommendations after seeing the scores of standardized tests taken by their students.
To the best of our knowledge, there is no consensus on what the goal of secondary school track advice should be or what defines a “better" recommendation.
The metric we adopt is the one based on Principal Stratification frangakis2002principal\footnote{For a discussion on Principal Stratification and how it is rooted in the Instrumental Variable literature, see MealliMattei2012Refreshing.} that Imai2023 use to study whether algorithms can help judges in deciding which arrested individuals should be released while waiting for their trial: in this case, a decision is better than another if it avoids releasing subjects at risk of repeating a crime.
In the Dutch setting, a reform was introduced that gives teachers the right (but not the obligation) to revise their initial track assignment upwards if the student performs better on a standardized test than the initial recommendation would imply. This reform offers an ideal setting to apply the metric that we adopt\footnote{The setting offered by the reform was also used in Oosterveen2022. }.
The nature of this metric in the context of track advising can be intuitively described with reference to an education system featuring two tracks: low (e.g., vocational) and high (e.g., academically oriented). We make
two assumptions that we maintain throughout the analysis: (1) if a student is able to complete a more difficult track, it is better she is allowed to do so for herself (because she can obtain more desirable lifetime outcomes) and for society (inasmuch as there are collective gains from a more educated community); (2) changing between tracks is costly for students and for the education system.
Note that given these two assumptions, even in the absence of negative congestion externalities in the educational process and in the labor market\footnote{For evidence and a discussion of these externalities see ichino2025.} it would be sub-optimal to send all students to the high track because it would be costly for some of them not to complete it and have to change to the low track.
Let's divide the population of students into groups -- principal strata frangakis2002principal -- that differ by how the education outcome depends on track recommendations.
The first group in which we are interested comprises students for whom the track recommendation determines unequivocally the completed track: if the student is recommended the low track in the first year, she will graduate in the same track at the end of secondary school, while if recommended the high track, this will be her final graduation outcome. These are students for whom the teacher's recommendation is a self-fulfilling prophecy, thus becoming crucially important. In light of this consideration, we call these subjects “Helpable” (H) to signal that they are the ones who can be {\it helped} by a high first-year track recommendation instead of a low one.\footnote{They correspond to the group of cases that, in the context of Imai2023, are labeled as “Preventable” crimes. Similarly, Oosterveen2022 label them as “Trapped-in-Track” (TT).} Adapting the methodology proposed by Imai2023 to this context, we can establish whether the information provided by test scores helps teachers improve their recommendations by measuring the fraction of students in the H group that are recommended the high track by teachers and see if it increases when teachers can decide to upgrade based on test score information. In the Principal Stratification framework, this is the Average Principal Causal Effect ($APCE$) of “changing the source" of a recommendation (i.e., giving test score information to the source, in our context) within the principal stratum of H students. If this $APCE$ is positive, test scores help teachers provide better track advice by upgrading to more challenging tracks students who can effectively complete them.
We are also interested in two other groups of students: those who, independently of the track initially recommended to them, always finish secondary education in the low track, and those who, instead, always graduate from the high track independently of the initial recommendation. Adapting to our setting the notation of Oosterveen2022, we call these students Always Low (AL) and Always High (AH), respectively. The common characteristic of these two groups is that the final student's outcome is not affected by the teacher's recommendation; however, this outcome can be achieved in different ways, which are more or less costly, depending on the number of track changes they require. In these cases, the goal of a recommendation should be to minimize the number of track changes that these two types of students will experience if they start on a track that differs from the one they will ultimately complete. The best recommendation is then the one that sends all AL students to the low track and all AH students to the high track. Therefore, allowing teachers to upgrade based on test scores produces more desirable outcomes, the larger the $APCE$ is for AH students and the lower it is for their AL peers.\footnote{It is conceivable the existence of a fourth group of students who always finish secondary education in the {\it opposite} track with respect to the one that was initially recommended to them. This is discussed below. }
We estimate an $APCE$ for H students which is not lower than 6%, suggesting that when teachers can upgrade on the basis of test scores, the quality of their advice improves by at least 6% as measured by the fraction of these students that are recommended a more difficult track and can complete it successfully, while in the counterfactual case, they would have remained in a lower level track. For AH students, for whom an upgrade is also desirable to reduce the cost of track changes, the $APCE$ is estimated to have approximately the same lower bound.
In the case of AL students, for whom an upgrade is not desirable, the corresponding estimates are also positive with an $APCE$ of at least 7%. This suggests that a significant fraction of students who are unable to complete the more challenging curriculum are upgraded, and we cannot rule out that this fraction is even greater than that estimated for the H and AH strata. A possible reason for this finding is that primary school teachers do not bear any consequence related to track changes in secondary schools and are, therefore, less sensitive to the goal of reducing them in the case of AL students.
Given the opposite desirability of the findings concerning the effect of test scores on the quality of teacher advice for H and AH students on one hand and AL students on the other, it is natural to seek a method to obtain a comprehensive evaluation. An approach is offered by the classification framework proposed by benmichael2024does. This framework enables us to construct a weighted sum of the losses generated by mistakes in recommendations for the various types of students. We can then test whether this sum decreases when teachers can upgrade based on test scores, depending on the relative weight of the short-term losses suffered by students who must change track because of misplacement versus the lifetime losses of students who are not directed to a more challenging track they can eventually complete. We find that the relative weight of short-term losses would have to be unreasonably high to conclude that teachers should not be allowed to upgrade based on test scores.
Finally, building on Imai2022, a corollary contribution of this analysis is that it makes it possible to evaluate whether being allowed to upgrade based on test scores helps teachers improve the {\it fairness} of their recommendations with respect to protected attributes such as SES, race, or gender. For example, suppose that gender is a protected attribute with respect to which we want to assess how fair a recommender is. Then, in the stratum of H students, the recommender is fair if the fraction of high recommendations is the same for female and male students. However, the analogous fraction in other strata could be lower or higher as long as it is equal for both genders. Therefore, in the overall population, females may receive a different fraction of high first-year advice from this recommender; however, this would not constitute a violation of fairness, as it would be due to differences in the gender composition of the strata, not to discrimination based on gender within a stratum.
We find that allowing teachers to upgrade recommendations based on test scores improves fairness. Specifically, it raises the chances that immigrants and low-SES students, who are otherwise discriminated against, are recommended a more difficult track that they can complete relative to natives and high-SES peers, respectively.
The paper is organized as follows. In Section (ref), we describe the Dutch institutional setting and the quasi-randomized experiment that makes our analysis possible.
The statistical framework and its assumptions follow in Section (ref).
Section (ref) presents our results, while Section (ref) supports their robustness by showing that they also hold when the Exclusion Restriction assumption is substituted by alternative assumptions. In Section (ref) we move to the analysis of the fairness of track recommendations, and Section (ref) concludes.
\section{The Dutch school system and the quasi-experiment}
What needs to be known about the Dutch school system for the purpose of this paper is that it features five main secondary school tracks, of which three are vocational (VMBO-BL, VMBO-KL, VMBO-GT ranked in this order by level of difficulty) and two lead to university studies (HAVO, VWO, similarly ranked). In addition, tracks that are contiguous by level of difficulty can be combined in mixed tracks, so that there are in total nine tracks to which a student can be recommended in the first year of secondary school. These nine tracks are listed from the easiest to the hardest in the row headings of Table (ref).
\begin{table}[h]
\caption{Tracks of the Dutch school system}
\begin{center}
\begin{tabular}{lccccc}
\hline \hline
\\[-2ex]
\hspace {2cm} & \hspace{2cm}& \hspace{2cm} & \hspace{2cm}
\\ [-2ex]
Track & \multicolumn{2}{c}{\makecell[c]{Students initially \\assigned to the track \\denoted by the row}} & \multicolumn{1}{c}{\makecell[c]{Test score cut-\\off for upgrade \\ to next level}} & \multicolumn{2}{c}{\makecell[c]{Fraction above the cutoff\\ of students eligible\\for upgrade to next level}} \\
& 1& 2& 3 & 4 & 5 \\
\\ [-2ex]
\hline
\\ [-2ex]
V: BL & 9,585 & 6,929 & 519 & \hspace{0.5cm} 0.45 & 0.46\\
V: BL/KL & 3,781 & 3,828 & 526 & \hspace{0.5cm} 0.31 & 0.30\\
V: KL & 15,698 & 11,258 & 529 & \hspace{0.5cm} 0.33 & 0.36\\
V: KL/GT & 3,255 & 4,003 & 529 & \hspace{0.5cm} 0.49 & 0.54\\
V: GT & 29,117 & 21,154 & 533 & \hspace{0.5cm} 0.45 & 0.46\\
V: GT/HAVO & 8,540 & 9,620 & 537 & \hspace{0.5cm} 0.40 & 0.41\\
A: HAVO & 28,226 & 21,465 & 540 & \hspace{0.5cm} 0.47 & 0.50\\
A: HAVO/VWO & 9,932 & 11,136 & 545 & \hspace{0.5cm} 0.29 & 0.31\\
A: VWO & 27,750 & 24,469 & n.a & \hspace{0.5cm} n.a. & n.a.\\
\\ [-2ex]
\hline
\\ [-2ex]
Total & 135,884 & 113,862 & & \hspace{0.5cm} 0.42& 0.43\\
\\ [-2ex]
\hline
\\ [-2ex]
Cohort & 2015 & 2016 & 2015 \& 2016 & \hspace{0.5cm} 2015 & 2016 \\
\\ [-2ex]
\hline \hline
\end{tabular}
\end{center}
\vspace{-0.2cm}
\begin{minipage}{1\linewidth \setstretch{0.75}}
{\scriptsize Notes:
The row headings are the names of the secondary school tracks to which a student can be assigned in the Dutch system, ranked by level of difficulty. The first six tracks are vocational, while the last three lead to university studies (academic). Columns 1 and 2 report the number of students initially assigned to each track by their primary education teachers, in view of starting secondary school in the academic year 2015-16 or 2016-17. Column 3 reports the CITO test score cutoffs above which a student in the track denoted by the corresponding row may be upgraded to the next level by the primary school teacher; the test score takes on integer values $\in\{\underline{S}, \underline{S}+1, \underline{S}+2, \dots, \overline{S}\}$, with $\underline{S}=501$ and $\overline{S}=550$. The upgrade is not mandatory. Columns 4 and 5 report, for each initial track, the fraction of students who, in the years considered, may be upgraded due to a sufficiently high CITO test score.}
\end{minipage}
\end{table}
Using the database of the Dutch National Cohort Study on Education (NCO), we consider the population of students who begin their secondary school in the academic years
2015--16 and 2016--17 (hereafter, 2015 and 2016 cohorts).\footnote{This database can be accessed following the instructions provided at this \href{https://www.cbs.nl/nl-nl/onze-diensten/maatwerk-en-microdata/microdata-zelf-onderzoek-doen/microdatabestanden/nco-nationaal-cohortonderzoek-onderwijs}{link}.
For documentation in English, \href{https://www.nationaalcohortonderzoek.nl/sites/nco/files/media-files/20211022_NRO-NCOcodeboek_ENG_def.pdf}{ see here}). The track assignment regulations that generate the quasi-experiment we exploit became effective for students starting secondary school in the 2014-15 academic year, but the data for this cohort do not contain all the necessary information and thus cannot be used.
}
The timing implied by these regulations is described in Figure \ref{f:timeline}. Let $t=1$ denote the first year of secondary school.
In March of the previous year, $t=0$, the primary education teachers give their students an initial track recommendation.
The first and second columns of Table \ref{t:tracks} report the number of students in each cohort initially assigned to each track by their primary education teachers as a result of this provisional recommendation.
\begin{figure}[tbp]
\caption{Time-line for track recommendations and outcomes}
\label{f:timeline}
\vspace{-0.5 cm}
\begin{center}
\includegraphics[width = 14.5cm]{time_line_v10.png}
\end{center}
\vspace{-0.6cm}
\end{figure}
In April of the same year, $t=0$, the students rank the schools offering the track that was recommended to them (or a less difficult one if they prefer) and are assigned to one of them.\footnote{The more specific method of student allocation differs by municipality. In Amsterdam, for example, a deferred acceptance algorithm is used, with lottery numbers serving as tiebreakers (\citealp{Ketel2023}).} Then, in the following month, they take their end-of-primary education standardized test (CITO), which is graded by a computer with no intervention from primary education teachers. The scores obtained in this test map into track levels on the basis of nationwide, predetermined test score cutoffs. As in \cite{Oosterveen2022}, we refer to the mapping of a CITO test score to a track level as a ``test-based'' track assignment. If the test-based track assignment is higher than the teacher’s track assignment recorded in March of year $t=0$, the teacher is mandated by the law to consider an upward revision of the provisional recommendation. Such revision is only optional and dependent on the teacher’s judgment.
Column 3 of Table \ref{t:tracks} reports the CITO test score cutoffs above which a student, whose provisional recommendation is the track of the corresponding row, must be re-evaluated by the teacher for a possible upgrade. Columns 4 and 5 report, for each track, the fraction of students who
may be upgraded because of a sufficiently high CITO test score.
Let $R_{i}$ be a dummy variable equal to $0$ if the primary education teacher \textit{finally} recommends the low track to student $i$.
$R_{i}$ is, instead, equal to $1$ if the teacher recommends an upgrade to a higher track.
Note that the 52,219 students whose provisional recommendation is the highest track (VWO in the last row of Table \ref{t:tracks}) cannot be treated with $R_{i}=1$ because there is no higher track to which they can be upgraded. Our analysis focuses on the students for whom the track provisionally recommended in March of year $t=0$ by the teacher is one of those listed in the first 8 rows of Table \ref{t:tracks}, totaling 197,527 students.
Denote with $K^\tau_i$ the track that student $i$ in cohort $\tau \in \{2015, 2016\}$
is provisionally recommended by the teacher; $K^\tau_i$ can take on values $\in \{1,8\}$. Let $S^\tau_i$ be the score obtained by student $i$ in cohort $\tau$ in the CITO test; $S^\tau_i$ can take on integer values $\in\{\underline{S}, \underline{S}+1, \underline{S}+2, \dots, \overline{S}\}$, with $\underline{S}=501$ and $\overline{S}=550$. Let the assignment to treatment be $Z_i$, defined as follows:
\begin{equation}
Z_i = \begin{cases}
0 & \mbox{if } \hspace{0.3 cm}
S^\tau_i = C^\tau_{K^\tau_i} -1 \hspace{0.5 cm}
\\
1 & \mbox{if } \hspace{0.3 cm}
S^\tau_i = C^\tau_{K^\tau_i} \hspace{0.5 cm}
\\
\mbox{else} & \mbox{if } \hspace{0.3 cm}
S^\tau_i \neq \{C^\tau_{K^\tau_i}-1, C^\tau_{K^\tau_i}\} \hspace{0.5 cm}
\end{cases}
\end{equation}
where $C^\tau_{K^\tau_i}$ is the cutoff score that student $i$ in cohort $\tau$ provisionally assigned to track $K^\tau_i$ has to reach in order to be re-evaluated by her primary school teacher for a possible upgrade; ``else''
denotes students who do not score close to their respective cutoff score.
If we condition on the students in a particular cohort $\tau$, with a specific value of $K^\tau_i$ and with a value of $Z_i \in \{0,1\}$, we can assume that only randomness, possibly conditional on some covariates to capture the heterogeneity of primary schools as illustrated below, determines whether $S^\tau_i= C^\tau_{K^\tau_i}-1$ or $S^\tau_i=C^\tau_{K^\tau_i}$, that is, whether
$Z_i=0$ or $Z_i=1$.
This institutional setting generates the quasi-experiment we exploit to answer our research questions. Among students with a realized value of
$S^\tau_i \in \{C^\tau_{K^\tau_i}-1, C^\tau_{K^\tau_i}\}$, $Z_i$ randomly assigns similar subjects in each of the eight initial tracks to two different ``sources'' of first-year track recommendation: one is the ``teacher alone'', while the other is the ``teacher informed by the CITO test results'', who therefore has the possibility to upgrade students based on such results. This assumption is known in the Regression Discontinuity (RD) literature as ``local randomization'' (LR).\footnote{Local randomization has been first formalized by \cite{LiMatteiMealli2015} and \cite{CattaneoFrandsenTitiunik2015}. The LR approach to the analysis of RD designs has been later discussed in \cite{MatteiMealli2017ObsStudies, BransonMealli2019, CattaneoEtAl_book2020a,CattaneoEtAl_book2020b}, and extended to allow for randomization conditional on covariates in \cite{Forastiere2025}.} One of its advantages is that it allows the researcher to easily deal with discrete running variables without having to rely on the usual continuity-type assumption\footnote{See also \cite{EcklesEtAl2020} and \cite{LiMercatanti2020}}. We formalize the LR assumption for our context in Section \ref{s:stat}, where we also provide evidence to assess its plausibility.
Next, we denote with $Y_{i} \in \{0,1\}$ the outcome that we study. We define $Y_{i} = 1$ if student $i$ graduates from a higher track than the one provisionally recommended to him/her in March of $t=0$, before taking the CITO test, and $0$ otherwise. We consider a student as graduating in a given track even if he/she does so in more years than the required minimum.\footnote{Our conclusions do not change qualitatively if we assume instead that a student graduates in a given track only if he/she does so in the regular possible time (results available from the authors). Four years are typically required for a vocational track, while five and six years are needed for the two academic tracks (HAVO and VWO, respectively). Six years are also needed for students who change from VMBO to HAVO, and seven for those who change from HAVO to VWO.
The release of the NCO data at our disposal follows the 2015 cohort for 7 years, and the 2016 one for only 6 years after enrollment. Graduation is not observed for some students in these cohorts. We thus assume that the students of these cohorts who are still enrolled in secondary school in the last year of observation will graduate in the track in which they are enrolled.} Our goal is to evaluate which source of track recommendation (``teacher alone'' or ``teacher informed by test scores and allowed to upgrade")
is ``preferable'' according to two criteria: a recommendation is better if it (i) induces a student to enroll in the most difficult track that she/he is able to complete successfully and (ii) reduces to a minimum the chances of costly track changes.
An institutional complication of this quasi-experimental setting is that, on rare occasions, also students who score below the cutoff ($Z_i=0$) may be recommended for an upgrade by their teacher ($R_{i} = 1$).\footnote{As explained by the relevant authorities (\citealp{VanPoNaarVo2025}),
``[T]here may be situations in which a student does not obtain a higher test recommendation, but a school or teacher still believes that there is more knowledge and skills than expected. In such situations, the school may choose to adjust the provisional school advice. The school's judgment is central here. ... If parents do not agree with the advice, they will consult with the teacher and/or director of the primary school. If they cannot reach an agreement about the level at which the student will enter secondary education, parents can, as a last resort, use the school's complaints procedure.'' \label{foot_def}}
Moreover, as already mentioned, students with $Z_i=1$ may have $R_{i} = 0$ because their primary school teacher does not want to upgrade them even if an upgrade is possible. Therefore, as we will show in Section \ref{s:stat}, non-compliance may occur in both assignments to treatment conditions.
\section{Statistical framework}
\label{s:stat}
Consider a sample of $N^\tau_{k}$ units of cohort $\tau$ whose provisional recommendation is track $k$ with $Z_i\in \{0,1\}$, $i \in 1, \cdots, N^\tau_{k}$.
Let $\mathbf{Z}$ and $\mathbf{R}$ be the $N^\tau_{k}$-dimensional vectors of the possible values of $Z_i$ and $R_{i}$, with generic element $z_i$ and $r_{i}$ respectively. Denote with $R_{i}(\mathbf{Z}) \in \{0,1\}$ the recommendation that student $i$ receives from the teacher as a function of the assignment vector $\mathbf{Z}$. Similarly, denote with $Y_{i}(\mathbf{Z}, \mathbf{R}) \in \{0,1\}$ the potential track (low or high) completed by student $i$ at the end of secondary school as a function of the assignment vector $\mathbf{Z}$ and the vector of recommendations $\mathbf{R}$.
Since different primary schools may have different benchmarks to which they compare their students and decide their provisional track assignments, we must consider the possibility that students with different abilities are assigned to the same provisional track $k$ simply because they have more or less lenient primary teachers. We denote with $\pi_i$ the benchmark against which student $i$ has been evaluated to decide her/his provisional track assignment.
\subsection{Assumptions}
\label{s:assumptions}
The first assumption we need is the Stable Unit Treatment Value Assumption (SUTVA, see \citealp{Rubin1980}), which restricts the dependence of the potential outcomes for a student only to the assignment and enrolment of that student, ruling out interference:
\begin{assumption} \label{a:SUTVA}
Stable Unit Treatment Value Assumption: \\
$ R_{i}(\mathbf{Z})=R_{i}(\mathbf{Z'})
\hspace{0.5cm} \mbox{if } z_i=z'_i$ \\
\noindent
$ Y_{i}(\mathbf{Z}, \mathbf{R})=Y_{i}(\mathbf{Z'}, \mathbf{R'}) \hspace{0.5cm} \mbox{if } z_i=z'_i \hspace{0.5cm}\mbox{and
}
\hspace{0.5cm}
r_{i}=r'_{i}$.
\end{assumption}
Assumption \ref{a:SUTVA} postulates the existence of two potential versions of $R_{i}(z)$ with $z\in \{0, 1\}$ and four potential versions of $Y_{i}(z, r)$ with $z, r\in \{0, 1\}$. Given SUTVA, the next assumption is:
\begin{assumption} \label{a:rand_sourc}
Local Randomization of the source conditional on primary schools' leniency benchmarks and provisional track assignments:
for every cohort $\tau$
we have
$$ \{R_{i}(z), Y_{i}(z,r), X_i, U_i \} \perp\!\!\!\!\perp Z_i \hspace{0.1cm}
\vert \hspace{0.1cm} \pi_i , K_i $$
for $z \in \{0, 1\}$ and $r \in \{0, 1\}, $ where $X_i$ and $U_i$ are observable and unobservable covariates, respectively.
\end{assumption}
This assumption requires some discussion.
It formally says that for students assigned to the same track and whose teachers have a similar benchmark criterion $\pi_i$ to evaluate beliefs about students' ability, the reasons for them to score just below or at the cutoff are unrelated to either potential outcomes or observed, $X$, or unobserved, $U$, covariates.
Hence, in their case, $P(Z_i=1|R_{i}(z), Y_{i}(z,r), U_i, X_i, \pi_{i}, K_i)= P(Z_i=1|\pi_{i}, K_i)$.
If $\pi_i$ is observed, this is a testable implication in that if Assumption \ref{a:rand_sourc} holds, we can expect, within each provisional track assignment and for a given $\pi_i$, the observed characteristics of students to be well-balanced across the score cutoff \citep{MatteiMealli2017ObsStudies,Forastiere2025, {AngristRokkanen2015}}.
Moreover, if Assumption \ref{a:rand_sourc} holds for all
tracks, we can view the data as coming from a block-randomized experiment, with blocks defined by the tracks,
and analyze them accordingly \citep{ImbensRubin2015}.
The problem is that while we observe the initial track assignment and we can condition on scoring just below or at the cutoff, we do not observe the benchmark $\pi_i$ of a student's primary teachers. A solution would be to condition on primary school fixed effects, assuming that teachers in the same school collectively agree on the same benchmark $\pi$. However, we do not have enough students per primary school to conduct such an analysis. What we can do is to proxy $\pi_i$ with the fraction of students in the primary school of $i$ who score above the cutoffs. This fraction measures the average teacher severity in that school because, as more students score above the cutoffs, the stricter their primary school teachers must have been in their provisional track recommendations. We define this variable as
$\tilde{Z}_i = \frac{\sum_{j\in School_{i}}\mathbb{I}(S^\tau_j \geq C^\tau_{K^\tau_j})}{\# School_{i}}$, which takes on the same value for all students of a given school
in a given cohort. $\tilde{Z}_i$ will thus be included in all the specifications of our analysis.
\begin{landscape}
\begin{table}
\caption{Balancing of covariates}
\label{t:balance}
\vspace{-0.1cm}
\begin{center}
\setlength{\tabcolsep}{2pt}
\begin{tabular}{lccccccccccc}
\hline \hline
\\[-2ex]
& \multicolumn{11}{c}{Panel A: 2015 cohort} \\
\\[-2ex]
\hline
\\[-2ex]
Track & Female & Immigrant & College & Missing & Household & Urbanity & Primary & Primary & Student's & Obs. for & Obs. for \\
& & origin & mother & college & income & primary & school & school \# & age in & \(Z=0\) & \(Z=1\) \\
& & & & mother & & school & confession & students & months & \\
& 1 & 2 & 3 & 4 & 5 & 6 & 7 & 8 & 9 & 10 &11\\
\\[-2ex]
\hline
\\[-2ex]
V: BL & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 519 & 511 \\
V: BL/KL & 1 & 0.7517 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 173 & 183 \\
V: KL & 1 & 0.1612 & 1 & 0.2872 & 1 & 0.1856 & 1 & 1 & 1 & 1199 & 860 \\
V: KL/GT & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 0.4632 & 274 & 219 \\
V: GT & 1 & 1 & 1 & 1 & 1 & 1 & 0.6102 & 1 & 1 & 1942 & 1948 \\
V: GT/HAVO & 1 & 0.8050 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 614 & 570 \\
A: HAVO & 1 & 1 & 1 & 1 & 1 & 0.4935 & 1 & 1 & 1 & 2149 & 2172 \\
A: HAVO/VWO & 0.0787 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 843 & 769 \\
\\[-2ex]
\hline
\\[-2ex]
& \multicolumn{11}{c}{Panel B: 2016 cohort} \\
\\[-2ex]
\hline
\\[-2ex]
Track & Female & Immigrant & College & Missing & Household & Urbanity & Primary & Primary & Student's & Obs. for & Obs. for \\
& & origin & mother & college & income & primary & school & school \# & age in & \(Z=0\) & \(Z=1\) \\
& & & & mother & & school & confession & students & months & \\
& 1 & 2 & 3 & 4 & 5 & 6 & 7 & 8 & 9 & 10 &11\\
\\[-2ex]
\hline
\\[-2ex]
V: BL & 1 & 0.7578 & 0.7132 & 1 & 1 & 1 & 1 & 1 & 1 & 396 & 458 \\
V: BL/KL & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 195 & 195 \\
V: KL & 0.4012 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 0.5522 & 656 & 630 \\
V: KL/GT & 1 & 1 & 0.8141 & 1 & 1 & 1 & 1 & 1 & 1 & 261 & 284 \\
V: GT & 1 & 1 & 0.0193 & 0.9419 & 1 & 1 & 1 & 0.8404 & 1 & 1386 & 1376 \\
V: GT/HAVO & 1 & 0.0915 & 1 & 1 & 1 & 1 & 1 & 1 & 0.3441 & 923 & 689 \\
A: HAVO & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1643 & 1728 \\
A: HAVO/VWO & 0.0976 & 0.3389 & 1 & 1 & 1 & 0.4697 & 1 & 1 & 0.1691 & 983 & 952 \\
\\[-2ex]
\hline \hline
\end{tabular}
\end{center}
\vspace{-0.1cm}
\begin{minipage}{1\linewidth \setstretch{0.75}}
{\scriptsize Notes: The row headings are the names of the secondary school tracks to which a student can be assigned in the Dutch system, ranked by level of difficulty. Columns 1 to 9 report, for each track and covariate, the Holm-Bonferroni p-value of a test for the equality of means in the two assignment-to-treatment groups defined by $Z_i\in\{0,1\}$.
The definition of the covariates whose names are not self-explanatory is as follows.
Urbanity of a primary school is a categorical variable indicating the population density at the location of the primary school. Primary school confession is a categorical variable indicating the type of confession in primary schools (Protestant, Catholic, Public, or Other). Age in months refers to the age of the student at the time she/he receive the first recommendation.
The last two columns report the number of observations in the two assignment groups. Standard errors are always clustered at the primary school level.
}
\end{minipage}
\end{table}
\end{landscape}
Columns 1 to 9 of Table \ref{t:balance} support the validity of Assumption \ref{a:rand_sourc} by showing that nine observed
characteristics of students are balanced in the immediate vicinity of the cutoff. These covariates are: gender, migrant origin, mother's education, missing mother's education, household income, urban location of the primary school, confession of the primary school, number of students in the primary school, and age in months. Under Assumption \ref{a:rand_sourc}, our estimates should not be (and in fact are not, as we will show below) affected by the inclusion or exclusion of these balanced covariates because they \textit{are not} correlated with $Z_i$, conditional on $\tilde{Z}_i$. However, including them may increase efficiency since they \textit{are} likely to be correlated with $R_i$ and $Y_i$, and for this reason, they will appear in our preferred specification.
It must be noted that the validity of Assumption \ref{a:rand_sourc} does not imply that the proportion of students scoring above or below the cutoff is around $50\%$, as shown in the last two columns of Table \ref{t:balance}. The CITO test score is the result of the students responding correctly to a series of questions. To simplify this data-generating process, think about the score being generated (approximately) by a binomial distribution and suppose that units in the sub-sample that we consider have the same probability $\chi$ to respond correctly to a question in the test. If so, these units also have the same probability (propensity score) to score at the cutoff ($S^t_i=C^t_{K^\tau_i}$), yet this probability is not equal to the probability of scoring below ($S^\tau_i=C^\tau_{k^\tau_i}-1$) because these probabilities follow a binomial distribution. The difference in these two probabilities will depend on $\chi$ and on the number of questions in the test, but it is by no means an indication of the failure of Assumption \ref{a:rand_sourc}. In the data, we may also observe differences in this probability
from track to track because the cutoff is different or because the difficulty of the test (i.e., the probability of a correct response to single questions) is different. These findings would not represent violations of Assumption \ref{a:rand_sourc}.\footnote{This binomial example is, of course, over-simplistic; most likely, not all the questions have the same difficulty, and the students may not have the exact same probability of responding correctly to all questions. It is only meant to illustrate why, with a discrete score, the probability of being below or above the cutoff may differ even if the mechanism that generated $Z_i$ is random and unrelated to potential outcomes or covariates.}
Imai et al. (2023) invoke two additional assumptions that are necessary to answer our research question.
\begin{assumption} \label{a:ec}
Exclusion restriction: \\
$Y_{i}(z,r) = Y_{i}(z',r) = Y_{i}(r) \hspace{0.3 cm} \mbox{for } z, z',r \in \{0,1\}$.
\end{assumption} This assumption requires that the potential outcomes $Y_{i}(z,r)$
depend only on the recommendation $r$ and not on the source $z$ of the recommendation so that they can be written as $Y_{i}(r)$.
Intuitively, this means that the track in which the student graduates should only be affected by scoring above or below the cutoff through its effect on the teacher's recommendation. A violation of this restriction would occur, for example, if scoring above the cutoff increased the probability of graduating in the high track regardless of whether teachers decide to upgrade. In Section \ref{s:robust} we will show that we obtain similar results if we assume a Homogeneity and a Principal Ignorability assumption instead of the Exclusion Restriction (ER)
assumption, and therefore that our conclusions are robust and stable across these alternative identifying assumptions.
\begin{assumption}
\label{a:strata_monoton}
Monotonicity of $Y_{i}$ with respect to $R_{i}$: \\
$Y_{i}(1) \geq Y_{i}(0).$
\end{assumption}
This assumption plausibly excludes the existence of students, whom we call ``Rebels", who would not complete the high track if it were the recommended track, but who would complete it if the low track were recommended. These students would be characterized by the following set of potential outcomes: $(Y(1),Y(0)) = (0,1)$.\footnote{\cite{Oosterveen2022} allow for the existence of a similar stratum, which they label ``Slow Starters". When estimating lower and upper bounds for the corresponding proportion, they cannot reject that this proportion is zero.
}
Finally, a different type of monotonicity assumption needs to be discussed in our setting.
\begin{assumption}
\label{a:iv_monoton}
Monotonicity of $R_i$ with respect to $Z_i$: \\
$R_{i}(1) \geq R_{i}(0).$
\end{assumption}
This assumption is typically made in Instrumental Variable (IV) settings with non-compliance\footnote{ See \cite{Angrist1994} and \cite{Angrist1996}.} and excludes the existence of ``Defiers'', that is, students who would be recommended the low track when scoring at or above the cutoff and the high track when scoring below. \cite{Imai2023} do not make this assumption. In their context, a judge who decides with the help of an algorithm may release an arrested subject while the same judge, not helped by the algorithm, keeps the same subject in prison, or vice versa. More generally, every set of counterfactual decisions with and without the help of the algorithm is possible in their context.
To put it differently, they do not have (and do not want to have) any prior on which source of a decision is preferable: the fraction of ``Preventable'' cases kept under arrest may be higher when the judge is helped by the algorithm or when the judge decides alone, and their goal is precisely to assess for which one of the two sources the fraction is higher.
\begin{table}[h]
\caption{Non-compliance statistics}
\label{t:non_comp}
\begin{center}
\vspace{0.1cm}
\begin{tabular}{lcccc}
\hline \hline
\\ [-2ex]
TRACK
& $P(R=1|Z=0)$& $P(R=1|Z=1)$& $P(R=1|Z=0)$& $P(R=1|Z=1)$\\
&1 & 2 & 3 & 4 \\
\\ [-2ex]
\hline
\\ [-2ex]
V: BL & .0039& .0352& .0025& .0677\\
V: BL/KL & .0231& .1148& .0205& .1897\\
V: KL & .0209& .1767& .0122& .1952\\
V: KL/GT & .0109& .1233& .0153& .1655\\
V: GT & .0036& .056& .0065& .0574\\
V: GT/HAVO & .0163& .1684& .0206& .1858\\
A: HAVO & .0051& .0603& .0085& .0758\\
A: HAVO/VWO & .0237& .1821& .0193& .188\\ \hline
COHORT & 2015& 2015& 2016& 2016\\
\\[-2ex]
\hline \hline
\end{tabular}
\end{center}
\vspace{-0.2cm}
\begin{minipage}{1\linewidth \setstretch{0.75}}
{\scriptsize Notes: The row headings are the names of the secondary school tracks to which a student can be assigned in the Dutch system, ranked by level of difficulty. The other columns report the probabilities indicated in the column headings. }
\end{minipage}
\end{table}
In our case, the institutional setting is such that an upgrade to a higher track should be possible only if students score at or above the cutoff. Therefore, Defiers should not exist by construction because no student below the cutoff should be eligible for an upgrade. For the same reasons, also Always Takers should not exist. On the other hand, teachers may decide not to upgrade students who score at or above the cutoff, so Never Takers may exist in our setting.\footnote{Table \ref{t:NTandZtilde} in the Online \nameref{ap:NTandZtilde} shows that $\tilde{Z}_i$ correlates positively with the proportion of Never Takers. This finding supports our choice of $\tilde{Z}_i$ as a proxy for the leniency $\pi_i$ of student $i$ teachers.} As anticipated in Section \ref{s:inst}, however, there are rare instances of non-compliance for students scoring below the cutoff. This is shown in Table \ref{t:non_comp}: in the first and third columns, the probability that a student is upgraded ($R_{i} = 1$) even if scoring below the cutoff ($Z_i = 0$) is low but positive in all tracks, ranging between 0.3\% for BL in the 2016 cohort to 2.4\% for HAVO/VWO in the 2015 cohort, and therefore Always Takers do exist in our setting. In the second and fourth columns, the probability that a student is upgraded ($R_{i} =1$) when scoring above the cutoff ($Z_i = 1$) is significantly lower than 1 in all tracks, ranging from 3.5\% for BL in the 2015 cohort to 19.5\% for KL in the 2016 cohort. So, and this is no surprise, Never Takers also exist and are frequent. In general, bridge tracks show a significantly higher probability of upgrading than the other tracks, which makes intuitive sense as $R_{i}$ is defined to be equal to one if the student is upgraded to the higher component of the mixed track.
This evidence indicates that non-compliance is certainly present in our setting. However, we do not see reasons to think that in the rare cases in which students scoring below the cutoff are upgraded (see footnote \ref{foot_def} for the possible reasons), these same students would not be upgraded if scoring above, which makes them Always Takers but not Defiers. Hence, Assumption \ref{a:iv_monoton} of Monotonicity of $R$ with respect to $Z$, excluding the possibility of Defiers, appears plausible. Even if not required to identify causal effects in our setting, we will refer to it for the interpretation of our results.
\begin{table}[ht]
\caption{First stage}
\label{t:first_stage}
\begin{center}
\vspace{0.1cm}
\begin{tabular}{lcccc}
\hline \hline
\\ [-2ex]
Track & $\delta$ & $\delta$& $\delta$ & $\delta$ \\
&1 & 2 & 3 & 4 \\
\\ [-2ex]
\hline
& \hspace{1.5cm} & \hspace{1.5cm} \\ [-2ex]
V: BL & .0319& .0322& .07& .0712\\
& \footnotesize (.0093)& \footnotesize (.0092) & \footnotesize (.0147)& \footnotesize (.0146)\\
V: BL/KL & .0931& .0872& .1722& .1615\\
& \footnotesize (.0265)& \footnotesize (.0276) & \footnotesize (.0308)& \footnotesize (.0304)\\
V: KL & .1639& .1607& .187& .1892\\
& \footnotesize (.0152)& \footnotesize (.0149) & \footnotesize (.0176)& \footnotesize (.0177)\\
V: KL/GT & .1152& .11& .1535& .1564\\
& \footnotesize (.0248)& \footnotesize (.0244) & \footnotesize (.0244)& \footnotesize (.0247)\\
V: GT & .0564& .0561& .0537& .0537\\
& \footnotesize (.0058)& \footnotesize (.0058) & \footnotesize (.0074)& \footnotesize (.0074)\\
V: GT/HAVO & .1579& .1533& .1719& .1675\\
& \footnotesize (.0186)& \footnotesize (.0182)& \footnotesize (.0162)& \footnotesize (.0162)\\
A: HAVO & .0581& .058& .0707& .0705\\
& \footnotesize (.0063)& \footnotesize (.0063) & \footnotesize (.0076)& \footnotesize (.0076)\\
A: HAVO/VWO & .168& .1667& .1732& .1764\\
& \footnotesize (.0167)& \footnotesize (.0166)& \footnotesize (.0158)& \footnotesize (.0155)\\
\\[-2ex]
\hline
\footnotesize COVARIATES $X_i$ & \footnotesize NO & \footnotesize YES& \footnotesize NO & \footnotesize YES \\
\\[-2ex]
\hline
\footnotesize{COHORT} & \footnotesize 2015& \footnotesize 2015& \footnotesize 2016& \footnotesize 2016 \\
\hline \hline
\end{tabular}
\end{center}
\vspace{-0.2cm}
\begin{minipage}{1\linewidth \setstretch{0.75}}
{\scriptsize Notes: The table reports, separately for each track, estimates of parameter $\delta$ in the first stage regression (\ref{e:first_stage}): $R_{i} = \gamma + \delta Z_i + \mu X_i + \xi \tilde{Z}_i + \epsilon_i$. The row headings are the names of the secondary school tracks to which a student can be assigned in the Dutch system, ranked by level of difficulty. The covariates $X$ included in columns 2 and 4 are those described in Table \ref{t:balance}, respectively, for the two cohorts.
}
\end{minipage}
\end{table}
Table \ref{t:first_stage} reports, separately for each track, estimates of the first stage regression
\begin{equation}
\label{e:first_stage}
R_{i} = \gamma + \delta Z_i + \mu X_i + \xi \tilde{Z}_i + \epsilon_i
\end{equation}
where $\tilde{Z}_i$ is the proxy for $\pi_i$ on which we need to condition for the validity of Assumption \ref{a:rand_sourc} and $X_i$ are the nine balanced covariates described in Table \ref{t:balance}.
As expected, the estimates of $\delta$ confirm the existence of imperfect compliance. Considering the covariates $X_i$, the stability of the first-stage estimates with and without them supports Assumption \ref{a:rand_sourc} that the assignment to treatment is as good as random.
\subsection{Students' strata}
\label{s:strata}
Under Assumption 1 (SUTVA) and 3 (Exclusion restriction), and according to \cite{frangakis2002principal}, students can be classified in four strata based on the joint values of the two binary potential outcomes: $Y_{i}(R_{i}=1), Y_{i}(R_{i}=0)$.
\begin{itemize}
\item[AH:] $(Y_{i}(1),Y_{i}(0)) = (1,1)$\\
These are {\it Always High} students who always complete the high track independently of the recommendation they receive.
\item[AL:] $(Y_{i}(1),Y_{i}(0)) = (0,0)$\\
These are {\it Always Low} students who always complete the low track independently of the recommendation they receive.
\item[H:] $(Y_{i}(1),Y_{i}(0)) = (1,0)$\\
These are {\it Helpable} students who complete the high track if this is the track recommended to them by primary school teachers, and who would not do the same otherwise.
\item[R:] $(Y_{i}(1),Y_{i}(0)) = (0,1)$\\
These are {\it Rebel} students who complete the high track if recommended for the low track by primary school teachers, and the low track if recommended for the high track.
\end{itemize}
These principal strata, which we denote with $J\in\{AL, AH, H, R\}$, will play a crucial role in the definition of the metric we propose in Section \ref{s:metric} to measure the quality of track recommendations.
\begin{landscape}
\begin{table}[ht]
\caption{The Principal Strata}
\label{t:strata2}
\begin{center}
\vspace{0.1cm}
\begin{tabular}{ccccccc}
\hline \hline
\\ [-2ex]
Strata
& \scriptsize $\underbrace{Y_{i}(0,0)= Y_{i}(1,0)= Y_{i}(0) }$
& \scriptsize $\underbrace{Y_{i}(0,1)= Y_{i}(1,1)= Y_{i}(1) }$
& Strata
& \scriptsize $R_{i}(0)$
& \scriptsize$R_{i}(1)$
& Notes
\\
\scriptsize $R_{i} \rightarrow Y_{i} $
& Exclusion restriction
& Exclusion restriction
& \scriptsize $Z_i \rightarrow R_{i} $
&
&
&
\\
\\ [-2ex]
1 & 2 & 3 & 4 & 5 & 6 &7\\
\\ [-2ex]
\hline
\\ [-2ex]
AL & 0 & 0 & NT & 0 & 0
& \\
\\ [-2ex]
& & & AT & 1 & 1 \\
\\ [-2ex]
& & & C & 0 & 1\\
\\ [-2ex]
& & &\color{blue} D & \color{blue} 1 & \color{blue} 0
& \color{blue} May $\exists$ w/out Monotonicity $Z_i \rightarrow R_{i}$
\\
\\ [-2ex]
\hline
\\ [-2ex]
AH & 1 & 1 & NT & 0 & 0
& \\
\\ [-2ex]
& & & AT & 1 & 1 \\
\\ [-2ex]
& & & C & 0 & 1\\
\\ [-2ex]
& & & \color{blue} D & \color{blue} 1 & \color{blue} 0
& \color{blue} May $\exists$ w/out Monotonicity $Z_i \rightarrow R_{i}$ \\
\\ [-2ex]
\hline
\\ [-2ex]
H & 0 & 1 & NT & 0 & 0
& \\
\\ [-2ex]
& & & AT & 1 & 1 \\
\\ [-2ex]
& & & C & 0 & 1\\
\\ [-2ex]
& & &\color{blue} D &\color{blue} 1 &\color{blue} 0
& \color{blue} May $\exists$ w/out Monotonicity $Z_i \rightarrow R_{i}$ \\
\\ [-2ex]
\hline
\\ [-2ex]
\color{red} R &
\color{red} 1 &
\color{red} 0 &
\color{red} NT &
\color{red} 0 &
\color{red} 0
&
\color{red}
$\nexists$ because Monotonicity $R_{i} \rightarrow Y_{i}$ holds
\\
\\ [-2ex]
& & & \color{red} AT & \color{red} 1 & \color{red} 1 \\
\\ [-2ex]
& & & \color{red} C & \color{red} 0 & \color{red} 1\\
\\ [-2ex]
& & & \color{red} D &
\color{red} 1 & \color{red} 0
& \\
\\ [-2ex]
\hline \hline
\end{tabular}
\end{center}
\vspace{-0.2cm}
\begin{minipage}{1\linewidth \setstretch{0.75}}
{\scriptsize Notes: The four panels refer to the four Principal Strata listed in the first column of the table and defined by the possible values of $R_{i}$ and $Y_{i}$: the Always-Low (AL), the Always-High (AH), the Helpable (H) and the Rebels (R). Columns 2 and 3 describe the values of the potential outcomes $Y_{i}$ when the treatment $R_{i}$ is equal to $0$ or $1$, respectively, and the Exclusion Restriction assumption holds. Each one of these Principal Strata is divided into additional strata defined by the possible values of $Z_i$ and $R_{i}$, as listed in the fourth column: the Never Takers (NT), the Always Takers (AT), the Compliers (C) and the Defiers (D). Columns 5 and 6 describe the values of the potential treatments $R_{i}$ when the assignment to treatment $Z_i$ is equal to $0$ or $1$, respectively. The Notes in the first three rows of the last column make clear that Defiers may or may not exist depending on whether the Monotonicity Assumption \ref{a:iv_monoton} of $R_{i}$ with respect to $Z_i$ holds. The last note in the fourth row states that given the maintained Monotonicity Assumption \ref{a:strata_monoton} of $Y_{i}$ with respect to $R_{i}$, the Rebels do not exist.
}
\end{minipage}
\end{table}
\end{landscape}
Under the same assumptions, the more conventional classification in strata that characterizes compliance types is also possible, based on the joint values of the two binary potential treatment values: $R_{i}(Z_i=1),R_{i}(Z_i=0)$:
\begin{itemize}
\item[AT:] $(R_{i}(1),R_{i}(0)) = (1,1)$\\
These are {\it Always Takers} who are always recommended the high track, regardless of their score around the cutoff.
\item[NT:] $(R_{i}(1),R_{i}(0)) = (0,0)$\\
These are {\it Never Takers} who are never recommended the high track, independently of where they score around the cutoff.
\item[C:] $(R_{i}(1),R_{i}(0)) = (1,0)$\\
These are {\it Compliers} who are recommended the high track if they score at the cutoff, and the low track if they score below.
\item[D:] $(R_{i}(1),R_{i}(0)) = (0,1)$\\
These are {\it Defiers} who are recommended the low track if they score at the cutoff, and the high track if they score below.
\end{itemize}
We denote with $G\in\{AT, NT, C, D\}$ these strata.
Table \ref{t:strata2} describes the relationship between these basic principal strata. Under Assumption \ref{a:strata_monoton} (Monotonicity of $Y_{i}$ with respect to $R_{i}$), the stratum of the Rebels, $J=R$ in red, does not exist irrespective of the student belonging to any of the strata $G$, while under Assumption \ref{a:iv_monoton} (Monotonicity of $R_{i}$ with respect to $Z_{i}$), the stratum of Defiers, $G=D$ in blue, does not exist in any of the strata $J$.
\subsection{The metric to compare recommendations}
\label{s:metric}
Following \cite{Imai2023}, the metric that we propose to compare recommendations and establish which one is preferable relies on the classification of students in Principal Strata described in the previous section, combined with the following criteria:
\begin{enumerate}
\item If a student is able to complete a more difficult track, it is better that she is allowed to do it for both
\begin{itemize}
\item[1a)] herself, because she can obtain more desirable lifetime outcomes
\item[1b)] and for society, if there are collective gains from a more educated community.
\end{itemize}
\item Changing tracks is costly for students and for the education system.
\end{enumerate}
The first criterion defines what constitutes a better recommendation for Helpable (H) students. In their case, the goal of teachers should be to recommend the high track to the largest fraction of these students because, as a result of this recommendation, they will complete precisely this track, achieving a better outcome for themselves as well as for society. Therefore, we need to estimate the difference between ``the fraction of H students who are recommended a high first-year track by teachers who can upgrade based on students' test scores" and ``the analogous fraction when teachers do not see test scores". If this difference is positive, test scores help teachers improve the track advising they provide to their students.
The second criterion defines instead what constitutes a better recommendation for the Always Low (AL) and Always High (AH) students. In their cases, track advising should aim to minimize the number of track changes that these two types of students experience during secondary school, if they start on a track that differs from the one they will ultimately complete. Therefore, the best recommendation is to send all AL students to the low track and all AH students to the high track.
To formalize this intuitive characterization of what constitutes a better recommendation, consider the following Average Principal Causal Effect:
\begin{equation}
\label{e:apce_J}
APCE_{J} = E(R_{i}(1) - R_{i}(0) | \mbox{ student $i$ is in } J)
\end{equation}
The $APCE_{J}$ is the difference between the fraction of students in stratum $J$ who are recommended the high track by teachers who can upgrade based on test scores ($Z_i=1$) minus the fraction receiving the same recommendation by teachers who have not seen the test scores ($Z_i=0$). To put it differently, the $APCE_{J}$ is the Intention To Treat effect of the source of recommendations on their content in stratum $J$.
Note that each $APCE_{J}$ is also equal to the difference between the fractions of Compliers and Defiers:
\[APCE_{J}
=P(G_i=\mbox{C}|\mbox{ student $i$ is in } J)-P(G_i=\mbox{D}|\mbox{ student $i$ is in } J).\]
If Monotonicity of $R_{i}$ with respect to $Z_i$ (Assumption (ref)) holds, Defiers do not exist, and the three $APCE_{J}$ are all non-negative by construction, being equal to the proportion of Compliers. In this case we could conclude\footnote{As already mentioned when we introduced Assumption (ref) in the previous section, we emphasize that this assumption allows to interpret $APCE_{J}$ as the proportion of Compliers in stratum $j$, but it is not necessary for identification.} that test scores help teachers improve their recommendations if the $APCE_H$ and the $APCE_{AH}$ are positive while the $APCE_{AL}$ is equal to zero. Having established that the $APCE_{J}$ are informative about whether test scores help teachers improve the quality of track advising, the next section shows how these population parameters can be identified using data generated by the quasi-experiment under study.
\subsection{Identification of the $APCE_J$}
\subsubsection*{Estimands for Helpable (H) students}
Adapting Theorem 1 of Imai2023 to this context, it is possible to show that\footnote{For the derivation of this equation and all the others in this section, see the proofs in the Online \nameref{ap:apce_proof}.
}
\begin{equation}
APCE_{H} = \frac
{E(Y_{i}|Z_i=1) - E(Y_{i}|Z_i=0) }
{Pr(Y_{i}(1)=1) - Pr(Y_{i}(0)=1)}.
\end{equation}
where the terms in the numerator do not depend on missing potential outcomes and can be estimated with their sample analogs.
The denominator is instead the proportion of Helpable students, given that Rebels are excluded because of the Monotonicity Assumption (ref), and can only be partially identified. In some settings, even if principal strata are latent, their proportions can be point identified under specific assumptions. This is the case, for example, of the proportion of compliers in IV settings, because $Z$ is random. In our case, because the principal strata are defined by the joint values of $Y(r)$ and $R$ is confounded, the strata proportions are not point-identified unless we impose further assumptions on the distribution of potential outcomes and the recommendation $R$.
To address this problem, Imai2023 propose to derive non-parametric bounds for the $APCE_{H}$
by bounding the terms $Pr(Y_{i}(R)=1)$ in the denominator of ((ref)).
Using Assumption 1 and the Law of Total Probability, it is possible to show that these bounds are:
\begin{eqnarray}
\max_z Pr(Y_{i} = 1,R_{i} =0 | Z_i=z)
\leq
& Pr(Y_{i}(0)=1 )&
\leq
\min_z Pr(Y_{i} = 1 | Z_i=z)
\\ \nonumber
\max_z Pr(Y_{i} = 1 | Z_i=z)
\leq
& Pr(Y_{i}(1)=1) &
\leq
1- \max_z Pr(Y_{i} = 0, R_{i}=1 | Z_i=z)
\end{eqnarray}
These results imply that the lower bound of $APCE_{H}$ is
\begin{equation}
\frac{E(Y_{i}|Z_i=1)-E(Y_{i}|Z_i=0)}{1- \max_z Pr(Y_{i} = 0, R_{i}=1 | Z_i=z) - \max_z Pr(Y_{i} = 1,R_{i} =0 | Z_i=z)}
\end{equation}
while the upper bound is
\begin{equation}
\frac{E(Y_{i}|Z_i=1)-E(Y_{i}|Z_i=0)}{\max_z Pr(Y_{i} = 1 | Z_i=z) -\min_z Pr(Y_{i} = 1 | Z_i=z) }.
\end{equation}
An alternative identification and estimation strategy that we can adopt is instead based on the assumption of unconfoundedness of $R$ with respect to $Y$:
\begin{assumption}
Unconfoundedness of $R$: \\
$Y_i(r) \perp\!\!\!\!\perp R_i(z) \mid X_i = x, Z_i=z$ and \\
$0 <Pr(R_i(z) = r \mid X_i = x, Z_i =z) < 1 \quad \forall z \in \{0,1\}, x \in \mathcal{X}, r \in \{0,1\} $
\end{assumption}
This assumption, which allows one to point identify the $APCE_J$, implies that, conditioning on observable covariates $X$, a teacher's decision to upgrade is independent of potential outcomes. In other words, covariates $X$ contain all the information teachers use when making their decision, and the content of the recommendation given this information is random.\footnote{This assumption would be violated
if, for example, teachers also relied on the student’s behavior in class, which we do not observe, when taking their upgrade decisions. As we will see, we achieve similar results under the different identification assumptions we adopt.
}
Identification under this assumption is formally discussed in the Online \nameref{ap:apce_proof}.
\subsubsection*{Estimands for Always High (AH) students}
As for the AH students, we show in
the Online \nameref{ap:apce_proof} that
\begin{eqnarray}
ACPE_{AH} = \frac{Pr(R_{i}=0,Y_{i}=1 \mid Z_i =0) - Pr(R_{i}=0,Y_{i}=1 \mid Z_i =1)}{Pr(Y_{i}(0) = 1)},
\end{eqnarray}
where, once again, the terms in the numerator do not depend on missing potential outcomes and can be estimated using their sample analogs, whereas further assumptions are needed for the denominator. To bound this denominator, note that:
\begin{equation}
\begin{aligned}
\max_{z} Pr(Y_{i} = 1, R_{i} = 0 \mid Z_i = z) & \leq Pr(Y_{i}(0)=1) \leq \min_{z} Pr(Y_{i} = 1 \mid Z_i = z) \\
\end{aligned}
\nonumber
\end{equation}
Therefore the lower bound of $APCE_{AL}$ is
\begin{equation}
\frac{Pr(R_{i}=0,Y_{i}=1 \mid Z_i =0) - Pr(R_{i}=0,Y_{i}=1 \mid Z_i =1)}{\min_{z} Pr(Y_{i} = 1 \mid Z_i = z)}
\end{equation}
while the upper bound is
\begin{equation}
\frac{Pr(R_{i}=0,Y_{i}=1 \mid Z_i =0) - Pr(R_{i}=0,Y_{i}=1 \mid Z_i =1)}{\max_{z} Pr( R_{i} = 0,Y_{i} = 1 \mid Z_i = z)}
\end{equation}
Also, in the case of AH students, as with their H peers, we will compare the bounds obtained under the above identification strategy with point estimates obtained by assuming the unconfoundedness of $R$ (Assumption (ref)), as shown in the Online Appendix \nameref{ap:apce_proof}.
\subsubsection*{Estimands for Always Low (AL) students}
Finally, considering the AL students, the Online \nameref{ap:apce_proof} shows that
\begin{eqnarray}
APCE_{AL} &=& \frac{Pr(R_{i}=1,Y_{i}=0 \mid Z_i=1) - Pr(R_{i}=1,Y_{i}=0 \mid Z_i=0)}{1 - Pr(Y_{i}(1)=1)}.
\end{eqnarray}
where, also in this case, the terms in the numerator do not depend on missing potential outcomes and can be estimated with their sample analogs, while the estimation of the denominator requires further assumptions. To bound this denominator, note that:
\begin{equation}
\begin{aligned}
\max_{z} Pr(Y_{i} = 1 \mid Z_i = z) & \leq Pr(Y_{i}(1) = 1 ) \leq 1 - \max_{z} Pr(Y_{i} = 0, R_{i} = 1 \mid Z_i =z)
\end{aligned}
\nonumber
\end{equation}
Therefore, the lower bound of $APCE_{AL}$ is
\begin{equation}
\frac{Pr(R_{i}=1,Y_{i}=0 \mid Z_i=1) - Pr(R_{i}=1,Y_{i}=0 \mid Z_i=0)}{1 - \max_{z} Pr(Y_{i}\mid Z_i = z) }
\end{equation}
while the upper bound is
\begin{equation}
\frac{Pr(R_{i}=1,Y_{i}=0 \mid Z_i=1) - Pr(R_{i}=1,Y_{i}=0 \mid Z_i=0)}{\max_{z} Pr( R_{i} = 1,Y_{i} = 0 \mid Z_i =z)}.
\end{equation}
Once again, as for the $APCE_{H}$ and the $APCE_{AH}$, also the $APCE_{AL}$ can be point identified under unconfoundedness of $R$ (Assumption (ref)), as shown in the Online \nameref{ap:apce_proof}.
\section{Test scores and quality of teachers' advice: evidence}
\subsection{Estimates for the three strata}
\subsubsection*{Estimates for Helpable (H) students}
In the first column of Table (ref), we present estimates of the numerator of the $APCE_H$ estimand in the RHS of equation ((ref)), separately for each track and combined for the three aggregate tracks: Vocational, Academic and All.\footnote{Here and in what follows, each population estimand has been estimated separately for each track and cohort. The estimates reported in the tables for aggregate tracks have been obtained as weighted averages with weights based on the relative sample sizes of each cell. For this reason, the remaining tables report only one set of results, rather than separate results by cohort.
Standard errors (reported in parentheses) are bootstrapped (with 1000 repetitions) separately in each track, and then appropriately aggregated with the same weights. The estimates in the tables reported in the text include the covariates $\tilde{Z}_i$ and $X_i$ (see Table (ref)).
The Online \nameref{ap:3strata} reports tables with corresponding estimates obtained without the balanced covariates $X_i$ (but including $\tilde{Z}_i$). The inclusion or exclusion of the balanced covariates does not significantly alter the estimates, as expected. This evidence supports the validity of Assumption (ref) (Randomization of the source).}
The point estimates of this numerator are strictly positive in all tracks. Students scoring at the cutoff, who can be upgraded by their primary school teachers based on test scores, are 1.8--to--10% more likely to complete a higher track than students who score below the cutoff. These point estimates are statistically different from zero except for the BL track. The estimates for the three aggregate tracks in the last rows of the table are also significant and equal to 4.4%.
\begin{table}
\caption{Estimates for the $APCE_{H}$, with covariates}
\begin{center}
\begin{tabular}{l*{5}{c}}
\hline\hline
\\ [-2ex]
Track & Numerator& Lower & Upper & Lower & Upper\\
& $APCE_H$ & bound & bound & bound & bound \\
& & denominator & denominator& $APCE_H$ & $APCE_H$ \\
& & $APCE_H$ & $APCE_H$& & \\
& 1& 2& 3 & 4 & 5 \\
\\ [-2ex]
\hline
\\ [-2ex]
V: BL & .0251& .0251& .6534& .0381& 1\\
& { (.0167)}& { (.0154)}& { (.0126)}& { (.0263)}& { -}\\
V: BL/KL & .0773& .0773& .2877& .2757& 1\\
& { (.0294)}& { (.0279)}& { (.0222)}& { (.1029)}& { -}\\
V: KL & .0996& .0996& .6984& .1426& 1\\
& { (.0159)}& { (.0159)}& { (.0112)}& { (.0229)}& { -}\\
V: KL/GT & .0256& .0256& .3967& .0644& 1\\
& { (.0238)}& { (.0217)}& { (.0232)}& { (.0564)}& { -}\\
V: GT & .0182& .0182& .7607& .0237& 1\\
& { (.0077)}& { (.0077)}& { (.0076)}& { (.0104)}& { -}\\
V: GT/HAVO & .0526& .0526& .5089& .1021& 1\\
& { (.0164)}& { (.0159)}& { (.0138)}& { (.0309)}& { -}\\
A: HAVO & .0349& .0349& .8082& .0432& 1\\
& { (.0076)}& { (.0076)}& { (.0065)}& { (.0096)}& { -}\\
A: HAVO/VWO & .0642& .0642& .5451& .117& 1\\
& { (.0156)}& { (.0154)}& { (.0126)}& { (.0276)}& { -}\\
\hline
VOCATIONAL & .0445& .0445& .6486& .0686& 1\\
& { (.0058)}& { (.0059)}& { (.005)}& { (.009)}& { -}\\
ACADEMIC & .0442& .0442& .7251& .0609& 1\\
& { (.0072)}& { (.0071)}& { (.0059)}& { (.0099)}& { -}\\
ALL & .0444& .0444& .6796& .0653& 1\\
& { (.0045)}& { (.0045)}& { (.0038)}& { (.0066)}& { -}\\
\hline\hline
\end{tabular}
\end{center}
\begin{minipage}{1\linewidth \setstretch{0.75} }
{\scriptsize Notes: The row headings are the names of the secondary school tracks to which a student can be assigned in the Dutch system, ranked by level of difficulty. The last three rows of the table aggregate the Vocational (VMBO*), the Academic (HAVO*), and All tracks, respectively. With reference to the $APCE_H$ estimand defined in equation ((ref)), the table reports for the outcome $Y$ and separately for each track, estimates of the numerator (in column 1), of the lower and upper bounds of the denominator (in columns 2 and 3), and of the lower bound of the $APCE_H$ (in column 4).
Column 5 reports the upper bound of the $APCE_H$ as equal to 1 because it is the ratio between $E(Y_{i}|Z_i=1)-E(Y_{i}|Z_i=0)$ and $\max_z Pr(Y_{i} = 1 | Z_i=z) -\min_z Pr(Y_{i} = 1 | Z_i=z)$ which are exactly equal under Assumptions (ref) and (ref). See footnote (ref) for the intuition.
The definitions of these two last statistics are in equations ((ref)) and ((ref)), respectively. The aggregations for the last three rows are performed as follows: point estimates of the numerator and of the lower and upper bounds of the denominator are created by weighing each track in the corresponding aggregate by the relative population in each track. For the lower and upper bounds of the $APCE_H$, the aggregated estimates for the lower and upper bounds of the denominators are used directly. Standard errors are computed using a block bootstrap at the school level with 1000 repetitions; each quantity in every iteration is calculated in the same way.
The specifications in all columns include the covariates $\tilde{Z}_i$ and $X_i$ (see Table (ref)).
The Online \nameref{ap:3strata} reports tables with corresponding estimates obtained without the balanced covariates $X_i$ (but including $\tilde{Z}_i$). The inclusion or exclusion of the balanced covariates does not change the estimates in a relevant way, as expected, supporting the validity of Assumption (ref) (Randomization of the source).}
\end{minipage}
\end{table}
Recall that the denominator of the $APCE_H$ estimand in equation ((ref)) is certainly positive and equal to the proportion of Helpable students, given Assumption (ref) of Monotonicity of $Y_{i}$ with respect to $R_{i}$. Therefore, a positive numerator supports the conclusion that the information provided by test scores helps recommenders give better advice in the sense that, thanks to this information, they become better at detecting Helpable students and directing them into the most challenging secondary school track that they can successfully complete.
To confirm this conclusion and assess also quantitatively by how much test scores help teachers achieve this goal, estimates of the $APCE_H$ are needed, but as explained in the previous section, the denominator in the RHS of equation ((ref)), which is the proportion of $H$ students, can only be bounded without additional assumptions. Estimates of the lower and upper bounds of this denominator are reported in columns 2 and 3 of Table (ref). Focusing on the aggregate tracks in the last three rows, the proportion of $H$ ranges between approximately 4% and 73%.
As a consequence of the large gap between the estimated bounds of the denominator, also the bounds of the $APCE_H$ reported in columns 4 and 5 of Table (ref) are substantially different one from the other, ranging between 6% and 100% for the three aggregate tracks.\footnote{Note that the upper bound of the $APCE_H$ derived in equation ((ref)) and reported in column 5 of Table (ref) is 1 by construction because it is the ratio between $E(Y_{i}|Z_i=1)-E(Y_{i}|Z_i=0)$ and $\max_z Pr(Y_{i} = 1 | Z_i=z) -\min_z Pr(Y_{i} = 1 | Z_i=z)$, which are exactly equal under the two Monotonicity Assumptions (ref) and (ref). The intuition is that the intention-to-treat (${E}[Y_i \mid Z_i=1]-{E}[Y_i\mid Z_i=0]$) at the numerator of this ratio identifies the proportion of students who are compliers and are also helpable. Hence, the minimum plausible share of all Helpable students at the denominator of this ratio cannot be smaller than the share of Helpable who are also compliers at the numerator, and the upper bound of the $APCE_H$ must be 1 by definition.
Note also that when the share of Helpable who are also compliers is effectively equal to zero or very small, we may estimate it to be negative because of small sample variability. In tracks where this occurs, we round this proportion to zero. However, even when this proportion is infinitesimally small, it is still the case that the upper bound for the $APCE_H$ is 1 for the reason explained above.
}
However, what matters most from the viewpoint of our research question is the lower bound of the $APCE_H$, which allows us to conclude that the information provided by test scores improves the quality of recommendations by at least 6% as measured by their capacity to push Helpable students into the high track.
This is a remarkable finding, particularly if the proportion of Helpable students in the population can be as high as 80%, as suggested by the estimated upper bound of the $APCE_H$ denominator in column 3 of Table (ref).
\begin{table}
\caption{Estimates of the $APCE_{AH}$, with covariates}
\begin{center}
\begin{tabular}{l*{5}{c}}
\hline\hline
\\ [-2ex]
Track & Numerator& Lower & Upper & Lower & Upper\\
& $APCE_{AH}$ & bound & bound & bound & bound \\
& & denominator & denominator& $APCE_{AH}$ & $APCE_{AH}$ \\
& & $APCE_{AH}$ & $APCE_{AH}$& & \\
& 1& 2& 3 & 4 & 5 \\
\\ [-2ex]
\hline
\\ [-2ex]
V: BL & .0004& .317& .3189& .0011& .0011\\
& { (.0118)}& { (.0111)}& { (.0111)}& { (.036)}& { (.036)}\\
V: BL/KL & 0227& .6858& .705& .0323& .0332\\
& { (.0243)}& { (.0209)}& { (.0214)}& { (.0336)}& { (.034)}\\
V: KL & .0083& .2289& .2392& .0368& .0379\\
& { (.0101)}& { (.0086)}& { (.0091)}& { (.0401)}& { (.0416)}\\
V: KL/GT & .0745& .5681& .581& .1288& .1319\\
& { ( .0297)}& { (.0206)}& { (.0198)}& { (.0499)}& { (.05)}\\
V: GT & .0097& .1938& .1938& .0482& .0482\\
& { (.006)}& { (.0055)}& { (.0055)}& { (.03)}& { (.03)}\\
V: GT/HAVO & .0439& .4169& .4262& .1023& .1041\\
& { (.0165)}& { (.0113)}& { (.0115)}& { (.0375)}& { (.0379)}\\
A: HAVO & 0& .1283& .1283& 0& 0\\
& { (.0006)}& { (.004)}& { (.004)}& { (.0047)}& { (.0047)}\\
A: HAVO/VWO & .0281& .3667& .3777& .074& .0762\\
& { (.0146)}& { (.0105)}& { (.0109)}& { (.037)}& { (.0379)}\\
\\ [-2ex]
\hline
\\ [-2ex]
VOCATIONAL & .0188& .2988& .3044& .0618& .063\\
& { (.0049)}& { (.004)}& { (.0041)}& { (.0157)}& { (.0159)}\\
ACADEMIC & .0089& .2035& .207& .0429& .0436\\
& { (.0046)}& { (.0042)}& { (.0043)}& { (.0219)}& { (.0223)}\\
ALL & .0148& .2601& .2649& .0558& .0568\\
& { (.0034)}& { (.0029)}& { (.003)}& { (.0126)}& { (.0128)}\\
\\ [-2ex]
\hline\hline
\end{tabular}
\end{center}
\begin{minipage}{1\linewidth \setstretch{0.75} }
{\scriptsize Notes: The row headings are the names of the secondary school tracks to which a student can be assigned in the Dutch system, ranked by level of difficulty. The last three rows of the table aggregate the Vocational (VMBO*), the Academic (HAVO*), and All tracks, respectively. With reference to the $APCE_{AH}$ estimand defined in equation ((ref)), the table reports for the outcome $Y$ and separately for each track, estimates of the numerator (in column 1), of the lower and upper bounds of the denominator (in columns 2 and 3), and of the lower and upper bounds of the $APCE_{AH}$ (in columns 4 and 5). The definitions of these two last statistics are in equations ((ref)) and ((ref)), respectively. The aggregations for the last three rows are performed as follows: point estimates of the numerator and of the lower and upper bounds of the denominator are created by weighing each track in the corresponding aggregate by the relative population in each track. For the lower and upper bounds of the $APCE_{AH}$, the aggregated estimates for the lower and upper bounds of the denominators are used directly. Standard errors for each quantity are computed using a block bootstrap at the school level with 1000 repetitions; each quantity in every iteration is calculated in the same way.
The specifications in all columns include the covariates $\tilde{Z}_i$ and $X_i$ (see Table (ref)).
The Online \nameref{ap:3strata} reports tables with corresponding estimates obtained without the balanced covariates $X_i$ (but including $\tilde{Z}_i$). The inclusion or exclusion of the balanced covariates does not change the estimates in a relevant way, as expected, supporting the validity of Assumption (ref) (Randomization of the source).}
\end{minipage}
\end{table}
\subsubsection*{Estimates for Always High (AH) students}
Moving to the stratum of Always High students, the estimates for them are reported in Table (ref). In their case, the best source of advice is the one that recommends the high track to the highest number of them (possibly to all of them) in order to minimize the cost of track changes they would incur if they started secondary school in the low track. Recall, in fact, that these students always complete the high track no matter where they start.
Focusing for brevity just on Columns 4 and 5 of the table, the gap between the lower and the upper bounds of the $APCE_{AH}$ is remarkably small in almost all tracks, allowing for a very precise estimation of this population parameter.\footnote{In the case of HAVO track, the sample analog of the numerator in the RHS of equation ((ref)) is close to zero but has a negative sign, possibly due to small sample variability. We approximate it with 0.} The intuition for why we obtain such a precise result is that we can be confident in the population proportion of AH students (i.e. the denominator) as those with $R = 0$ and $Y = 1$ are for sure AH, and measuring this under the cutoff (i.e. for those with $Z = 0$) gives a precise estimate of the population parameter, showing that the population proportion of AH is approximately 26%.\footnote{More specifically, because it is extremely rare to have $Z = 0$ and $R = 1$ (which can only occur due to the institutional complications explained in footnote (ref)) we know that among those with $Z = 0$, those with $R = 0$ and $Y=1$ are AH for sure, while those with $R = 0$ and $Y = 0$ are for sure not AH, and these two categories together form virtually the whole population. This means that the fraction of AH students can be estimated very precisely, as $Z$ is randomized and the same fraction should exist above the threshold. } In the eight basic tracks, the point estimates of the $APCE$ range from 0 to 13%. When the tracks are aggregated in the Vocational, Academic, and All tracks (last three rows of the table), the $APCE_{AH}$ is estimated to be about 6.2%, 4.3%, and 5.6%, respectively, and statistically significant. These results indicate that when teachers are allowed to upgrade recommendations based on test scores, they send a significantly higher fraction of AH students to the high track, thereby reducing the costly track changes this group would otherwise incur.
\subsubsection*{Estimates for Always Low (AL) students}
In the case of Always Low students, instead, our estimates of the $APCE_{AL}$ suggest that test score information does not help teachers give better recommendations. These estimates are
\begin{table}[!]
\caption{Estimates of the $APCE_{AL}$, with covariates}
\begin{center}
\begin{tabular}{l*{5}{c}}
\hline\hline
\\ [-2ex]
Track & Numerator& Lower & Upper & Lower & Upper\\
& $APCE_{AL}$ & bound & bound & bound & bound \\
& & denominator & denominator& $APCE_{AL}$ & $APCE_{AL}$ \\
& & $APCE_{AL}$ & $APCE_{AL}$& & \\
& 1& 2& 3 & 4 & 5 \\
\\ [-2ex]
\hline
\\ [-2ex]
V: BL & .0271& .028& .6576& .0409& .9765\\
& {(.0056)}& {(.0055)}& {(.0125)}& {(.0085)}& {(.0266)}\\
V: BL/KL & .0239& .0266& .2177& .1126& .8981\\
& {(.0084)}& {(.008)}& {(.0208)}& {(.0404)}& .{(1399)}\\
V: KL & .0658& .0727& .6612& .0991& .9026\\
& {(.0071)}& {(.0069)}& {(.0126)}& {(.0105)}& {(.03)}\\
V: KL/GT & .0335& .0352& .3934& .0853& .9569\\
& {(.0084)}& {(.0083)}& {(.0194)}& {(.0212)}& {(.0599)}\\
V: GT & .0377& .041& .7899& .0477& .9193\\
& {(.0038)}& {(.0036)}& {(.0072)}& {(.0048)}& {(.0253)}\\
V: GT/HAVO & .0644& .0742& .5213& .1238& .8682\\
& {(.0082)}& {(.0077)}& {(.0128)}& {(.0156)}& {(.0409)}\\
A: HAVO & .0487& .0539& .8464& .0576& .9035\\
& {(.0042)}& {(.004)}& {(.0058)}& {(.005)}& {(.0267)}\\
A: HAVO/VWO & .0788& .0883& .5581& .141& .8928\\
& {(.0077)}& {(.0074)}& {(.0126)}& {(.0133)}& {(.0278)}\\
\\ [-2ex]
\hline
\\ [-2ex]
VOCATIONAL & .0458& .0506& .652& .0703& .9059\\
& {(.0027)}& {(.0025)}& {(.005)}& {(.004)}& {(.016)}\\
ACADEMIC & .0582& .0648& .7554& .0771& .8988\\
& {(.0037)}& {(.0035)}& {(.0055)}& .0048& {(.0192)}\\
ALL & .0508& .0563& .694& .0733& .9026\\
& {(.0022)}& {(.0021)}& {(.0037)}& {(.0032)}& {(.0126)}\\
\\ [-2ex]
\hline\hline
\end{tabular}
\end{center}
\begin{minipage}{1\linewidth \setstretch{0.75} }
{\scriptsize Notes: The row headings are the names of the secondary school tracks to which a student can be assigned in the Dutch system, ranked by level of difficulty. The last three rows of the table aggregate the Vocational (VMBO*), the Academic (HAVO*), and All tracks, respectively. With reference to the $APCE_{AL}$ estimand defined in equation ((ref)), the table reports for the outcome $Y$ and separately for each track, estimates of the numerator (in column 1), of the lower and upper bounds of the denominator (in columns 2 and 3), and of the lower and upper bounds of the $APCE_{AL}$ (in columns 4 and 5). The definitions of these two last statistics are in equations ((ref)) and ((ref)), respectively. The aggregations for the last three rows are performed as follows: point estimates of the numerator and of the lower and upper bounds of the denominator are created by weighing each track in the corresponding aggregate by the relative population in each track. For the lower and upper bounds of the $APCE_{AL}$, the aggregated estimates for the lower and upper bounds of the denominators are used directly. Standard errors for each quantity are computed using a block bootstrap at the school level with 1000 repetitions; each quantity in every iteration is calculated in the same way.
The specifications in all columns include the covariates $\tilde{Z}_i$ and $X_i$ (see Table (ref)).
The Online \nameref{ap:3strata} reports tables with corresponding estimates obtained without the balanced covariates $X_i$ (but including $\tilde{Z}_i$). The inclusion or exclusion of the balanced covariates does not change the estimates in a relevant way, as expected, supporting the validity of Assumption (ref) (Randomization of the source).}
\end{minipage}
\end{table}
reported in Table (ref). The best source of advice for these students is the one that recommends the high track to the {\it lowest} number of them (possibly to none), in order to minimize the cost of the track changes they would incur if they started secondary school in the high track. Recall that these students can only complete the low track, no matter where they start.
Focusing again for brevity just on the last two columns of the table, Column 4 shows that the estimated lower bound of the $APCE_{AL}$ is positive and statistically different from zero in many tracks, ranging from 4% for BL to 14% for HAVO/VWO. Considering the aggregate tracks in the last three rows of the table, the $APCE_{AL}$ is at least as high as 7% and significantly different from zero in all of them.
Therefore, it appears that the information provided by test scores leads teachers to upgrade AL students into the high track. The consequence is a deterioration of the quality of recommendations because AL students cannot complete the high track.
A possible reason for this finding is that primary school teachers do not incur any costs resulting from secondary school track changes and are therefore less sensitive to these consequences of the advice they offer to their students.
\subsection{A comprehensive evaluation}
The evidence in the previous section suggests that, in the case of Helpable and Always High students, test score information helps teachers improve the quality of their recommendations in terms of the two criteria that we have adopted (directing students towards the most difficult track they can successfully complete and reducing track changes). On the contrary, in the case of the Always Low, the positive $APCE_{AL}$ estimate indicates that the same information reduces this quality because it induces AL students to enroll in the high track that they later cannot complete.
The classification framework proposed by benmichael2024does provides a comprehensive evaluation of whether it is a good idea to offer teachers the possibility to upgrade based on test score information, depending on how much the policy-maker cares about H and AH students compared to their AL peers. Define with $l_{ry}$ the loss deriving from a recommendation $R=r \in\{0,1\}$ given to a student with potential outcome $Y(1)= y \in\{0,1\}$. The four possible values of this loss are represented in the following Confusion Matrix:
\begin{center}
\textbf{Confusion Matrix} \\
\begin{tabular}{c|c|c}
& Negative ($R=0$) & Positive ($R=1$) \\ \hline
Negative ($Y(1) = 0$) & \begin{tabular}[c]{@c@}{True Negative}\\ $l_{00}=0$\end{tabular} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}{False Positive}\\ $l_{10}>0$\end{tabular}} \\ \hline
Positive ($Y(1) = 1$) & \begin{tabular}[c]{@c@}{False Negative}\\ $l_{01}>0$\end{tabular} & \multicolumn{1}{c|}{\begin{tabular}[c]{@c@}{True positive}\\ $l_{11}=0$\end{tabular}} \\ \hline
\end{tabular}
\end{center}
The True Negative is a student whom the teacher correctly identifies as Always Low, while the True Positive is correctly identified as Always High or Helpable. In all these cases, the recommendation of the teacher does not cause any loss: $l_{00} = l_{11} = 0$.\footnote{To simplify the analysis, we abstract here from the possibility of negative losses (gains), which may differ between H, AL, and AH students.} The False Negative is a Helpable or Always High student who is not recognized as such by the teacher and thus receives a low recommendation. In her/his case, the loss $l_{01}$ is positive because the student could complete the high track following the opposite advice (in the case of H) and complete the high track regardless, but with a track change (in the case of AH). Similarly positive is the loss for the last type of student, False Positive $l_{10}$, who is an Always Low not recognized as such by the teacher. Therefore, she/he receives a high recommendation without being later able to complete the high track.
Let's normalize to $l_{01} = 1$ the loss suffered by a False Negative student. It is then reasonable to assume that $l_{10}\leq l_{01} = 1$.
This is because $l_{10}$ is the “short-term” loss generated by the cost of changing track that a False Positive AL student incurs if she/he is recommended a too-difficult track. In light of this, if $l_{01}$ is the loss of a False Negative AH, then it is also a short-term cost of having to change track, which can be assumed to be comparable for AL and AH students, and therefore the above inequality holds weakly: $l_{10} = l_{01} = 1$. If, instead, $l_{01}$ is the loss of a False Negative H student, he/she suffers a long-term loss extending well beyond the end of secondary school, because she/he earns, in expectation, the lower lifetime income of a low graduate, while with the opposite advice she/he would earn the higher one of a high graduate. Therefore, in the case of this student, we can assume that $l_{10} < l_{01} = 1$ strictly.
We can then write the overall expected loss of a recommendation given by a source $Z=z \in \{0,1\}$ as
\begin{equation}
\mathcal{L}(l_{10},z ) = \Pi_{01}(z) + l_{10}\Pi_{10}(z),
\end{equation}
where $\Pi_{01}(z) = Pr(R(z)=0; Y(1)=1)$ is the probability of a False Negative (H or AH) and $\Pi_{10}(z) = Pr(R(z)=1; Y(1)=0)$ is the probability of a False Positive (AL). Our goal is to compare the overall loss $\mathcal{L}(l_{10}, z )$ when the source is a “teacher alone” $(z=0)$ or a “teacher informed by test scores and allowed to upgrade"” $(z=1)$, computed at different values of the relative weight $l_{10}$. This weight is at most equal to $1$, when the short-run loss of AL students is considered by the policy-maker as relevant as the long-run loss of H students, so that $l_{10}=1$, and decreases towards $0$ as the former loss becomes less relevant $(l_{10}<1)$.
In the Online \nameref{ap:ben-michael} we show that
\begin{equation}
\mathcal{L}(l_{01},1) - \mathcal{L}(l_{01},0) = \Pi_{11}(0) - \Pi_{11}(1) + l_{10}(\Pi_{10}(1) - \Pi_{10}(0))
\end{equation}
where the RHS is identified and can be estimated thanks to Assumption (ref) (Randomization of the source). The left panel of Figure (ref) reports for each track the P-values of the tests of the null that this difference is negative for values of the relative weight $l_{01}$ in the set $\{0, 0.05, 0.1, ...,1\}$:
\begin{equation}
H_0:\mathcal{L}(l_{01},1) - \mathcal{L}(l_{01},0)\leq 0 ;
H_1:\mathcal{L}(l_{01},1) - \mathcal{L}(l_{01},0)> 0
\end{equation}
Failure to reject the null hypothesis of this test (high P-value)
indicates that we cannot rule out the possibility that allowing teachers to upgrade their recommendations based on test scores is a better (or equally effective) system because we do not have enough evidence to conclude that the provision of test scores leads to larger losses.
Conversely, if we can reject $H_0$ (with a small P-value), we can rule out the possibility that allowing teachers to upgrade based on test scores generates smaller losses and thus improves recommendations.
With the exception of three tracks (V:BL, V:GT and A:HAVO), we cannot reject that the difference is negative even when the short-run loss of AL students is considered as relevant as the long-run loss suffered by H students $(l_{10}= 1)$ and AL students have the maximum relative weight.
For the remaining V:BL, V:GT and A:HAVO tracks, we reject that the difference is negative only when the short-term loss suffered by AL students has an (implausibly) high relative weight -- higher than about 0.8, 0.5, and 0.3, respectively, at the 5% significance level. Similar evidence is reported in the right panel of the figure for the three aggregate tracks. We never reject the null at the same significance level for the All and Vocational tracks, even if the relative weight on the short-term loss of AL students is the highest. In the case of the Academic track, we reject only when this weight is higher than about 0.75.
\begin{figure}[ht]
\caption{Comprehensive evaluation that test score information improves teachers' recommendations }
\begin{minipage}{0.95\linewidth \setstretch{0.75}}
{\scriptsize Notes: The figure reports p-values of the test that the difference $ \mathcal{L}(l_{01},1) - \mathcal{L}(l_{01},0)$ between the overall losses with and without test score information (equation (ref)) is negative, at different values of the relative weight $l_{10}$ of AL students. Failure to reject the null hypothesis in this test (high P-value) indicates that we cannot rule out that the loss is larger when teachers do not see the test score information. The solid line indicates the 5% significance level of the P-values above which the null cannot be rejected with sufficient statistical confidence. }
\end{minipage}
\end{figure}
In light of this evidence, we conclude that, overall, the possibility of upgrading based on test scores helps primary teachers improve the quality of their secondary school track recommendations. For this not to be the case, the short-term loss of AL students (who have to change track because they are recommended too difficult a track that in the end they cannot complete) would have to be considered implausibly almost as relevant as the life-time loss of H students (who are prevented from completing a high track they would be able to complete).
\section{An alternative to assuming the Exclusion Restriction
}
In this section, we explore the identification and estimation of the Average Principal Causal Effects without relying on the Exclusion Restriction (ER) assumption. This assumption may fail to hold if the causal effect of $Z$ on $Y$ is not mediated solely by $R$. Consider, for instance, the Never-Takers stratum. Since these students are never upgraded, under the ER their expected value of $Y$ should, in principle, be the same on both sides of the cutoff.
However, scoring above the cutoff may boost the confidence of these students and encourage their parents to advocate for an upward track change after initial enrollment in secondary school. More generally, it may have a direct positive effect on these students' performance in secondary school, even in the absence of an upgrade at the moment of the initial track enrollment. In these cases, the Exclusion Restriction would not hold.
We study such a possibility by deriving bounds on the APCEs when
a specific form of direct effect of $Z$ on $Y$ for given $R$ is allowed to play a role. Finding that even in the presence of this effect, our results are unchanged would suggest that the ER assumption is not strictly necessary for our conclusions about the preferability of informing teachers about the test score results of their students, or, in other words, that at least this deviation from the Exclusion Restriction assumption is not changing our results (see also Angrist1996).
To this end, we first assume that the direct effect of $Z$ on $Y$ is additive and equal for all students:
\begin{assumption} Homogeneity of the direct effect of $Z$ on $Y$ \\
$\mathbb{E}[Y_{i}(1,r)] =\mathbb{E} [Y_{i}(0,r)] + \eta \hspace{0.3 cm} \mbox{for } r \in \{0,1\}$,
\end{assumption}
Under Assumption (ref), we derive new expressions for the APCEs and their bounds as a function of $\eta$, as described in the Online \nameref{ap:robust}.\footnote{Mealli2013 study another example of bounds for principal causal effects when Exclusion Restriction assumptions are relaxed.} These expressions generalize the original framework proposed by Imai2023. Indeed, it can be shown that as $\eta$ goes to 0, i.e. the level at which the Exclusion Restriction holds, the new bounds converge to those derived in Section (ref) under the ER.
More generally, the bounds of the APCEs that we obtain are functions of the unknown parameter $\eta$, which serves as a \emph{sensitivity} parameter (e.g., Imbens2003Sensitivity), and as such cannot be identified and estimated without invoking further assumptions.
In order to gain some insights into the plausible magnitude of $\eta$, we can invoke the Principal Ignorability assumption described below that allows to identify and estimate $\eta$.
This enables us to establish not only the extent to which the ER assumption is violated but also the stability of our results if the violation indeed occurs and identification is achieved under an alternative assumption.
To understand the nature of this Principal Ignorability assumption, consider the effect of $Z$ on $Y$ for Never-Takers, which under the Homogeneity Assumption (ref) is equal to $\eta$:
\begin{equation}
\eta = \mathbb{E}[Y \mid Z=1 , R(1) = 0, R(0) = 0] - \mathbb{E}[Y \mid Z=0 , R(1) = 0, R(0) = 0].
\end{equation}
The first term on the RHS of equation ((ref)) is identified and can be easily estimated because students who score at the cutoff but are not upgraded by the teacher are certainly Never-Takers and are observed. The second term, however, is not identified because the group of students scoring below the cutoff (and not upgraded) comprises both Compliers and Never Takers, who cannot be distinguished. We solve this problem by invoking the following assumption:
\begin{assumption}
Principal Ignorability: \\
$Y_i(0) \perp\!\!\!\!\perp R_i(1) \mid R_i(0) = 0, X_i$
\end{assumption}
This assumption implies that, conditional on covariates, the distribution of potential outcomes below the cutoff is the same for Compliers and Never-Takers.\footnote{This assumption is, of course, as debatable as the ER assumption. Moreover, similarly to the unconfoundedness assumption, its plausibility crucially depends on the information contained in the observed covariates (see feller2017principal, mattei2023assessing, and ding2017principal) However, our goal here is simply to assess how stable our results are when invoking this assumption instead of the ER.} Therefore, the second term on the RHS of equation ((ref)) is identified and $\eta$ can be estimated.\footnote{Using observations above the cutoff, where Never-Takers are identified, we can estimate the probability of being a Never-Taker conditional on covariates (the estimated principal score $\hat p_{nt}$). Then, $\mathbb{E}[Y_i \mid Z_i=0 , R_i(1) = 0, R_i(0) = 0] $ can be estimated by the average $Y_i$ for units with $Z_i=0$ and $R_i=0$, weighted by the principal score. That is, under Assumption (ref), $\mathbb{E}[Y_i \mid Z_i=0 , R_i(1) = 0, R_i(0) = 0] $ is identified and can be estimated by $ \frac{\sum_{i: Z_{i} = 0, R_{i}=0} \hat p_{i,nt}Y_{i}}{\sum_{i: Z_{i} = 0, R_{i}=0}\hat p_{i,nt}}$. See mattei2023assessing for further details.}
Table (ref) in the Online \nameref{ap:robust} reports estimates of $\eta$ obtained with the above procedure.
Some point estimates of $\eta$ are quantitatively sizeable. For example, $\hat \eta$ for students in track V:BL/KL in the 2016 cohort is 0.1013 (s.e.: 0.0465), which is larger than the estimated numerator of the $APCE_H$ for Helpable students in the same track, i.e. 0.0773.
In light of this finding, we use the point estimates $\hat \eta$ to estimate the bounds for the APCEs when the Homogeneity Assumption (ref) and the Principal Ignorability Assumption (ref) hold instead of the ER.
Figure (ref) compares, for the three aggregate tracks and for the H, AH, and AL strata,
the point identified $APCE_J$ estimates obtained under unconfoundedness (Assumption (ref), red dot), that are derived in the Online \nameref{ap:apce_proof}; the bounds of the $APCE_J$ obtained under the ER (Assumption (ref), solid black), that are derived in Section (ref); and the bounds of the $APCE_J$ obtained under PI and Homogeneity (Assumptions (ref) and (ref), dashed light-blue), that are derived in this section.
In the case of H and AL students, unconfoundedness delivers estimates that are located near the lower bounds of the partially identified alternative estimates. For the AH students, both with and without unconfoundedness, the $APCE_J$ estimates are precise and close one to the other.
\begin{figure}[tbp]
\caption{Overview of the $APCE_J$ estimates}
\begin{minipage}{0.9\linewidth \setstretch{0.75}}
{\scriptsize Notes: The figure reports, for the three aggregate tracks and for the H, AH and AL strata: the point identified $APCE_J$ estimates obtained under unconfoundedness (Assumption (ref), red dot), that are derived in the Online \nameref{ap:apce_proof}; the bounds of the $APCE_J$ obtained under the ER (Assumption (ref), solid black), that are derived in Section (ref); and the bounds of the $APCE_J$ obtained under PI and Homogeneity (Assumptions (ref) and (ref), dashed light-blue), that are derived in Section (ref).
}
\end{minipage}
\end{figure}
This graphical comparison under the different assumptions indicates that the implications in terms of \textit{lower bounds} of the relevant effects are very similar in all cases. We conclude that our main results are robust with respect to these alternative identification strategies.\footnote{The estimates of all the bounds obtained without the Exclusion Restriction and under the alternative PI and Homogeneity assumptions can be found in Online \nameref{ap:robust}.}
\section{Fairness of teachers' recommendations}
Within the same statistical framework described in the previous sections, Imai2022 and Imai2023 propose the concept of “Principal Fairness," which can be interestingly used to compare the fairness of two sources of first-year track assignment.
\begin{definition} Principal Fairness:\\
A source $Z_i = z \in \{0,1\}$ of first-year track assignment $R_{i}(z)$ satisfies Principal Fairness with respect to a protected attribute $B_i$ (e.g., SES, race, gender), if the first-year recommendations deriving from this source are conditionally independent of $B_i$ within each principal stratum $J_i \in \{AH,AL,H\}$
\begin{equation}
Pr(R_{i}(z) | J_i, B_i) = Pr(R_{i}(z) | J_i)
\end{equation}
\end{definition}
According to this definition, a recommender is fair as long as her recommendations are independent of the protected attribute among students in the same stratum. The analogous fraction in other strata can be lower or higher if it is equal for both genders within each stratum. A test for Principal Fairness of a source of secondary school track recommendations can then be designed as follows. Let the two sources of recommendations that we would like to compare be: “teacher alone" ($z=0$) versus “teacher informed by test scores and allowed to upgrade" ($z=1$).
Given two values $b$ and $b'$ of a protected attribute $B_i$, the Principal Fairness of source $ z$ in stratum $j$ is given by:
\begin{equation}
\Delta_j(z) = Pr\{R_i(z)=1 | B_i=b, J_i=j\} - Pr\{R_i(z)=1 | B_i=b', J_i=j\},
\end{equation}
where $R_i(z) = R \in \{0,1\}$ is the recommendation given by source $z$ to student $i$. Note that source $z$ is perfectly fair if $ \Delta_j(z)=0$.
Otherwise, source $z$ is {\it unfair} because its probability of recommending the high track to students in stratum $j$ changes with the values of $B_i$.
It is important to note that, considering the protected attribute $B_i$ without conditioning on the other covariates requires to interpret carefully the reason for a possible lack of fairness. A fairness criterion that does not hold unconditionally may hold conditionally if the covariates correlate with the protected attribute (see Imai2022, for a discussion). For instance, suppose that teachers do not discriminate based on immigrant status but do discriminate based on SES, and immigrants have, in general, lower SES. In this case, if the protected attribute $B_i=b$ in equation ((ref)) is immigrant status, this attribute should actually be considered as a proxy for SES. However, from a descriptive viewpoint, it would still be the case that immigrants are treated differently, even if in each stratum defined by SES they are not. As argued by Imai2023, the choice between marginal principal fairness and conditional principal fairness is not statistical.
The sign of $\Delta_j(z)$ is also important, as it indicates the direction of discrimination. Suppose again that $B_i=b$ denotes immigrant students. Then finding, for example, that $\Delta_j(1) > 0$ while $\Delta_j(0) <0$, would mean that when recommenders are allowed to upgrade their advice based on test scores, immigrant students are positively discriminated into the high track, while when the recommenders give advice without seeing test scores, immigrants are negatively discriminated. Note that in our setting, where teachers are officially allowed to upgrade only students who score above the cutoff, all the $\Delta_{j}(0)$ should be equal to 0 by definition. Therefore, a $\Delta_{j}(0) = 0$ cannot be interpreted as evidence of fairness in the provisional track recommendation made by teachers without the test scores. On the other hand, a $\Delta_j(0)$ different from 0 would indicate different rates of non-compliance below the cutoff across protected attributes in stratum $j$.
\begin{figure}[ht]
\caption{Fairness for immigrants}
\begin{minipage}{0.9\linewidth \setstretch{0.75}}
\scriptsize Notes: the figure plots, for the three aggregate tracks and for the H, AH and AL strata, the point estimates of the statistic
$\Delta_{j}(z) = Pr(R(z)=1 \mid \text{Immigrant}) - Pr(R(z) = 1 \mid \text{No Immigrant})$ and the correspondent 95% confidence intervals.
\end{minipage}
\end{figure}
\begin{figure}[ht]
\caption{Fairness by SES}
\begin{minipage}{0.9\linewidth \setstretch{0.75}}
\scriptsize Notes: the figure plots, for the three aggregate tracks and for the H, AH and AL strata, the point estimates of the statistic
$\Delta_{j}(z) = Pr(R(z)=1 \mid \text{Low SES}) - Pr(R(z) = 1 \mid \text{High SES})$ and the correspondent 95% confidence intervals.
\end{minipage}
\end{figure}
Under assumptions (ref) (SUTVA), (ref) (Local Randomization), (ref) (Exclusion Restriction), (ref) (Monotonicity of $Y$ with respect to R), and (ref) (Unconfoundedness of R with respect to Y), as shown by Imai2023, we can estimate equation (ref). Figure (ref) plots, for the three aggregate tracks and for the H, AH, and AL strata, $\Delta_j(1)$ and $\Delta_j(0)$, considering immigrant status as the protected attribute $B_i=b$.
In all these cases, $\Delta_j(0)$ is not statistically different from zero and is small in size. $\Delta_j(1)$ is instead clearly positive and significantly larger than $\Delta_j(0)$ in all strata. This evidence
suggests that the information provided by test scores considerably increases the chances that immigrant students are upgraded.
The interpretation of this finding is that test scores induce a larger revision of the teachers' posterior beliefs regarding the ability of immigrant students compared to those of native students. Note that this outcome is good news for immigrant students in the H and AH strata but not for immigrant students in the AL stratum, for whom this positive discrimination causes track changes.
Figure (ref) reports the analogous quantities considering low socioeconomic status (SES) as the protected attribute, where low SES is defined as having a household income below the median. For the academic tracks, none of the $\Delta_{j}(z)$ is significantly different from 0.
In vocational tracks, however, the $\Delta_{j}(0)$ are negative and the $\Delta_{j}(1)$ are positive and economically significant, albeit smaller in size than in the case of immigrant students (the scale of the vertical axis changes between Figures (ref) and (ref)). The fact that all the $\Delta_{j}(0)$ are negative suggests a higher rate of non-compliance below the cutoff among high-SES students. A plausible explanation is that parents of high-SES students may be more likely to pressure teachers to upgrade their children when they are provisionally recommended for vocational tracks. On the other hand, the positive $\Delta_{j}(1)$ values indicate that the provision of test scores induces upgrades for low-SES students more often than for high-SES ones.
We view this finding as a consequence of test scores inducing a larger shift in teachers’ posterior beliefs for low-SES students.
Finally, the analogous evidence for gender is shown in Figure (ref) of the Online \nameref{ap:fair}.
We do not find any $\Delta_{j}(z)$ significantly different from 0, meaning that teacher recommendations are fair when the protected attribute is gender.
\section{Conclusions}
Using a quasi-experimental setting offered by Dutch educational institutions, we show that when track recommendations given to students at the end of primary education can be upgraded based on standardized test scores, the quality of advice improves by at least 6% as measured by the fraction of students that are recommended a more challenging track they can complete successfully, while in the counterfactual case, they would have remained in a lower level track. An improvement of about the same size is also estimated for students who should be recommended to more challenging tracks, as they are always able to complete them independently of where they start.
We also find, however, that the possibility of upgrading based on test scores results in a significant number of students who are indeed upgraded but are unable to complete a more challenging track. In the case of these students, the goal should be to direct them immediately toward the low track that they can complete successfully, thus minimizing costly track changes. However, an overall evaluation that combines these opposite findings in a weighted objective function indicates that the relative weight of the short-term losses suffered by students who must change track because of misplacement when teachers are given the opportunity to upgrade based on test scores would have to be unreasonably high to conclude that it is better if teachers do not see these scores.
The implications for policy of these results are particularly relevant in countries where tracking is the established way to organize secondary school studies. It is surprising that no experimental evidence exists to guide educational policymakers of these countries in designing the best system of track advising. We fill in this gap in several ways.
First, we define a reasonable metric to evaluate the quality of teachers' recommendations. This metric values two objectives: reducing costly track changes and
directing towards more challenging tracks students who are able to complete them, but who need to be “pushed" towards these tracks and be convinced that they can complete them successfully. Differently from the case of a weather forecast, predicting the school track in which a student would have the best performance and recommending this track may have an effect on the object of the prediction itself, i.e., the choice of the student and her performance in the chosen track, as well as later in life. In other words, track recommendations may be self-fulfilling prophecies, and it is crucial to consider this feature in the optimal design of their implementation. The metric we propose puts this feature at the center stage.
Second, our results show that allowing teachers to upgrade their recommendations based on standardized test scores helps them to improve the quality of track advising.
Ideally, one would like to compare recommendations given by teachers with and without the information provided by test scores, which is not what the quasi-experiment we study can offer.
However, our evidence suggests that test scores do help teachers provide better recommendations in general, not just to upgrade previous ones. Therefore, appropriate controlled experiments should be implemented to further study and formally test the more general hypothesis.
Third, a large literature (see footnote (ref)) has demonstrated, in different contexts, that teachers’ recommendations do not typically reflect only the previous academic achievement and the future ability potential of a student, but are also highly correlated with characteristics like SES, gender, race, and behavior in class. The bias affecting recommendations that are determined by the conscious or unconscious use of these variables by teachers is typically deemed unacceptable and has reinforced the opposition to school tracking in general. Our evidence suggests that test score information also has an impact on the fairness of recommendations with respect to protected attributes. We find that the provision of test scores induces more frequent upgrades of recommendations for immigrant and low SES students. We interpret these results as indicating that the provision of test scores information induces a larger shift in teachers' posterior beliefs regarding these students.
Together with the fact that high-SES students are more likely to be upgraded when originally recommended a vocational track and scoring below the cutoff,
we conclude that test scores are a valuable tool for increasing the fairness of track recommendations.
Independent of track-advising, but no less importantly, our results contribute to the recent and growing literature on algorithm-assisted human decisions (see, for example, Imai2022, Imai2023, Imai_2024, and Rambacan2024). It is increasingly common for decision-makers in various fields (justice, medicine, finance, politics, and education, to name a few) to use big data processed by multiple algorithms to make high-stakes decisions in highly uncertain contexts. However, it is still unclear how to make the best use of these algorithmic and data-driven tools in human decision-making. Our results contribute to this literature by providing novel quasi-experimental evidence and by improving existing methods to assess algorithm-assisted human decisions.
spacing{0.8}
{
}
\setcounter{page}{1}
center[center omitted — 320 chars of source]
spacing{1}