Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
71,819 characters · 19 sections · 56 citation commands
Teacher bias or measurement error?
\begingroup \footnote{ We thank S{\"o}nke Matthewes, Joppe de Ree, Astrid Sands{\o}r, and Dinand Webbink for valuable comments. The paper has also benefited from the comments of participants at the EQOP meeting at the University of Oslo, the workshop on exam performance and test-taking behavior at University Carlos III of Madrid, KVS new paper sessions, LMG Lisbon, the EALE, LEER, LESE, and PPE conference, and the brown-bag seminars at the University of the Balearic Islands, the University of Bergen, and the Utrecht University School of Economics. The data used in this study are housed within the secure virtual environment at Statistics Netherlands and cannot be publicly shared due to privacy restrictions. However, all the underlying data supporting the presented results are securely accessible through Statistics Netherlands. National and international researchers can request access to the administrative data through Statistics Netherlands, and to the test score data through the Netherlands Cohort Study on Education (NCO) by contacting [email removed]. Matthijs Oosterveen gratefully acknowledges financial support from FCT, Fundação para a Ciência e a Tecnologia (Portugal), national funding through research grant UIDB/04521/2020 & UID06522. The authors have no relevant or material financial interests that relate to the research described in this paper. All omissions and errors are our own. Author names are ordered alphabetically. } \addtocounter{footnote}{-1} \endgroup }
\thispagestyle{empty}
\linespread{1.50}
\setcounter{page}{1}
Teacher bias may be central to inequalities in educational attainment and later life outcomes. To estimate the extent of such bias, researchers typically compare student assessments in which teachers exercise discretion, such as teacher-assigned grades or track recommendations, to more objective indicators of student ability, such as standardized test scores alesina2024revealing,falk2023mentoring. However, the presence of systematic gaps in subjective teacher evaluations between groups, conditional on objective ability measures, is not necessarily evidence of teacher bias. This approach faces two main concerns: omitted variables and measurement error in the ability measure. Studies often address the first concern by controlling for additional factors burgess2013test. In contrast, the second concern is typically overlooked or addressed under strict assumptions that may not hold in practice, a shortcoming shared by the broader literature on discrimination.
This paper focuses on the second concern: measurement error in the objective measure of ability. In classical testing theory, test scores are typically based on the number of correct answers and provide an unbiased but noisy estimate of ability. The resulting “classical” measurement error in the test score is considered purely random and hence uncorrelated with ability. In modern item response theory (IRT), statistical models produce test scores that may be biased estimates of ability but are unbiased as predictors. The resulting “non-classical” measurement error may correlate with ability. For instance, in case the test score is estimated with a posterior mean score, students with a low and high signal are drawn towards the population mean, such that the error correlates negatively with ability, yet is uncorrelated with the test score.
To examine the implications, consider a regression in which a teacher’s subjective evaluation of a student is the dependent variable, and the independent variables include a group characteristic, such as family socioeconomic status (SES), and a test score as a measure of student ability. In case of classical measurement error, the test score coefficient is biased towards zero (the well-known attenuation bias). Moreover, if student ability is associated with SES and matters for teacher evaluations, this bias also contaminates the coefficient of the conditional SES gap. In particular, when ability is positively associated with both SES and teacher evaluations, the conditional SES gap is overestimated. In case of non-classical measurement error, the attenuation bias in the test score coefficient may be partially or fully eliminated, though this leads to a biased coefficient of the association between ability and SES. This bias then also contaminates the estimate for the conditional SES gap.
We first show how both classical and certain forms of non-classical measurement error, particularly relevant to modern test scores, generate identical contamination bias in the coefficient measuring the conditional gap between groups. We then discuss three empirical strategies to address measurement error. First, we consider an instrumental variable (IV) approach, a well-established method in econometrics to correct for measurement error hausman2001mismeasured. Second, we discuss a method in the spirit of errors-in-variables (EIV) models, an approach that is common in psychometrics fuller2009measurement. The additional parameter required for the EIV strategy, compared to OLS, is the reliability ratio of the test that is used to measure ability. We argue that, under weaker assumptions than the IV strategy, the IV first stage can be used to identify this ratio. Third, we combine the EIV strategy with a novel approach to identify the reliability ratio under yet weaker assumptions.
We apply these three strategies to estimate conditional gaps in teacher track recommendations in the Netherlands, focusing on differences by family SES. The Dutch tracking regime provides a compelling setting for this analysis. While nearly all OECD countries have some form of educational tracking, the Netherlands assigns students to secondary-school tracks at a relatively young age of twelve. Assignment is based on binding track recommendations from the primary school teacher wvo2014. These recommendations may be highly consequential, as roughly 70% of students does not experience track mobility four years into secondary education deree2023quality. Track placement determines the level of the curriculum and peer group abilities, and may affect student achievement matthewes2021better. As only the highest tracks provide access to university, track placement may also have long-term consequences for attainment and income dustmann2017long,borghans2019long. Moreover, negatively biased recommendations signal low teacher expectations, which can in turn harm student outcomes carlana2019implicit,papageorge2020teacher.
Given the consequential nature of these recommendations, potential teacher bias has been a major policy concern in the Netherlands. This concern gained momentum in 2016, when a report by the Dutch inspectorate2016 demonstrated that students from low SES families receive significantly lower track recommendations, conditional on scores from a standardized end-of-primary school test. This finding fueled an ongoing public debate on teachers as gatekeepers of opportunities and contributed to a national policy reform, effective from 2023-24. The reform reduces the role of teachers in track placement, and increases the role of the end-of-primary school test. Similar findings in academic studies reinforce concerns about unequal opportunities in education, both in the Netherlands zumbuehl2025can and elsewhere carlana2022implicit,falk2023mentoring.
Using new administrative data, we first replicate previous findings from the Netherlands, and find that children with lower family SES receive systematically lower track recommendations, conditional on standardized test scores. We discuss that our test scores are derived from an IRT model, with ability parameters estimated using weighted maximum likelihood sanders1993psychometrie. This method is known to produce largely unbiased estimates of student ability warm1989weighted, and the measurement error is considered classical. We address concerns about measurement error using three strategies. Our novel strategy combines the EIV method with an unbiased estimate for the test's reliability ratio under weak assumptions, based on students’ standardized test scores across all primary school grades. Across the three approaches, we find that measurement error can explain between 35 and 43% of the conditional SES gap.
Our paper contributes to the broad literature on measurement error schennach2016recent. In particular, we focus on contamination bias due to measurement error. As modalsli2022spillover point out, “[a]lthough the notion of bias in one coefficient arising from error in another regressor is a well-known econometric result, it is seldom addressed in practice with empirical studies.” modalsli2022spillover examine the implications of classical measurement error in multigenerational income regressions and use an IV approach to mitigate the resulting bias. Similarly, gillen2019experimenting use IV to show that the gender gap in competitiveness is well explained by risk attitudes and overconfidence once classical measurement error is addressed. Our study extends this literature by applying three strategies to address contamination bias and by incorporating both classical and non-classical measurement error. The latter is pervasive in education, where test scores are often derived from IRT models jacob2016measurement. It is also common in the labor market, as many individuals misreport hours worked, leading to biased estimates of wage discrimination and inequality borjas2024. More generally, it arises in survey data when respondents base their answers on incomplete information hyslop2001bias
We also contribute to the literature on teacher bias by providing new evidence on the conditional SES gap in track recommendations. Teacher bias in track recommendations has been studied in Germany falk2023mentoring, Italy carlana2022implicit, and the Netherlands zumbuehl2025can. Previous studies have also examined teacher bias in grading by gender cornwell2013noncognitive,lavy2018origins,terrier2020boys, ethnicity burgess2013test, and migration background alesina2024revealing. Consistent with the observation above, contamination bias due to measurement error has been largely ignored in this literature. Three notable exceptions are botelho2015racial, ferman2022assessing, and zhu2024, who study teacher bias in grading in Brazil and the US. All three studies find that the conditional gaps are substantially reduced or even reversed when addressing measurement error in test scores. However, as these studies assume classical measurement error and exclusively apply an IV strategy, it remains unclear whether the results hold under non-classical error and are robust to alternative strategies relying on weaker assumptions.
In the Netherlands, the transition to tracked secondary education takes place around the age of twelve when students leave the sixth and final grade of primary education. The secondary school system consists of three main tracks: (i) a university track (vwo), which has a duration of six years and gives students access to university education, (ii) a college track (havo), which has a duration of five years and prepares students for college education, and (iii) a vocational track (vmbo), which has a duration of four years and serves as a preparation for vocational education. The vocational track consists of four sub-tracks, which can be ordered from more practice-oriented to more theory-oriented.
The allocation of students to tracks is based on the recommendation provided by the sixth-grade primary school teacher. The recommendation process consists of two main steps. First, by March of the sixth grade the teacher provides an initial track recommendation. Then, in April or May, each student takes a standardized end-of-primary education test, which the teacher may use to consider an upward revision wpo2014. Hence, the final track recommendation may be higher (but not lower) than the initial one. The teacher track recommendation is binding in the sense that, as a rule, secondary schools cannot place students in a higher track than indicated by the recommendation wvo2014.
We use the initial, instead of the final, track recommendation as our measure for the teacher's subjective evaluation. This choice reflects that the initial recommendation is often the most influential, as many secondary schools begin allocating students to first-year classes as early as March, relying on these initial recommendations. Secondary schools may face challenges when changing their allocation by the time the final track recommendation is provided in May inspectorate2019. Moreover, in practice, teachers upgrade only 1 in 10 students, meaning the initial and final track recommendations are identical for 9 out of 10 students deree2023quality.
Teachers are expected to make holistic track recommendations by considering both objective standardized test results and their subjective judgments on factors such as motivation, attitude towards school, and classroom behavior. While there are no strict national rules describing how to combine these types of information, general guidelines from the ministry2022 suggest that standardized test results provide a “good starting point”. A recent survey among 400 primary school teachers shows that 85% of the respondents attach “considerable” to “a lot of” value to standardized test results when formulating their recommendations Duo2023.
All Dutch primary school students, from grade one to six, are expected to take two standardized tests per year. The first test is administered in the middle of the school year around February and the second at the end of the school year in June. We will refer to these tests as the midterm and end term test, respectively. We use the fifth-grade end term test as the main measure of student ability to estimate the conditional SES gap in track recommendations. This test can be considered the last score that teachers observe before making the initial recommendation. Alternatively, one could use the end-of-primary school test score (the sixth-grade end term test) as the ability measure and the final track recommendation as the teacher's subjective evaluation. This alternative approach is problematic since the final track recommendation can only be adjusted upwards based on performance on the end-of-primary school test. Hence, this test is low (high) stakes for students who are (un)satisfied with the initial recommendation.
Primary schools have to select one of the standardized test systems that are approved by the Ministry of Education. Most schools opt for the test system provided by Cito, which has a market share of roughly 80% in the period we study. The biannual Cito tests focus on two domains: mathematics and reading. The tests consist of approximately 70 to 100 items, include both multiple-choice and open-ended questions, and are scored by the teacher using standardized answer keys. The item scores are entered into the Cito software and used to generate students' test scores through an item response theory (IRT) model.
An IRT model specifies the probability that a student answers a test item correctly as a function of latent student ability and item parameters. The IRT model used by Cito is the Birnbaum model Cito2017, which includes two item parameters: difficulty and discrimination. The item difficulty parameter reflects the student ability level at which there is a 50% probability of answering the item correctly. The discrimination parameter reflects how sharply the item distinguishes between students of different ability levels. These parameters may be estimated for a selection of items to form a so-called calibrated item bank, which can be used to create multiple tests of varying difficulty that aim to measure student ability on a common scale. This enables, for instance, observing the development of student ability from grade 1 to 6, one of the main goals of standardized testing in the Netherlands.
Cito constructs its calibrated item bank through national norming studies. First, sample tests are administered to representative student populations. Second, the item parameters are estimated: the discrimination parameters are held fixed while the difficulty parameters are estimated using conditional maximum likelihood. Both parameters are then iteratively adjusted to improve model fit. The resulting calibrated item bank enables the construction of the biannual tests for various grade levels. Although more difficult items are used in tests for higher grades, student ability can supposedly be measured on a common scale.
Latent student ability is estimated from these biannual tests as a “fixed effect” using weighted maximum likelihood sanders1993psychometrie. In this approach, the two item parameters are held fixed at their calibrated values, and the likelihood function is weighted by the test information function. This function indicates how much information the test provides about student ability, with items of difficulty close to the student’s ability offering more information. warm1989weighted shows that this procedure yields (nearly) unbiased estimates of student ability. The estimated ability scores obtained through this procedure are the test scores we use in our analysis. Unbiased, yet noisy, estimates of ability implies that the measurement error in the biannual Cito test scores is assumed to be classical.
We use proprietary administrative data from Statistics Netherlands. The data on the biannual standardized tests is gathered by the Netherlands Cohort Study on Education (NCO) and subsequently made available to Statistics Netherlands. The NCO collected the test score data only from schools using the tests developped by Cito. Since 2018–19, data were obtained by individually contacting primary schools, each of which had to grant permission. As a result, the test score data covers about 50% of all primary schools in the Netherlands. See haelermans2020using for details on the NCO data collection. This test score data can be linked to standard administrative data using an anonymized personal ID, allowing us to also observe teacher track recommendations, parental income and education, and several other student background characteristics.
The test score history, from grade one to six, is available for the students who were enrolled in sixth grade from 2018-19 onward. Our sample includes four cohorts: students who are in sixth grade in the school years 2018-19, 2019-20, 2020-21, and 2021-22. The standard administrative data covers 686,309 students across these four cohorts, with almost no missing values. The biannual test score data includes 318,630 students, although not all students took both standardized tests in each grade. After merging these two datasets, we observe 276,841 students in 5,802 schools. After dropping students with one or both grade 5 test scores missing, as well as those with extreme household parental income values (discussed below), our main estimation sample consists of 148,019 students in 3,295 schools.
The first four columns in (ref) present the descriptive statistics for our estimation sample. Panel A focuses on the measures used for the subjective teacher evaluation. Our primary measure is a dummy variable that equals one if the initial track recommendation is at least equal to the college track. College or university track enrollment is highly policy relevant, since it implies the student is tracked towards higher education. Nearly half (49.7%) of the students receive a track recommendation equal to, or higher than, the college track.
Panel B presents the statistics for the different measures of SES. Our baseline SES measure is gross annual household parental income, which averages \euro64,390. We dropped observations with a yearly parental income above (below) the 99.9th (0.1st) percentile. In our analysis below we standardize parental income per cohort. For robustness checks, we use three alternative SES measures: parental years of schooling of the highest educated parent, a binary income indicator denoting whether parental income exceeds the student's cohort median, and a binary education indicator reflecting whether the highest educated parent holds at least a college degree.
Panel C shows the statistics for the biannual Cito tests in fifth grade, per domain and the average across domains. The domain-specific scores are the raw scores estimated on a common scale by the IRT model. The average score is calculated by first standardizing the domain-specific scores per cohort, and subsequently taking the mean across domains. (ref) in Appendix D provides the descriptive statistics for the raw domain-specific scores on the biannual tests in grades two through four. We exclude grade 1 data since a relatively high fraction of the test scores for that grade are missing. Panel D presents the statistics for the variables used as controls in the robustness checks.
We also compare the statistics of the students in our estimation sample with those from three larger samples representing the different steps of our selection process. Columns (5) to (7) correspond to the samples from the standard administrative data covering the full population, the test score data, and the merged dataset, respectively. The statistics in these columns are based on different numbers of observations due to missing values in some variables, particularly in the grade 5 test scores of columns (6) and (7). The comparison indicates that our estimation sample closely resembles the student population in the Netherlands.
Let $y_t$ be a variable that measures a teacher's subjective evaluation of a student in time period $t$, $SES$ a variable that measures the socioeconomic status of the student, and $s_t$ a variable that reflects student ability in time period $t$. The long regression equation equals,
Throughout the paper we suppress the subscript for student $i$ and the intercept in regression equations. Scholars and policy makers are generally interested in the parameter $\beta$. That is, they are interested in whether subjective teacher evaluations differ systematically by SES, or another group characteristic such as gender or ethnicity, conditional on student ability. If individuals with similar student ability experience similar causal effects of the subjective evaluation on outcomes of interest, then $\beta$ can also be interpreted as a measure of principal fairness, as recently introduced by imai2023principal.
We cannot estimate (ref) as we do not observe ability. Instead, students produce an unbiased but noisy signal of ability by taking a test, denoted by $\widetilde{s}_{t}$. The signal equals ability plus measurement error, $\widetilde{s_t}=s_t+m_t$, where $m_t$ is the signal's measurement error. An unbiased signal, meaning one that equals ability in expectation, implies that the error is classical.
This classical error reflects, for instance, that students may guess answers, are more or less familiar with the specific selection of test items, or feel (un)well on the day of the test. According to classical testing theory, measures such as the fraction of correctly answered items follow (ref) and are unbiased but noisy estimates of ability.
In modern assessment systems, test scores are rarely that simple. Instead, they are often the product of statistical IRT models designed to extract more information from the pattern of responses. For instance, jacob2016measurement discuss that IRT test scores in many assessment systems in the US, such as the ECLS, the NELS:88, the ELS, and the HSLS are constructed from posterior means. Let $s^m_t$ denote the test score, which is then computed by
Students are treated as a random sample from a population of ability values $s_t$, characterized by the posterior density function $f(s_t \mid \widetilde{s}_t)$ from the IRT model. The posterior mean score is the mean of that population. Hence, instead of using the signal as the test score directly, the signal is used as a predictor for ability to generate a test score. This prediction is not perfect, so we can similarly write that the test score equals ability plus the score's measurement error, $s^m_{t}=s_{t}+{m}^{m}_{t}$.
Posterior mean scores can be interpreted as Empirical Bayes estimates that, roughly, shrink the student’s (unbiased) maximum likelihood score toward the population mean in proportion to the noisiness of the maximum likelihood score jacob2016measurement. Building on this interpretation, we approximate the posterior mean in (ref) by the kelley1947fundamentals estimator in (ref), where $\widetilde{\lambda}$ is the shrinkage factor. With $\widetilde{\lambda} \in [0,1)$, students with low (high) ability signal are pushed up (down) towards the population mean. As a result, the measurement error of the test score $m_t^m$ correlates negatively with ability. This feature of the non-classical error in (ref) is similar to the error generated by any posterior mean in (ref).
We will analyze contamination bias in $\beta$ when replacing ability in (ref) by test scores measured through (ref). Our results hold (approximately) for two types of test scores. The first are test scores estimated classically or through an IRT model using (weighted) maximum likelihood under a “fixed effect” approach. Both provide an unbiased estimate of ability. This corresponds to (ref) with $\widetilde{\lambda}=1$: the test score is equal to the signal, $s^m_t=\widetilde{s}_t$, and measurement error is classical, $m^{m}_t=m_t$. The second are test scores estimated from an IRT model under a “random-effects” approach, where the scores are posterior means. Equation (ref) holds exactly only for the Kelley estimator, which corresponds to normality for both the prior and likelihood distributions. Generally, it is an approximation, and depending on the choice of these two distributions, the posterior mean in (ref) may even lack a closed form. Because the Kelley estimator corresponds to an OLS regression of $s_t$ on $\widetilde{s}_t$, by the properties of OLS it provides the best linear approximation to the (possibly nonlinear) posterior mean.
Our results speak less to settings where test scores are reported as posterior medians or modes, where multiple plausible values are drawn from a student’s posterior distribution, or where the prior distribution on ability incorporates student characteristics. The latter approach is used in the major US and international assessments NAEP and PISA. An alternative strategy is to jointly estimate the research model in (ref) alongside the IRT model, treating the latter as a direct model of measurement junker2012use. This requires access to item-level response data, which is typically unavailable. IRT models often produce noisier test scores at the tails of the distribution, reflecting the limited information available to these parametric methods when student ability does not align well with item difficulty. Our approach allows for the potential difference between so-called local and global reliability.
Non-classical measurement generated by (ref) with $\widetilde{\lambda} \in [0,1)$ is also considered by hyslop2001bias in survey data. They consider a setting in which respondents are asked questions like “what is the value of $s_t$?” given an information set that consists of the unbiased but noisy signal $\widetilde{s}_t$. Respondents may passively report the “flawed” but unbiased value ($\widetilde{s}_t$), or may actively seek to provide an “optimal” response ($s^m_t$) given their information set. An argument for the latter is that respondents are likely aware of the lack of precision in the signal. Our results on the contamination bias below may thus also be relevant for the analysis of survey data more generally.
The medium regression equation replaces $s_t$ with $s^m_t$,
To analyze what $\beta^m$ and $\gamma^m$ identify, we also introduce the balancing regression equation,
Note that $(1-\widetilde{\lambda} )\mathbb{E}[\widetilde{s}_t]$ is fixed and gets soaked up by the intercept of the regression. The following proposition formalizes what OLS identifies.
All proofs are provided in Appendix A to C. First consider (ref) and (ref) under classical measurement error, where $\widetilde{\lambda}=1$. Classical error on the left-hand side does not introduce bias in the estimate of $\delta^m$. However, as shown in (ref), the estimate of $\gamma^m$ is subject to the well-known attenuation bias. The multivariate reliability (or signal-to-total variance) ratio of the test, denoted by $\lambda$, decreases from one toward zero as the measurement error variance increases, leading to attenuation bias in $\gamma^m$. Now consider the case of non-classical measurement error, where $\widetilde{\lambda} \in [0,1)$. (ref) shows that the attenuation bias in $\gamma^m$ may be smaller and disappears completely when the shrinkage factor equals the reliability ratio, $\widetilde{\lambda} = \lambda$. However, (ref) shows that non-classical error on the left-hand side results in attenuation of $\delta^m$.
(ref) clarifies that both types of measurement error contaminate the estimate of $\beta^m$. Under classical measurement error, attenuation of $\gamma^m$ spills over to $\beta^m$. Under non-classical measurement error, the attenuation bias in $\gamma^m$ may be corrected, but at the expense of attenuation in $\delta^m$, which again spills over to $\beta^m$. Importantly, the contamination is identical for both types of measurement error, as the bias term involves the product of $\gamma^m$ and $\delta^m$, where the former is divided and the latter is multiplied by $\widetilde{\lambda}$. Although this may seem intuitive, to our knowledge this result has not been reported in the literature. In particular, hyslop2001bias analyze (ref) and (ref) under both classical and non-classical measurement error, while pei2019poorly,gillen2019experimenting,ferman2022assessing examine (ref) in the case of classical measurement error only. Note that the contamination bias grows larger as $\gamma$ and $\delta$ increase and $\lambda$ decreases.
The address measurement error, the instrumental variables (IV) strategy aims to use a second objective measure of ability in period $t$ as an instrument for the first measure. In practice, a lagged test score from period $t-1$ is often used instead botelho2015racial,ferman2022assessing,zhu2024. Similar to above, students produce an unbiased but noisy signal of ability by taking a test in $t-1$, where $\widetilde{s}_{t-1} = s_{t-1}+m_{t-1}$. The corresponding test score satisfies (ref), where $s_{t-1}^m=\mathbb{E}[s_{t-1}|\widetilde{s}_{t-1}]$ and $s^m_{t-1}=s_{t-1}+m^m_{t-1}$. We do not index the shrinking parameter $\widetilde{\lambda}$ by time, thereby assuming it remains constant across tests. Our main results hold without this simplification, though at the expense of additional notational clutter and reduced intuition. Using $s_{t-1}^m$ as an instrument for $s_{t}^m$ generates the following first and second stage regression equations respectively,
To show how this IV strategy can address measurement error, it is useful to also introduce the balancing regression for $t-1$,
and the variable $\Delta_{s_{t}}=s_t-s_{t-1}$, which captures the change in unobserved ability between time period $t$ and $t-1$. We describe two additional assumptions that place increasing restrictions on $\Delta_{s_{t}}$.
(ref)(a) requires the change in ability to be the same, on average, across low and high SES students. (ref)(b) requires this for low and high ability students in time $t$.
(ref) requires that the change in ability is the same for each student, which is stronger than (ref). The following proposition summarizes what the IV strategy identifies while combining (ref) with either (ref) or (ref).
Under classical measurement error with $\widetilde{\lambda}=1$, (ref) formalizes that the IV strategy yields consistent estimates under (ref). This result is not new. For instance, gillen2019experimenting use IV to address contamination bias from classical error in (ref) and review similar applications in the literature. Intuitively, a constant $\Delta{s_{t}}$ implies that the test score in period $t-1$ is truly a second objective measure of ability in period $t$. The first stage regression of $s_{t}^m$ upon $s_{t-1}^m$ measures to which extent this relationship is diluted due to measurement error, so that it reveals the reliability ratio $\lambda$. Furthermore, the reduced form equation is similar to the medium regression equation (ref) while replacing $s_{t}^m$ with $s_{t-1}^m$. With a constant $\Delta{s_{t}}$ the reduced form coefficient reveals the estimate $\gamma^m=\gamma\lambda$, which was attenuated towards zero by the reliability ratio. Dividing the reduced form by the first stage corrects for this attenuation. Since $\gamma^{iv}=\gamma$ and $\delta^m=\delta$, we also have that $\beta^{iv}=\beta$.
In the presence of non-classical error with $\widetilde{\lambda}\in[0,1)$, the IV strategy also addresses contamination bias under (ref), though for a different reason. To illustrate, suppose the shrinkage factor equals the reliability ratio, such that $\widetilde{\lambda}=\lambda$. In this case, the OLS coefficient on test scores $\gamma^m$ does not suffer from attenuation bias. However, the IV strategy still divides by the first stage, which means that the IV coefficient $\gamma^{iv}$ is overestimated by a factor equal to the inverse of the reliability ratio. As noted by hyslop2001bias, IV has a bias away from zero in the presence of non-classical measurement error. However, this upward bias in $\gamma^{iv}$ is exactly what is needed to correct for the contamination bias to $\beta^{iv}$, which arises from attenuation in the estimate of $\delta^m$ from the balancing regression. Since the contamination bias term involves the product of the test score and balancing coefficients, the upward bias in the former offsets the downward bias in the latter. As a result, the IV strategy also effectively addresses the contamination bias in the presence of non-classical measurement error.
Based on the results of (ref), the errors-in-variables (EIV) strategy directly formulates expressions for the parameters of interest,
All the elements on the right-hand side are known under (ref), except for the reliability ratio. The EIV strategy then plugs a reliable estimate for $\lambda$ into (ref) and (ref) to estimate the parameters of interest.
(ref) shows that the first stage can be used as an estimate for $\lambda$ under (ref). The intuition for this result follows more easily in a univariate setting without $SES$ as an independent variable. The univariate reliability ratio is defined by replacing $\sigma_{u_{t}}^2$ in $\lambda$ with $\sigma_{s_{t}}^2$. (ref)(b) implies that $\mathrm{Cov}[s_t,s_{t-1}]=\sigma_{s_{t}}^2$, so the coefficient from the univariate first stage regression of $s^m_{t}$ on $s^m_{t-1}$ is proportional to the univariate reliability ratio. This result extends to the multivariate setting since together (ref)(a)-(b) imply that $\mathrm{Cov}[u_t,u_{t-1}]=\sigma_{u_{t}}^2$.
Our EIV first stage (EIV FS) strategy uses the first stage as an estimate for $\lambda$. The following proposition formalizes what EIV FS identifies under (ref) and (ref).
Similar to the IV strategy, EIV FS corrects attenuation bias in the test score coefficient under classical measurement error, overestimates it under non-classical error, and addresses contamination bias in the SES coefficient in both cases. Unlike the IV strategy, EIV FS achieves these results under the weaker (ref). It relies directly on the medium regression estimate, whereas the IV strategy uses the reduced form estimate. The two are equivalent only if $\mathrm{Cov}[e_{t},\Delta_{s_t}]=0$, which holds when $\Delta_{s_t}$ is constant. Whereas (ref) only seems likely when the two tests are administered arbitrarily shortly after one another, (ref) appears generally more plausible. In particular, under (ref) the teachers may pay attention to $\Delta_{s_t}$ in their subjective evaluations, so that it is part of the error term $e_{t}$, but $\Delta_{s_t}$ cannot correlate with the included variable $SES$ and unobserved ability $s_t$.
We propose a novel EIV test score history (EIV TH) strategy that recovers the reliability ratio under (ref) only, while making use of data on students' test score histories. The intuition of this strategy is to approximate the ideal conditions of the test-retest method, which estimates the reliability ratio using a test administered at time $t-\epsilon$, arbitrarily short before time $t$ Cito2017. With this ideal test from period $t-\epsilon$, the first stage equation is,
where we include a time subscript on the coefficients. For a test administered arbitrarily close to time $t$, (ref) holds trivially, and (ref) shows that $\pi_{t-\epsilon}'= \lambda$. While a test score that close to time $t$ may not be available, we may observe multiple test scores further away. It is plausible that such scores satisfy (ref), even if (ref) and (ref) may not hold. Nevertheless, the EIV TH strategy leverages these test scores observed at earlier points in time to identify the reliability ratio.
We begin by repeating the first stage equation, sequentially replacing the test score from period $t-1$ on the right-hand side with scores from earlier periods ,
It follows from the proof of (ref) that under (ref) only, $\pi_{t-\tau}'$ is a biased estimate for the reliability ratio,
with $B_{\pi_{t-\tau}}=\Bigg( \frac{\mathrm{Cov}[u_{t},u_{t-\tau}]}{\sigma_{u_{t}}^2}\Bigg)$. Therefore, we use $\pi_{t-\tau}'$ as a dependent variable in a forecasting regression equation of the form,
Subsequently, we make a one-period out-of-sample forecast towards $\tau=\epsilon$, denoted by $ \widehat{f}(t-\epsilon)$. Without forecasting error this resembles the ideal conditions of the test-retest method, such that $\widehat{f}(t-\epsilon)={f}(t-\epsilon)= \pi_{t-\epsilon}'=\lambda$. In the absence of forecasting error, the chosen polynomial in (ref) accurately captures the changes in $\pi'_{t-\tau}$ over time. While the EIV FS strategy estimates $\gamma^{eiv}$ and $\beta^{eiv}$ by replacing $\lambda$ with $\pi'$ in (ref) and (ref), EIV TH estimates them by replacing $\lambda$ with ${f}(t-\epsilon)$.
The baseline OLS estimates are presented in columns (1) and (2) of (ref). Column (1) reports estimates from a regression where the dependent variable is a dummy equal to one if the teacher’s track recommendation is at or above the college track, and the independent variables are parental income (SES) and the fifth-grade end term test score. Consistent with previous evidence from the Netherlands, the findings show that students with higher income parents are more likely to receive at least a college track recommendation, conditional on the fifth-grade end term test. The estimate for the conditional SES gap implies that a one standard deviation increase in parental income is associated with a 2.8 percentage point increase in receiving at least a college track recommendation.
Column (1) further shows that test scores are positively associated with teacher track recommendations. As our test scores are subject to classical measurement error, this estimate is attenuated. Whether and how this contaminates the OLS estimate for the conditional SES gap depends upon the relationship between test scores and SES. This relationship is presented in column (2). As expected, SES is positively associated with the fifth-grade end term test score. We conclude that OLS overestimates the conditional SES gap, where the magnitude of this positive contamination bias also depends on the reliability ratio of the test.
(ref) in Appendix D shows OLS estimates of the conditional SES gap while including additional control variables: gender, migration dummies, age, and cohort fixed effects. The estimated gap only changes marginally, from 0.028 to 0.029. It may seem unsurprising that the estimate changes little, given that most included controls correlate weakly with parental income, with migration background as a clear exception. In general, these results suggest that our main estimates are robust to controlling for additional variables typically observed in administrative data.
Columns (3) through (5) of (ref) show the results of the IV strategy, which uses the fifth-grade midterm test as an instrument for the fifth-grade end term test. The first stage in column (3) shows that the estimated coefficient on the midterm test is equal to $0.887$. This estimate measures the reliability ratio after multiplication with the fraction containing the variances of the error terms from the balancing regression. Hence, the first stage estimate for the reliability ratio is $0.887 \times 0.992 = 0.880$.
The IV strategy addresses measurement error by using the first stage to correct the reduced form estimate on the midterm term test shown in column (4). The second stage estimate on the end term test in column (5) can be calculated by dividing the reduced form by the first stage, $\frac{0.383}{0.887}=0.432$. The corresponding IV estimate for the conditional SES gap is equal to 0.016, which implies that a one standard deviation increase in parental income is associated with a 1.6 percentage point increase in receiving at least a college track recommendation.
The EIV strategy requires a value for the reliability ratio of the end term test. From the IV strategy we know that the first stage estimate for the reliability ratio is 0.880. Reliability ratios near 0.90 are considered high Cito2017. Depending on the topic, the GMAT is advertised to have a reliability ratio between 0.89 and 0.90 GMAT2023, the GRE between 0.87 and 0.95 ETS2023, and the SAT between 0.89 and 0.93 collegeboard2013.
The EIV FS strategy uses the first stage as an estimate for the reliability ratio to directly address the bias in the OLS estimates. In particular, with the reliability ratio equal to $0.880$, the OLS estimate on the fifth-grade end term test is attenuated by a factor of $0.880$, and the conditional SES gap is overestimated by a value of $\gamma^m \delta^m \big(\frac{1-\lambda}{\lambda}\big)=0.381\times0.229\times\big(\frac{1-0.880}{0.880}\big)=0.012$. Hence, the EIV FS estimates are equal to $\big(\frac{0.381}{0.880}\big)=0.433$ and $0.028-0.012=0.016$, respectively.
The novel EIV TH strategy uses an alternative estimate for the reliability ratio, which is obtained by using the biannual test score data throughout grade two to five. (ref) presents the estimates from the first stage regressions that separately use the midterm and end term tests from grade two to five as independent variable. Instead of using the rightmost estimate on the fifth-grade midterm test as the reliability ratio, the EIV TH approach relies on the first stage estimates of all previous test scores to predict the first stage estimate of a test score very close to the fifth-grade end term test. This approximates the estimate obtained under the ideal conditions of the test-retest method. We use a linear polynomial for the forecasting regression equation. With an $R^2$ of $0.981$ this provides a good fit for the development of the first stage estimates. The estimate for the reliability ratio is the one-period out-of-sample forecast and is equal to $0.899$. The EIV TH strategy produces an estimate for the fifth-grade end term test of 0.424 and an estimate for the conditional SES gap of $0.018$.
(ref) shows the estimates with 95% confidence intervals for the OLS strategy together with all three strategies that aim to address measurement error. The left figure presents the estimates for the end term test and visualizes the attenuation bias of OLS with classical measurement error. The IV and EIV FS strategy inflate the test score estimate by a comparable amount, since the reduced form estimate on the midterm test scores ($0.383\times 0.992=0.380$) is similar to the medium regression estimate on the end term test scores ($0.381$). The right figure presents the estimates for the conditional SES gap. Similar attenuation bias translates into a similar contamination bias, and so the IV and EIV FS strategy also produce essentially the same estimate for the conditional SES gap. The attenuation and contamination bias are somewhat smaller for the EIV TH strategy, since the estimate for the reliability ratio is somewhat higher when using the one-period out-of-sample forecast.
The formula $\Bigg(\frac{\beta^m-\beta^{\bullet}}{\beta^m} \Bigg) \times 100$ quantifies the share of the conditional SES gap in track recommendations attributable to measurement error. Here, $\beta^m$ is the OLS estimate and $\beta^{\bullet}$ refers to the IV and EIV estimates. The IV strategy attributes 42.14% of the SES gap to measurement error, while the EIV strategies yield estimates ranging from 35.13% to 42.78%, depending on how the reliability ratio is estimated. Overall, the results highlight the importance of addressing contamination bias and that, in our setting, the three strategies produce broadly consistent findings.
(ref)(a) and (ref) have direct testable implications. Define $\Delta_{s^m_{t}}=s_{t}^m-s_{t-1}^m$ as the change in grade 5 test scores, and let $\alpha_x = \big(\frac{\mathrm{Cov}[\Delta {s^m_{t}},x]}{\sigma_{x}^2}\big)$ be the OLS estimator of a model that regresses $\Delta_{s^m_{t}}$ upon a single student characteristic $x$. The classical measurement error in our test scores, when they appear on the left-hand side, does not affect the OLS estimate. Hence, (ref)(a) requires that $\alpha_{SES}$ is zero and (ref) requires that $\alpha_{x}$ is zero for every student characteristic $x$, including $SES$.
(ref) in Appendix D shows OLS estimates of the change in grade 5 test scores on SES and all additional control variables. Column (1) shows that the estimate of SES is statistically significant at the 1% level. Although we can reject (ref)(a), the SES estimate of 0.005 seems economically small. In particular, the estimate is reduced by 97.82% compared to the SES estimate from the regression of the fifth-grade end term test score on SES in (ref). Column (2) to (5) further show that the gender dummy, the immigrant dummies, and age also correlate with the change in test scores at the 1% significance level. The coefficients for the immigrant dummies seem of non-negligible size. These results cast doubt on (ref). In general, (ref) seems unlikely in settings where the difference between time period $t$ and $t-1$ is large, especially if the variable is highly dynamic.
Though it is difficult to empirically test (ref)(b), one may discuss its implications for the change in ability over time. In particular, if $\Delta_{s_t}$ is uncorrelated with $s_t$, the following two equalities directly follow:
Low ability students in $t-1$ must have had larger positive changes in ability than high ability students in $t-1$ (from (ref)) and the variance of ability must be larger in $t-1$ than in $t$ (from (ref)). Whether these equalities are likely depends on the setting at hand. In an education context, (ref) and (ref) are consistent with diminishing marginal returns to ability.
The EIV TH strategy also relies, in addition to (ref), on an accurate prediction of changes in the first stage estimate over time. (ref) supports this, showing that a linear polynomial explains 98.1% of the variation in the first stage estimates. This also suggests that the factors driving the first stage away from the reliability ratio, which are $\mathrm{Cov}[\Delta{s_{t}},SES]\neq 0$ and $\mathrm{Cov}[\Delta{s_{t}},s_{t}]\neq 0$, diminish over time in a predictable way. This is further supported by results presented in (ref) in Appendix D, which plots OLS estimates of $\Delta{s^m_{t}}$ on $SES$ across time and shows that $\mathrm{Cov}[\Delta{s^m_{t}},SES]$ converges to zero in an approximately linear fashion.
We test the robustness of our findings in four ways. First, (ref) in Appendix D shows the results when using two alternative measures for the teacher track recommendation: a dummy that equals one if the teacher recommendation is equal to or higher than the theory-oriented vocational track ((ref)(a)) or the university track ((ref)(b)). Second, (ref) in Appendix D shows the results when using three alternative SES measures: a dummy that equals one if parental income is above the median of a student's cohort ((ref)(a)), years of schooling of the highest educated parent ((ref)(b)), and a dummy that equals one if the highest educated parent completed at least college education ((ref)(c)). Third, (ref) in Appendix D shows the results when separately estimating the models per cohort ((ref)(a)-(d)). All results are similar to our baseline findings.
Fourth, as SES is strongly correlated with migration background, a potential concern is that the conditional SES gap at least partly reflects a gap by migration background. To test whether this is the case, we estimated our baseline models for immigrants and natives separately. (ref) in Appendix D shows that the SES estimates are similar in both subgroups.
Teacher bias is typically assessed by examining gaps in teacher's subjective evaluations, conditional on objective measures of student abilities. A key challenge in identifying such bias is the presence of measurement error in test scores, which are used as proxies for student ability. We show that both classical and certain forms of non-classical measurement error introduce equivalent bias in the estimated conditional gap in subjective evaluations. We discuss three different strategies to address this contamination bias. While an IV approach is more commonly used in applied econometric studies, we demonstrate that an EIV approach can address measurement error under weaker assumptions.
The empirical analysis focuses on conditional SES gaps in teacher track recommendations in the Netherlands. This is a setting where potential teacher bias is highly consequential and findings on conditional gaps with respect to parental SES are widely accepted. Using new administrative data, we find that measurement error in test scores can explain a substantial portion of the conditional SES gap, between 35 and 43%.
It may seem striking that measurement error explains such a large share of the conditional SES gap, especially since our test scores are reliable by conventional norms. Yet when student ability is strongly correlated with SES and matters for teacher track recommendations, contamination bias can still be substantial. While our findings show that teacher bias may be smaller than previously reported, the strong correlation between ability and SES points to significant inequalities of opportunity before and during primary education.
\addcontentsline{toc}{section}{References}