Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
14,457 characters · 4 sections · 11 citation commands
A Note on Identification of Match Fixed Effects as Interpretable Unobserved Match Affinity
Evaluating the quality of matching and affinity between entities is common in empirical research. Affinity is divided into observed and unobserved components. Observed affinity is typically measured by the coefficient of interaction terms involving observable characteristics, while unobserved affinity is captured through match fixed effects using dummy variables. This approach is widely used in labor and education economics to analyze and interpret match quality across pairs such as teacher-student and worker-company relationships. Our study aims to clarify the identification process of match fixed effects.
As an illustrative example, inoue2023teachers investigate the impact of a teacher's major on students' achievement using the following econometric model:
where \( Y_{ifj} \) represents the science test score of student \( i \) in subfield \( f \) within class \( j \). The parameter \( \beta \) denotes the coefficient of interest, and \( Major_{fj} \) is an indicator variable that indicates whether the teacher's major field in natural science matches the subfield of the student's test score. \( \delta_{f} \) denotes the fixed effect specific to subfield \( f \), while \( \eta_{ij} \) represents the student-teacher fixed effects, accounting for any subfield-invariant determinants of science test scores between student \( i \) and teacher \( j \), thereby capturing their unobserved affinity. The authors find a significant increase in R-squared upon introducing student-teacher fixed effects, underscoring the importance of match fixed effects in their analysis.
Similar methodologies are utilized in various studies, including pitcher-catcher fixed effects on strikeout likelihood biolsi2022task, worker-company fixed effects on income mittag2019simple and turnover rates ferreira2011measuring, student-school fixed effects on student performance ovidi2022parents, teacher-school fixed effects on student test scores jackson2013match, and student-university fixed effects on post-graduation income dillon2020consequences. Additionally, some studies account for two-way fixed effects by separately considering the fixed effects on each side.
Despite their common use, the interpretation of match fixed effects as affinity indicators for all matches remains unclear. Our paper addresses this by distinguishing between relative and absolute match fixed effects within a standard model. We argue that absolute match fixed effects are not identifiable without specific restrictions, whereas relative match fixed effects are identifiable. We propose location normalization conditions to enable the identification and interpretation of absolute match fixed effects, facilitating comparisons of unobserved match quality.
Applying this approach to the data from inoue2023teachers, we demonstrate the distribution of unobserved match quality between students and teachers in a specific school. Our analysis reveals that one teacher excels with high-achieving students but underperforms with lower-achieving ones, while another teacher shows the opposite pattern in terms of unobserved match affinity.
We consider the following typical setting. Suppose that we can observe $I$ students indexed by $i$ and $J$ teachers indexed by $j$ at time $t=1,\cdots,T$. Student $i$ makes some outcome with teacher $j$ at time $t$. For avoiding later notational complexity, we do not include time fixed effects like a panel regression, but the inclusion does not affect our findings. We consider the following regression model:
where $Y_{ijt}$ is the test score of student $i$ with teacher $j$ at time $t$, $X_{ij}$ is $d$-dimensional covariates consisting of observed characteristics of student $i$ and teacher $j$ and its interaction at time $t$, $\alpha_{i}$ is student $i$'s fixed effect, $\beta_{j}$ is teacher $j$'s fixed effect, and $\mu_{ij}$ is $(i,j)$-match fixed effect which is of our interest, $\gamma$ is a $d$-dimensional vector of parameters, $1(\cdot)$ is an indicator function, and $\varepsilon_{ijt}$ is an error term assumed to be drawn i.i.d from standard normal distribution. Note that we explicitly decompose affinity between student $i$ and teacher $j$ into two parts, that is, $X_{ijt}'\gamma$ and $\mu_{ij}$ as the observed and unobserved affinities. For later discussion, we call $\mu_{ij}$ the absolute match effect of student $i$ and teacher $j$. Similarly, we call $\alpha_{i}$ and $\beta_{j}$ the absolute fixed effects.\footnote{Using Equation (ref), matrix representation is described as $Y=X\gamma+1_{I}\alpha+1_{J}\beta+1_{IJ}\mu+\varepsilon=X\gamma+\tilde{X}\delta + \varepsilon$ where $Y$ is $IJT\times 1$, $X$ is $IJT\times d$, $\gamma$ is $d\times 1$, $1_{I}$ is $IJT\times I$, $\alpha$ is $I\times 1$, $1_{J}$ is $IJT\times J$, $\beta$ is $J\times 1$, $1_{IJ}$ is $IJT\times IJ$, $\mu$ is $IJ\times 1$, $\varepsilon$ is $IJT\times 1$, and denote $\tilde{X}=[1_{I} 1_{J} 1_{IJ}]$ and $\delta=[\alpha^T \beta^T \mu^T]^T$. Then, define $I$ as $IJT\times 1$ one vector and $M=I-\tilde{X}(\tilde{X}^T\tilde{X})^{-1}\tilde{X}^{T}$ as the annihilator matrix for $\tilde{X}$. The OLS estimator of $\delta$ is obtained as $\hat{\delta}=(\tilde{X}^T M\tilde{X})^{-1}(\tilde{X}^T MY)$, so the standard full rank condition for $\delta$ is $Rank(\tilde{X}^T M\tilde{X})=(I+J+IJ)$. See hansen2022econometrics Chapter 3.16 for reference. However, the condition does not hold due to multicollinearity.}
In the context of the standard fixed effect model, when there are \( I \) groups, typically \( I - 1 \) fixed effects are incorporated, alongside a constant term, to avoid multicollinearity that would arise from including the \( I \)-th group. The match fixed effect case is more complex. Similarly, for a student \( i \neq 1 \) and teacher \( j \neq 1 \), Equation (ref) can be rewritten as
where the first line is a constant parameter normalized to student 1 and teacher 1 which are arbitrarily chosen, the second line is called teacher $j$'s relative fixed effect which is the fixed effect relative to teacher 1's fixed effects, the third line is called student $i$'s relative fixed effect which is the fixed effect relative to student 1's fixed effects, the third line is called the relative match effect of student $i$ and teacher $j$ relative to student 1 and teacher 1. Avoiding multicollinearity, we can identify and estimate these relative fixed effects and $\gamma$ instead of absolute fixed effects. Statistical software automatically reports the estimates of the relative fixed effects.
Relative fixed effects indicate how a match compares to a specific reference match, rather than categorizing it as "good" or "bad" compared to all other matches. For controlling match fixed effects or obtaining overall affinity without distinguishing observed from unobserved components, relative fixed effects are adequate. However, for measuring and understanding unobserved affinity, such as personality match quality, relative fixed effects are insufficient. In these cases, absolute fixed effects are crucial for meaningful comparison and interpretation of unobserved affinity.
Our central question posed is: “Can we derive the absolute fixed effects from the estimated relative fixed effects without imposing any restrictions?" In mathematical terms, “Can we solve the system of equations involving relative fixed effects for absolute fixed effects without restrictions?" Our conclusion is in the negative: No, we cannot achieve this without imposing constraints. Subsequently, we propose location normalization restrictions that are necessary for identifying the absolute fixed effects \( \mu_{ij} \), summarized in Proposition (ref).
See the proof and illustrative example in Appendix (ref). Intuitively, the conditions outlined are effective for location normalization. The restricted match effect indicates how much better the match is compared to the average match, rather than a specific match. Consequently, it yields positive values when the match is relatively superior and negative values when it is relatively inferior compared to the average match. This approach ensures interpretability across students and teachers, facilitating meaningful comparisons of unobserved affinity.
To illustrate our approach, we use data from the 2007 Trends in International Mathematics and Science Study (TIMSS), as in inoue2023teachers. This dataset includes middle school students' test scores in various science subfields (physics, chemistry, biology, and Earth science) and teacher-related variables. For details, refer to inoue2023teachers. We focus on data from a specific school (ID 228) with five teachers and 90 students, using Equation (ref) to estimate absolute match fixed effects with location normalization. We exclude subfield variables to avoid multicollinearity with student-teacher interaction dummies, as some subjects are taught exclusively by certain teachers. The author's GitHub page provides Monte Carlo simulation code and results, as well as replication files for our empirical exercises.
Figure (ref) presents a heatmap of absolute match fixed effects. Teacher ID 22802 shows better effects with high-performing students but poorer effects with low-performing students, while teacher ID 22804 displays the opposite pattern. This suggests that teacher 22802 excels with high achievers but underperforms with lower achievers, whereas teacher 22804 performs better with lower achievers. Furthermore, absolute match effects offer a more comparable measure than relative match effects. For instance, the absolute match effect for student ID 2280201 and teacher ID 22802 is 76.34, compared to 44.57 for student ID 2280115 and teacher ID 22804, indicating that the former match has 1.75 times higher unobserved affinity. Thus, absolute fixed effects are crucial for meaningful comparisons and interpretations of unobserved affinity.
We examine the estimation of affinities using match fixed effects, distinguishing between observed and unobserved components. We emphasize that, without proper restrictions, only relative fixed effects—fixed effects relative to normalized values—can be estimated, which do not provide interpretable measures of unobserved affinity. To enable meaningful interpretation of absolute match effects, we introduce theoretical constraints on parameters. Using 2007 TIMSS data on middle school students, we illustrate the distribution of absolute match fixed effects and underscore their significance.
\paragraph{Acknowledgments} We thank Shunya Noda, Shosei Sakaguchi, and Ryuichi Tanaka for their valuable advice. This work was supported by JST ERATO Grant Number JPMJER2301, Japan.