Vikram Krishnaveti, Saannidhya Rawat
arXiv 16 Sep 2024 · Econometrics · 1 citations (OpenAlex)
arXiv:2409.10750 · PDF · DOI · OpenAlex · Extracted main text
Scholastic Aptitude Test (SAT) is crucial for college admissions but its effectiveness and relevance are increasingly questioned. This paper enhances Synthetic Control methods by introducing "Transformed Control", a novel method that employs Large Language Models (LLMs) powered by Artificial Intelligence to generate control groups. We utilize OpenAI's API to generate a control group where GPT-4, or ChatGPT, takes multiple SATs annually from 2008 to 2023. This control group helps analyze shifts in SAT math difficulty over time, starting from the baseline year of 2008. Using parallel trends, we calculate the Average Difference in Scores (ADS) to assess changes in high school students' math performance. Our results indicate a significant decrease in the difficulty of the SAT math section over time, alongside a decline in students' math performance. The analysis shows a 71-point drop in the rigor of SAT math from 2008 to 2023, with student performance decreasing by 36 points, resulting in a 107-point total divergence in average student math performance. We investigate possible mechanisms for this decline in math proficiency, such as changing university selection criteria, increased screen time, grade inflation, and worsening adolescent mental health. Disparities among demographic groups show a 104-point drop for White students, 84 points for Black students, and 53 points for Asian students. Male students saw a 117-point reduction, while female students had a 100-point decrease.
appendix boundary found by appendix_titled_section at “Appendix A: Concordance Tables” · 62% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Raj Chetty, David Deming, and John Friedman (2023) Diversifying society's leaders? The causal effects of admission to highly selective private colleges | 0.843 | 4 | 3 | 75% |
| 2 | Caroline M. Hoxby (2009) The changing selectivity of american colleges | 0.737 | 4 | 2 | 75% |
| 3 | Andrew Gelman, John B. Carlin, Hal S. Stern, David B. Dunson, Aki Ve… (2013) Bayesian Data Analysis | 0.737 | 3 | 2 | 100% |
| 4 | Desnes Nunes, Ricardo Primi, Ramon Pires, Roberto Lotufo, and Rodrig… (2023) Evaluating GPT-3.5 and GPT-4 Models on Brazilian University Admission Exams | 0.644 | 3 | 2 | 67% |
| 5 | OpenAI (2023) GPT-4 Technical Report, 2023a | 0.644 | 3 | 2 | 67% |
| 6 | John K. Kruschke (2013) Bayesian estimation supersedes the t test | 0.644 | 2 | 2 | 100% |
| 7 | Alberto Abadie and Javier Gardeazabal (2003) The economic costs of conflict: A case study of the basque country | 0.644 | 2 | 2 | 100% |
| 8 | Brantly Callaway and Pedro H.C. Sant’Anna (2020) Difference-in-differences with multiple time periods | 0.585 | 3 | 1 | 100% |
| 9 | FairTest (2023) Test optional and test free colleges, 2022 | 0.511 | 2 | 2 | 50% |
| 10 | Christopher T. Bennett (2021) Untested admissions: Examining changes in application behaviors and student demographics under test-optional policies | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 28 scored citations.