Gabriel Okasa, Alberto de León, Michaela Strinzel, Anne Jorstad, Katrin Milzow, Matthias Egger, Stefan Müller
arXiv 25 Nov 2024 · Econometrics · publishedQuantitative Science Studies (2025) · 5 citations (OpenAlex)
arXiv:2411.16662 · PDF · DOI · OpenAlex · Extracted main text
Peer review in grant evaluation informs funding decisions, but the contents of peer review reports are rarely analyzed. In this work, we develop a thoroughly tested pipeline to analyze the texts of grant peer review reports using methods from applied Natural Language Processing (NLP) and machine learning. We start by developing twelve categories reflecting content of grant peer review reports that are of interest to research funders. This is followed by multiple human annotators' iterative annotation of these categories in a novel text corpus of grant peer review reports submitted to the Swiss National Science Foundation. After validating the human annotation, we use the annotated texts to fine-tune pre-trained transformer models to classify these categories at scale, while conducting several robustness and validation checks. Our results show that many categories can be reliably identified by human annotators and machine learning approaches. However, the choice of text classification approach considerably influences the classification performance. We also find a high correspondence between out-of-sample classification performance and human annotators' perceived difficulty in identifying categories. Our results and publicly available fine-tuned transformer models will allow researchers and research funders and anybody interested in peer review to examine and report on the contents of these reports in a structured manner. Ultimately, we hope our approach can contribute to ensuring the quality and trustworthiness of grant peer review.
appendix boundary found by appendix_command · 59% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Severin, A., Strinzel, M., Egger, M., Barros, T., Sokolov, A., Mouat… (2023) Relationship between journal impact factor and the thoroughness and helpfulness of peer reviews self | 1.000 | 9 | 3 | 100% |
| 2 | Forster, M., Schulz, C., Nokku, P., Mirsafian, M., Kasundra, J., and… (2024) The right model for the job: An evaluation of legal multi-label classification baselines | 0.928 | 4 | 3 | 100% |
| 3 | Ghosal, T., Kumar, S., Bharti, P. K., and Ekbal, A (2022) Peer review analyze: A novel benchmark resource for computational analysis of peer reviews | 0.874 | 5 | 2 | 100% |
| 4 | Pelaez, S., Verma, G., Ribeiro, B., and Shapira, P (2023) Large-scale text analysis using generative language models: A case study in discovering public value expressions in AI patents | 0.843 | 4 | 4 | 75% |
| 5 | Bucher, M. J. J. and Martini, M (2024) Fine-tuned `small' LLMs (still) significantly outperform zero-shot generative AI models in text classification | 0.737 | 4 | 2 | 75% |
| 6 | Gupta, A., Norberg, J., Schnidman, E., Viswanathan, S., Zhang, K., a… (2024) From West to the rest: Growing dispersion of AI jobs in America | 0.737 | 3 | 3 | 67% |
| 7 | Hren, D., Pina, D. G., Norman, C. R., and Marusić, A (2022) What makes or breaks competitive research proposals? A mixed-methods analysis of research grant evaluation reports | 0.737 | 3 | 2 | 100% |
| 8 | Fromm, M., Faerman, E., Berrendorf, M., Bhargava, S., Qi, R., Zhang,… (2021) Argument mining driven analysis of peer-reviews | 0.644 | 2 | 2 | 100% |
| 9 | Han, K., Rezapour, R., Nakamura, K., Devkota, D., Miller, D. C., and… (2023) An expert-in-the-loop method for domain-specific document categorization based on small training data | 0.644 | 2 | 2 | 100% |
| 10 | Langfeldt, L., Reymert, I., and Svartefoss, S. M (2024) Distrust in grant peer review — reasons and remedies | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 51 scored citations.