Rafael Macalaba, Aivin V. Solatorio, Patrick Michael Brock, Olivier Dupriez
arXiv 10 Sep 2026 · cs.CL
arXiv:2609.12107 · PDF · Extracted main text
Development and humanitarian organizations produce and support surveys, administrative registries, and other data resources to inform research, policy, and operations, yet systematically identifying where these datasets are referenced remains difficult. Such references are dispersed across research papers, project documents, humanitarian reports, and other unstructured text, limiting both the ability to trace data use and to identify potential gaps in data availability or dissemination. We present a weakly supervised framework for adapting dataset extraction to forced displacement and Fragile, Conflict, and Violence (FCV) documents without first constructing a large manually labeled training corpus. A lightweight model trained on general research literature generates candidate dataset mentions from unlabeled domain documents, which a frontier large language model (LLM) reviews in context, validating or rejecting candidates and correcting their extraction boundaries. The resulting annotations are supplemented with targeted synthetic and contrastive examples and used to fine-tune the lightweight model for large-scale extraction. We evaluate the resulting model on an independent gold-standard benchmark of 1,706 text passages spanning research, humanitarian, and operational documents. Across the full benchmark, the model achieves 74.1% precision and 70.5% recall at the mention level; among passages containing dataset references, precision reaches 89.5%. At the passage level, the model achieves 88.2% accuracy and 88.6% specificity in distinguishing passages with dataset references from those without them. These results demonstrate a practical approach for constructing domain-specific supervision when labeled data are limited, and provide a technical foundation for larger-scale analysis of data use and potential gaps in the displacement data landscape.
appendix boundary found by appendix_command · 79% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Rafael Macalaba and Aivin V. Solatorio (2026) AI for Monitoring and Classifying Data Used in Research Literature self | 0.843 | 3 | 3 | 100% |
| 2 | Buneman, Peter and Dosso, Dennis and Lissandrini, Matteo and Silvell… (2021) Data Citation and the Citation Graph | 0.644 | 2 | 2 | 100% |
| 3 | Hussain, Tayyaba and Akram, Muhammad Usman and Salam, Anum Abdul (2023) A Novel Data Extraction Framework Using Natural Language Processing (DEFNLP) Techniques | 0.644 | 2 | 2 | 100% |
| 4 | Mooney, Hailey and Newton, Mark P (2012) The Anatomy of a Data Citation: Discovery, Reuse, and Credit | 0.644 | 2 | 2 | 100% |
| 5 | Piwowar, Heather A. and Vision, Todd J (2013) Data Reuse and the Open Data Citation Advantage | 0.644 | 2 | 2 | 100% |
| 6 | Silvello, Gianmaria (2018) Theory and Practice of Data Citation | 0.644 | 2 | 2 | 100% |
| 7 | Zaratiana, Urchade and Tomeh, Nadi and Holat, Pierre and Charnois, T… (2023) GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer | 0.644 | 2 | 2 | 100% |
| 8 | Heddes, Jenny and Meerdink, Pim and Pieters, Miguel and Marx, Maarten (2021) The Automatic Detection of Dataset Names in Scientific Articles | 0.405 | 1 | 1 | 100% |
| 9 | (2020) Data Collection in Fragile States: Innovations from Africa and Beyond | 0.405 | 1 | 1 | 100% |
| 10 | Inter-Agency Standing Committee (2023) IASC Operational Guidance on Data Responsibility in Humanitarian Action | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 22 scored citations.