Christian M. Dahl, Torben S. D. Johansen, Emil N. Sørensen, Christian E. Westermann, Simon F. Wittrock
arXiv 5 Feb 2021 · cs.CV · 4 citations (OpenAlex)
arXiv:2102.03239 · PDF · DOI · OpenAlex · Extracted main text
Data acquisition forms the primary step in all empirical research. The availability of data directly impacts the quality and extent of conclusions and insights. In particular, larger and more detailed datasets provide convincing answers even to complex research questions. The main problem is that 'large and detailed' usually implies 'costly and difficult', especially when the data medium is paper and books. Human operators and manual transcription have been the traditional approach for collecting historical data. We instead advocate the use of modern machine learning techniques to automate the digitisation process. We give an overview of the potential for applying machine digitisation for data collection through two illustrative applications. The first demonstrates that unsupervised layout classification applied to raw scans of nurse journals can be used to construct a treatment indicator. Moreover, it allows an assessment of assignment compliance. The second application uses attention-based neural networks for handwritten text recognition in order to transcribe age and birth and death dates from a large collection of Danish death certificates. We describe each step in the digitisation pipeline and provide implementation insights.
appendix boundary found by appendix_titled_section at “Appendix” · 99% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Graves, Alex, Liwicki, Marcus, Bertolami, Roman, Bunke, Horst (2008) A novel connectionist system for unconstrained handwriting recognition | 0.874 | 6 | 2 | 100% |
| 2 | Lee, Chen-Yu, Osindero, Simon (2016) Recursive recurrent nets with attention modeling for ocr in the wild | 0.737 | 3 | 2 | 100% |
| 3 | Murphy, Kevin P (2012) Machine learning : a probabilistic perspective | 0.737 | 3 | 2 | 100% |
| 4 | Xu, Kelvin, Ba, Jimmy, Kiros, Ryan, Cho, Kyunghyun, Courville, Aaron… (2015) Show, attend and tell: Neural image caption generation with visual attention | 0.693 | 9 | 1 | 100% |
| 5 | Bay, Herbert, Ess, Andreas, Tuytelaars, Tinne, Van Gool, Luc (2008) Speeded-up robust features (SURF) | 0.644 | 2 | 2 | 100% |
| 6 | Gutmann, Myron P., Merchant, Emily Klancher, Roberts, Evan (2018) “Big Data” in Economic History | 0.585 | 3 | 1 | 100% |
| 7 | Myronenko, Andriy, Song, Xubo (2010) Point set registration: Coherent point drift | 0.585 | 3 | 1 | 100% |
| 8 | Csurka, Gabriella, Dance, Christopher, Fan, Lixin, Willamowski, Jutta (2004) Visual categorization with bags of keypoints | 0.511 | 2 | 1 | 100% |
| 9 | (2017) Mask R-CNN | 0.511 | 2 | 1 | 100% |
| 10 | Simonyan, Karen, Zisserman, Andrew (2015) Very Deep Convolutional Networks for Large-Scale Image Recognition | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 50 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | HANA: A HAndwritten NAme Database for Offline Handwritten Text Recognition | 0.405 | 1 | 1 |