Christian M. Dahl, Torben Johansen, Emil N. Sørensen, Simon Wittrock
arXiv 22 Jan 2021 · cs.CV · publishedExplorations in Economic History (2022) · 5 citations (OpenAlex)
arXiv:2101.10862 · PDF · DOI · OpenAlex · Extracted main text
Methods for linking individuals across historical data sets, typically in combination with AI based transcription models, are developing rapidly. Probably the single most important identifier for linking is personal names. However, personal names are prone to enumeration and transcription errors and although modern linking methods are designed to handle such challenges, these sources of errors are critical and should be minimized. For this purpose, improved transcription methods and large-scale databases are crucial components. This paper describes and provides documentation for HANA, a newly constructed large-scale database which consists of more than 3.3 million names. The database contain more than 105 thousand unique names with a total of more than 1.1 million images of personal names, which proves useful for transfer learning to other settings. We provide three examples hereof, obtaining significantly improved transcription accuracy on both Danish and US census data. In addition, we present benchmark results for deep learning models automatically transcribing the personal names from the scanned documents. Through making more challenging large-scale databases publicly available we hope to foster more sophisticated, accurate, and robust models for handwritten text recognition.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Deng, J., W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei (2009) ImageNet: A large-scale hierarchical image database | 0.644 | 2 | 2 | 100% |
| 2 | Abramitzky, R., L. P. Boustan, K. Eriksson, J. Feigenbaum, and S. Pé… (2020) Automated linking of historical data | 0.511 | 2 | 1 | 100% |
| 3 | Bailey, M., C. Cole, M. Henderson, and C. Massey (2020) How well do automated linking methods perform? Lessons from U.S. historical data | 0.511 | 2 | 1 | 100% |
| 4 | Van Rossum, G. and F. L. Drake (2009) Python 3 Reference Manual | 0.405 | 1 | 1 | 100% |
| 5 | Abramitzky, R., L. P. Boustan, and K. Eriksson (2012) Europe’s tired, poor, huddled masses: Self-selection and economic outcomes in the age of mass migration | 0.405 | 1 | 1 | 100% |
| 6 | Abramitzky, R., L. P. Boustan, and K. Eriksson (2013) Have the poor always been less likely to migrate? Evidence from inheritance practices during the age of mass migration | 0.405 | 1 | 1 | 100% |
| 7 | Abramitzky, R., L. P. Boustan, and K. Eriksson (2014) A nation of immigrants: Assimilation and economic outcomes in the age of mass migration | 0.405 | 1 | 1 | 100% |
| 8 | Abramitzky, R., L. P. Boustan, and K. Eriksson (2016) Cultural assimilation during the age of mass migration | 0.405 | 1 | 1 | 100% |
| 9 | Abramitzky, R., R. Mill, and S. Pérez (2020) Linking individuals across historical sources: A fully automated approach | 0.405 | 1 | 1 | 100% |
| 10 | Myronenko, A. and X. Song (2010) Point set registration: Coherent Point Drift | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 27 scored citations.