Carlos Rodriguez-Pardo, Massimo Tavoni
arXiv 31 Jul 2026 · Machine Learning
arXiv:2607.29527 · PDF · Extracted main text
A defining problem of the Anthropocene is to model the physical Earth and human societies as one coupled system, yet no learned representation spans their observational breadth. We argue the obstacle is geometric: the physical Earth is measured as continuous fields that ignore political borders, whereas societies are reported for administrative units. Earth-system foundation models serve the first geometry; coupling it to the second has required lossy averaging over borders. We introduce TerraNova, a foundation model trained on 1,024 physical and societal records in their native geometries: 512 gridded Earth-system fields and 512 national indicators. Dedicated encoders represent location, country, time and task, cross-modal transformers fuse them into a shared spatiotemporal state, and a hypernetwork generates a per-query decoder whose evidential head returns a predictive distribution. Two contrastive objectives couple the representation: a population-weighted alignment between each country and coordinates in its territory, and one to pretrained geospatial embeddings carrying image-derived semantics. Read out through that decoder, the representation is competitive with purpose-built geospatial encoders while spanning axes they do not represent (time, oceans and uncertainty) and supporting country-level capabilities. The frozen backbone reconstructs dense fields from sparse observations and adapts to unseen variables in minutes on consumer hardware.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Nis Meinert, Jakob Gawlikowski, and Alexander Lavin The unreasonable effectiveness of deep evidential regression | 0.843 | 3 | 3 | 100% |
| 2 | Jiale Kang and Qingyu Yin (2026) Bone: Block Affine Transformation as Parameter Efficient Fine-tuning Methods for Large Language Models | 0.811 | 4 | 2 | 100% |
| 3 | Alexander Amini, Wilko Schwarting, Ava Soleimany, and Daniela Rus (2020) Deep evidential regression | 0.737 | 3 | 2 | 100% |
| 4 | Kaifeng Bi, Lingxi Xie, Hengheng Zhang, Xin Chen, Xiaotao Gu, and Qi… (2023) Accurate medium-range global weather forecasting with 3D neural networks | 0.644 | 2 | 2 | 100% |
| 5 | Cristian Bodnar, Wessel P. Bruinsma, Ana Lucic, Megan Stanley, Anna… (2025) A foundation model for the Earth system | 0.644 | 2 | 2 | 100% |
| 6 | Vicente Vivanco Cepeda, Gaurav Kumar Nayak, and Mubarak Shah (2023) GeoCLIP: CLIP-inspired alignment between locations and images for effective worldwide geo-localization | 0.644 | 2 | 2 | 100% |
| 7 | Noah S Diffenbaugh and Marshall Burke (2019) Global warming has increased global economic inequality | 0.644 | 2 | 2 | 100% |
| 8 | Konstantin Klemmer, Esther Rolf, Caleb Robinson, Lester Mackey, and… SatCLIP: Global, general-purpose location embeddings with satellite imagery | 0.644 | 2 | 2 | 100% |
| 9 | Remi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Wirnsberge… (2023) Learning skillful medium-range global weather forecasting | 0.644 | 2 | 2 | 100% |
| 10 | Tung Nguyen, Johannes Brandstetter, Ashish Kapoor, Jayesh K Gupta, a… (2023) ClimaX: A foundation model for weather and climate | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 79 scored citations.