EconBase
← All papers

With big data come big problems: pitfalls in measuring basis risk for crop index insurance

Matthieu Stigler, Apratim Dey, Andrew Hobbs, David Lobell

arXiv 29 Sep 2022 · Econometrics

arXiv:2209.14611 · PDF · DOI · OpenAlex · Extracted main text

Abstract

New satellite sensors will soon make it possible to estimate field-level crop yields, showing a great potential for agricultural index insurance. This paper identifies an important threat to better insurance from these new technologies: data with many fields and few years can yield downward biased estimates of basis risk, a fundamental metric in index insurance. To demonstrate this bias, we use state-of-the-art satellite-based data on agricultural yields in the US and in Kenya to estimate and simulate basis risk. We find a substantive downward bias leading to a systematic overestimation of insurance quality. In this paper, we argue that big data in crop insurance can lead to a new situation where the number of variables $N$ largely exceeds the number of observations $T$. In such a situation where $T\ll N$, conventional asymptotics break, as evidenced by the large bias we find in simulations. We show how the high-dimension, low-sample-size (HDLSS) asymptotics, together with the spiked covariance model, provide a more relevant framework for the $T\ll N$ case encountered in index insurance. More precisely, we derive the asymptotic distribution of the relative share of the first eigenvalue of the covariance matrix, a measure of systematic risk in index insurance. Our formula accurately approximates the empirical bias simulated from the satellite data, and provides a useful tool for practitioners to quantify bias in insurance quality.

Citation extraction

37
references
51
in-text mentions
37
distinct cited
8
self-citations
6,476
main-text words

appendix boundary found by appendix_command · 90% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Stigler, M. and D. Lobell (2021) Optimal index insurance and basis risk decomposition: an application to Kenya, Tech self0.81142100%
2Ahn, J., J. S. Marron, K. M. Muller, and Y.-Y. Chi (2007) The High-Dimension, Low-Sample-Size Geometric Representation Holds under Mild Conditions0.7375260%
3Conradt, S., R. Finger, and R. Bokusheva (2015) Tailored to the extremes: Quantile regression for index-based insurance contract design0.64422100%
4Hall, P., J. S. Marron, and A. Neeman (2005) Geometric Representation of High Dimension, Low Sample Size Data0.64422100%
5Johnstone, I. M (2001) On the Distribution of the Largest Eigenvalue in Principal Components Analysis0.64422100%
6Deines, J. M., R. Patel, S.-Z. Liang, W. Dado, and D. B. Lobell (2021) A million kernels of truth: Insights into scalable satellite maize yield mapping and yield gap analysis from an extensive ground…0.51121100%
7Jin, Z., G. Azzari, C. You, S. Di Tommaso, S. Aston, M. Burke, and D… (2019) Smallholder maize area and yield mapping at national scales with Google Earth Engine self0.51121100%
8Koenker, R. and J. A. F. Machado (1999) Goodness of Fit and Related Inference Processes for Quantile Regression0.51121100%
9Carter, M., A. de Janvry, E. Sadoulet, and A. Sarris (2017) Index insurance for developing country agriculture: a reassessment0.51121100%
10Aoshima, M., D. Shen, H. Shen, K. Yata, Y.-H. Zhou, and J. S. Marron (2018) A survey of high dimension low sample size asymptotics0.40511100%

Showing the top 10 of 37 scored citations.