EconBase
← All papers

Proper Correlation Coefficients for Nominal Random Variables

Jan-Lukas Wermuth

arXiv 1 May 2025 · Statistics — Methodology

arXiv:2505.00785 · PDF · DOI · OpenAlex · Extracted main text

Abstract

This paper develops an intuitive concept of perfect dependence between two variables of which at least one has a nominal scale that is attainable for all marginal distributions and proposes a set of dependence measures that are 1 if and only if this perfect dependence is satisfied. The advantages of these dependence measures relative to classical dependence measures like contingency coefficients, Goodman-Kruskal's lambda and tau and the so-called uncertainty coefficient are twofold. Firstly, they are defined if one of the variables is real-valued and exhibits continuities. Secondly, they satisfy the property of attainability. That is, they can take all values in the interval [0,1] irrespective of the marginals involved. Both properties are not shared by the classical dependence measures which need two discrete marginal distributions and can in some situations yield values close to 0 even though the dependence is strong or even perfect. Additionally, I provide a consistent estimator for one of the new dependence measures together with its asymptotic distribution under independence as well as in the general case. This allows to construct confidence intervals and an independence test, whose finite sample performance I subsequently examine in a simulation study. Finally, I illustrate the use of the new dependence measure in two applications on the dependence between the variables country and income or country and religion, respectively.

Citation extraction

50
references
77
in-text mentions
50
distinct cited
3
self-citations
10,194
main-text words

appendix boundary found by appendix_command · 60% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Weiß, Christian H and Göb, Rainer (2008) Measuring serial dependence in categorical time series0.84333100%
2Pohle, Marc-Oliver and Wermuth, Jan-Lukas and Weiß, Christian H (2025) Asymptotic Inference for Rank Correlations self0.8229456%
3Pohle, Marc-Oliver and Wermuth, Jan-Lukas (2026) Proper Correlation Coefficients for Discrete Random Variables self0.81142100%
4Hajo Holzmann and Bernhard Klar (2024) Lancaster correlation – a new dependence measure linked to maximum correlation0.73732100%
5Pohle, Marc-Oliver and Dimitriadis, Timo and Wermuth, Jan-Lukas (2024) Measuring Dependence between Events self0.73732100%
6Zurlo, Gina A (2025) World Religion Database0.64422100%
7Gebelein, Hans (1941) Das statistische Problem der Korrelation als Variations- und Eigenwertproblem und sein Zusammenhang mit der Ausgleichsrechnung0.64422100%
8Geenens, Gery (2020) Copula modeling for discrete random vectors0.64422100%
9Hirschfeld, Herman Otto (1935) A Connection between Correlation and Contingency0.64422100%
10Tchen, André H (1980) Inequalities for distributions with given marginals0.64422100%

Showing the top 10 of 50 scored citations.