Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
69,184 characters · 23 sections · 53 citation commands
Proper Correlation Coefficients for Nominal Random Variables
Keywords: Goodman-Kruskal's Gamma; Confidence Intervals; Independence Tests; Statistical Dependence \newline
Most random variables featuring in statistical applications are real-valued or exhibit at least an ordinal scale. Excluding stochastic processes, which can be interpreted as function-valued random variables, the only remaining type of random variables that appear as data in reality are those with a nominal scale. That is, they map to a set $\Omega'$ with finite cardinality consisting of elements that do not allow for any meaningful ordering (see Subsection (ref) for details). While those variables are less interesting from a mathematical point of view as less measures of central tendency or dispersion and no cdf exists, they are still of immense practical relevance.
In particular, the question of how to characterize statistical dependence between random variables when at least one variable is measured on a nominal scale remains largely unexplored in the literature. Most existing work on dependence assumes continuous marginals, since the copula representation is particularly well-behaved in this case sklar1959, nelsenIntroduction2006, genestEverything2007, durante2016principles. Notable exceptions include geenens2020copula, neslehova2007rank, perroneGeometry2019, genestEstimation2013, and genest2007primer. However, these studies focus on random variables whose values exhibit a reasonable ordering.
For example, consider Table (ref) in which the absolute frequencies of Christians, Jews and Muslims living in Germany, Poland and Czechia are displayed. It is evident that the marginal frequencies do not allow for a concentration of either religion in one country, which would probably be the plainest type of (perfect) dependence possible. Instead, we have three primarily Christian countries which have Jewish and Muslim minorities of differing size. In fact, since Poland and Czechia have in the past enacted much stricter migration policies than Germany, the Muslims in this example are almost fully concentrated in Germany. The same holds true for Jews whose (re)settlement in Germany has been considered a great success after the Shoah. Taking the marginal frequencies as given, it is hard to imagine any bivariate distribution with a stronger dependence than the one displayed (besides removing Jews and Muslims completely out of Poland and Czechia). Note that fixing the marginals should always be the starting point for any dependence assessment since the two marginals and the dependence are three orthogonal objects, i.e.\ objects not carrying any information about each other geenens2023dependence. Still, a selection of dependence measures in Table (ref) indicates little to no dependence. This raises the question to which extent those dependence measures sensibly measure the dependence inherent in this table.
In mathematical terms, the nominal case creates several problems which explains why the stated dependence measures behave as they do. As no bivariate cdf exists, Fréchet-Hoeffding bounds frechet1951tableaux, hoeffding1940 and a (partial) dependence ordering tchen1980inequalities, yanagimoto1969partial are not available. Consequently, dependence measures for the nominal case such as contingency coefficients, Goodman-Kruskal's lambda and tau and the uncertainty coefficient cannot adhere to such concepts which causes serious problems in the interpretation of these measures. For example, it is easily possible to construct marginal distributions which constrain the range of these coefficients from $[0,1]$ to $[0,c)$ with $0\le c\le1$ (see Section (ref)). This is especially true for the introductory example and explains why the reported values are so small. However, since the marginal distributions have nothing to do with the dependence (see again geenens2023dependence), any dependence measure showing such behavior is surely suboptimal.
In addition to that, traditional measures critically rely on the assumption that both variables whose dependence shall be assessed are discrete. Thus, nominal-continuous comparisons such as country and income (on an individual level) or nominal-mixed comparisons such as country and precipitation (on a municipality level) are impossible.
The ambition of this paper therefore firstly is to offer a proposal for a definition of perfect dependence in the nominal case that is attainable irrespective of the marginals involved in Section (ref). Section (ref) collects a set of desirable properties whose fulfillment qualifies a dependence measure as proper, one of those being that the measure attains the upper bound 1 if and only if this concept of perfect dependence is satisfied. To the best of my knowledge, existing sets of desirable properties are almost exclusively tailored towards real-valued random variables Renyi1959, Schweizer1981, scarsini1984measures, Mari2001, Embrechts2002, Balakrishnan2009, fissler2023generalised. Appendix B in weiss2008measuring marks a noteworthy exception.
Section (ref) shows the improperness of all existing dependence measures by exploiting their inability to deal with continuities in one marginal distribution and their non-attainability, i.e.\ their characteristic that the upper bound 1 can only be attained if the marginal distributions exhibit a specific structure.
Section (ref) proposes a set of measures that cure the shortcomings of the improper measures in the nominal-nominal case and comply with the properties needed for properness. Additionally, it establishes an analogy between one of those new measures and the maximal correlation by gebelein1941maximal and hirschfeld1935maximal, an undirected dependence measure for real-valued random variables mapping to $[0,1]$. Moreover, it features a proposal for a dependence measure in the nominal-continuous case.
Section (ref) defines a consistent estimator for the measure(s) based on Goodman-Kruskal's $\gamma$ Goodman1954 and computes its asymptotic distribution, both with and without the assumption of independence between the involved random variables. These results can in turn be used to construct confidence intervals and an independence test. Both methods of statistical inference will subsequently (in Section (ref)) be examined with respect to their finite sample performance in a simulation study that features a selected number of DGPs.
The paper closes with two applications in Section (ref), one concerning the dependence between country and income and the other concerning the dependence between country and religion. Section (ref) concludes. The Appendix extends Section (ref), falsifies an alternative concept of perfect dependence and explains which algorithms lead to a fast computation of the estimator. Additionally, it contains all the proofs as well as additional figures.
\sloppy Moreover, I provide an R RCoreTeam package NCor with an implementation of the estimator (by using the algorithms introduced in Appendix (ref)), its confidence intervals and the independence test under https://github.com/jan-lukas-wermuth/NCor. The replication material for all the results in the paper is available under https://github.com/jan-lukas-wermuth/replication_NCor.
Consider two random variables $X$ and $Y$ defined on the same probability space $(\Omega, \mathcal{F}, \mathbb{P})$. On the one hand, at least one of the variables has a nominal scale, i.e.\ it maps to a set $\Omega'$ $(\Omega'')$ with $3\leq|\Omega'|<\infty$ $(3\leq|\Omega''|<\infty)$ consisting of arbitrary elements without a natural ordering.\footnote{It can, for example, be an alphabet, the set of all religions or the set of all nationalities.} We assume that these sets contain only elements with positive probability. The reason why the binary case is excluded is that you can always interpret such variables on an ordinal scale, i.e.\ as an event happening or not, and we have already treated this case in pohle2024measuring. On the other hand, at most one of the variables exhibits at least an ordinal scale and therefore maps to $\mathbb{R}$ (or, at least, can be coded by elements of $\mathbb{R}$). Consequently, the marginal distribution(s) of the nominal variable(s) is (are) discrete, and the marginal distribution of the potential other random variable can be either discrete, continuous, or mixed. We are interested in measuring the strength of dependence between these two random variables. Any attempt to measure the direction of dependence is doomed due to the presence of at least one nominal random variable, which makes such an interpretation impossible.\footnote{weiss2008measuring mention a possibility to define negative dependence also in the nominal case. However, a prerequisite for this definition is that the ranges of $X$ and $Y$ are identical. This requirement is considered too strict for our purposes.}
I now introduce a concept of perfect dependence for nominal random variables. Any such concept must be invariant with respect to the marginal distributions, i.e.\ it must be possible to attain perfect dependence irrespective of the shape and values that the involved marginals take pohle2024measuring, pohle2024. The view that this paper holds is that dependence within real-valued random vectors is not fundamentally different from the dependence within the vectors considered here. Therefore, we can start our analysis with an arbitrary $\mathbb{R}^2$-valued random vector $(X,Y)$ with at least one discrete marginal distribution. We assume that this vector exhibits perfect positive or negative dependence. This is equivalent to its CDF $F_{X,Y}$ being identical to the Fréchet-Hoeffding upper bound $(F_{X,Y}(x,y):=\min\{F_X(x),F_Y(y)\})$ or the Fréchet-Hoeffding lower bound $(F_{X,Y}(x,y):=\max\{0,F_X(x)+F_Y(y)-1\})$, respectively. The replacement of the values of the discrete component(s) by elements of $\Omega'$ $(\Omega'')$ shall now only eliminate the directional interpretation. The attribute of perfectness however is sustained such that $(X,Y)$ is perfectly dependent, without any notion of positive or negative dependence.
The following lemma shows that the above definition is unnecessarily complicated. In principle, we could drop one of the bounds from our definition.
We want to summarize the dependence between $X$ and $Y$ in a single number, a dependence measure.
As already mentioned before, a fundamental difference to the case of two real-valued random variables is the impossibility to interpret directions of dependence. Thus, this paper aims to construct a dependence measure that indicates only the strength of dependence, but not the direction. In order to perform a substantial assessment of whether a measure deserves the attribute of properness, we need a set of desirable properties.
Relative to the real-valued case, the most remarkable difference is the absence of negative values in the co-domain of the dependence measure. This is natural because without the option to measure direction of dependence, there is no need to distinguish between positive and negative values. Additionally, a (partial) dependence ordering similar to yanagimoto1969partial and tchen1980inequalities is in my opinion difficult to construct and therefore left for future research. As a consequence, I also do not include any property formalizing compliance of a dependence measure with such a concept. Finally, the attainability axiom uses the novel definition of perfect dependence and is thus fundamentally different from the attainability definition in the real-valued case.
This section together with Appendix (ref) reviews the most popular dependence measures for the nominal case and thus extends Appendix B in weiss2008measuring. More precisely, it considers several measures based on Pearson's mean square contingency coefficient. Appendix (ref) contains measures based on proportional reduction of predictive error and an entropy-based measure. Measures of agreement such as Cohen's $\kappa$ cohen1960agreement are left out of the discussion because their assumption that $X$ and $Y$ take exactly the same values is considered too strict for our purposes.
The most popular dependence measures for the nominal case arise as normalizations of Pearson's Mean Square Contingency $(MSC)$ coefficient (see Table (ref) and liebetrau1983measures for an overview). This coefficient is the population analogue of the test statistic in Pearson's $\chi^2$ independence test. It is only defined if both marginal distributions are discrete, i.e.\ if $X$ takes the values $x_1, ..., x_a$ and $Y$ takes the values $y_1, ..., y_b$:
Cramér's $V$ Cramer1945, Tschuprow's $T$ Tschuprow1925, Pearson's Contingency Coefficient $PC$ Pearson1904 and Sakoda's $S$ sakoda1977measures are defined as
Even though their improperness is immediate due to the necessity to have two discrete marginal distributions, it is insightful to analyze their violation of the attainability property in more detail. The following lemma is well-known in the literature.
If we interpret the previous lemma with a contingency table with $a$ rows and $b$ columns (see Table (ref)), the second statement is equivalent to each row (if $a\ge b$) or each column (if $a\le b$) containing only one non-zero entry. Equipped with this lemma, we can analyze the attainability properties of the other coefficients.
Given the conditions under which the coefficients attain their maximum value of 1, it is evident that it is easily possible to construct marginal distributions which do not allow a fulfillment of these conditions (see e.g.\ the introductory example). Hence, they can be regarded as too strict as the following lemma shows.
The essential idea behind the construction of proper dependence measures is the concept of perfect dependence explained in Subsection (ref) together with property (D) from Definition (ref).
Since Goodman-Kruskal's $\gamma$ is a proper dependence measure with straightforward inference procedures pohle2024inference, the subsequent analysis will focus on this special case.
Even though $\gamma^*$ is a measure for nominal random variables, the idea behind its construction is less connected with existing measures for the nominal case than with existing measures for $\mathbb{R}^2$-valued random vectors that map to $[0,1]$. Basically, those measures sacrifice the directional interpretation for a stronger interpretation of the 0 (the measures are 0 if and only if $X$ and $Y$ are independent) and a wider concept of perfect dependence which allows those measures to detect U-shaped relationships, for example. One particularly similar coefficient has been introduced by hirschfeld1935maximal and gebelein1941maximal and is since called the maximal correlation. It is defined as
where $r$ denotes the Pearson correlation and the supremum is taken over the set of all Borel-measurable functions $f,g:\mathbb{R}\to \mathbb{R}$. If we now consider
and restrict the set of functions over which we maximize to the set of all bijective functions, an interesting simplification emerges.
Intuitively, the rank-based nature of $\gamma$ makes it sufficient to consider all possible orderings of the values that $X$ and $Y$ can take. In the case of two discrete distributions, the innovation of our proposal is now to consider nominal random variables together with functions $f:\Omega'\to \mathbb{R}$ (and $g: \Omega''\to \mathbb{R}$)\footnote{holzmann2024lancaster already mention the ad hoc extension of the maximal correlation to $\mathcal{X}$- and $\mathcal{Y}$-valued random variables $X$ and $Y$, where $\mathcal{X}$ and $\mathcal{Y}$ are arbitrary measurable spaces.} and maximize $\gamma$, a correlation coefficient with superior properties relative to Pearson correlation pohle2024.
The population coefficient introduced in the previous section is only of practical use if it is possible to construct a consistent estimator for it. Additionally, knowledge about the asymptotic distribution could allow us to construct asymptotically valid confidence intervals as well as an independence test based on the coefficient. In the following, $\overset{d}{\to}$ denotes convergence in distribution and $\overset{p}{\to}$ denotes convergence in probability.
Our estimator simply replaces the theoretical distribution $\mathcal{L}_{X,Y}$ implicitly used in Definition (ref) with the empirical distribution $$\widehat{\mathcal{L}}_{X,Y}:=\frac{1}{n}\sum_{i=1}^{n}\delta_{X_i,Y_i}.$$ We make the following sampling assumption:
Note that however consistent, the estimator will in general be biased in finite samples (see Figure (ref) for a simulation study which shows small mean biases for moderate sample sizes). Due to the small size of the biases and the strong limits that are placed upon bias correction techniques by hirano2012impossibility, I refrain from proposing alternative estimators.\footnote{Also, the results by chernozhukov2013intersection constitute no remedy for this problem since half median unbiasedness is of limited value for a point estimate.}
Imagine the true (or one of potentially many true) population numbering(s) was known to the researcher. Then, an alternative estimator $\widehat{\gamma}(X^{s_k^{(*)}},Y)$ or $\widehat{\gamma}(X^{s_k^{(*)}},Y^{s_l^{(*)}})$ could be considered in which $s_k^{(*)}$ and $s_l^{(*)}$ are fixed. In this case, the standard asymptotic distribution of $\widehat{\gamma}$ could be used for statistical inference pohle2024inference. Unfortunately, this is in reality never the case and we have no choice but revert to the estimator introduced in Definition (ref). Since the maximum function in this estimator always chooses one specific numbering and can (also asymptotically) oscillate between several of those, the asymptotic distribution is non-standard. In order to calculate it, we first need the asymptotic distribution of the vector of inputs at the left hand side of ((ref)). Since the length of this vector depends on the number of values that $X$ (and $Y$) can take, this paper can only provide a strategy to compute the asymptotic variance and no closed-form solution.
Given Proposition (ref) (and the precise structure of the involved matrices, which is given in the proof), we can analyze the asymptotic distribution of our estimator. Since the maximum function is not differentiable, we cannot use the standard Delta Method to obtain our result.\footnote{However, a generalized Delta Method by fang2019inference is applicable and constitutes an alternative way to obtain the result.} Instead, the asymptotic distribution is non-standard and contains several case distinctions. In case our population exhibits perfect dependence in the sense of Definition (ref), the estimator is almost surely equal to 1 (if the sample is large enough such that there is variation in either of the two observed variables). Therefore, we exclude this case in the subsequent proposition.
Constructing confidence intervals for the new correlation coefficient is complicated due to the non-standard asymptotic distribution described in Proposition (ref). Although a complicated distribution could still be exploited (see, e.g., the confidence intervals for Cole's C cole1949 in pohle2024measuring constructed by an inverted testing procedure), the problem is that the researcher does not know which of the many possible asymptotic distributions prevails in her particular setting.
Facing a similar problem, holzmann2024lancaster propose several methods to construct confidence intervals using standard procedures or bootstrap. For some of those methods, they also take into account the case distinctions they have in their asymptotic distribution. I refrain from such an involved approach for the following reasons. Firstly, I do not consider bootstrap because due to the rapidly increasing nature of the factorial, estimating the coefficient is already computationally intensive enough (see also footnote (ref)). Secondly, the number of case distinctions in our asymptotic distribution is so large that I do not see any promising path to combine them in an attempt to construct (possibly conservative) confidence intervals (see Remark (ref)). Instead, we make a simplifying assumption.
With this assumption (and the assumption that $(X,Y)$ are not perfectly dependent), we are in the well-behaved first case of Proposition (ref) and can construct our confidence intervals with asymptotic coverage probability $1-\alpha$ as
where $z_{1-\alpha/2}=\Phi^{-1}(1-\alpha/2)$ and $\widehat{\sigma}_{\gamma^*}$ is equal to $\widehat{\sigma}_{\gamma}$ as defined in pohle2024inference relative to the numbering that $\widehat{\gamma}^*_n$ has chosen.
The simulation section shows that this approach delivers good coverage probabilities for a set of reasonable DGPs.
Independence tests for two discrete (or even nominal) random variables are readily available in the literature and have been in scientific use for decades. Notable examples are the Chi-square test of independence and the G-Test mcdonald2014handbook. However, if we consider a nominal variable together with a continuous or discrete-continuous variable, testing options are quite limited. One possibility would be to discretize the continuous part and revert to the G-Test, but the arbitrariness in choosing the width and number of bins is suboptimal. Another option would be to regress the continuous variable on dummy variables that jointly determine the value of the nominal variable and subsequently perform a global F-Test on all the slope coefficients.
If the real-valued variable is discrete-continuous, one could in principle extend this strategy to corner solution models such as Tobit.
Since we know the asymptotic distribution of our correlation coefficient under independence, we can add a further option based on this coefficient. Unlike as in the construction of confidence intervals, Assumption (ref) is clearly inappropriate. Therefore, the asymptotic distribution is more complicated.\footnote{For an analytical expression of its pdf, see Corollary 4 in arellano2008exact. Lemma 4 in holzmann2024lancaster also gives the cdf, albeit only in the bivariate case.}
In order to exploit this lemma for an independence test, we need to be able to compute p-values via the cdf of the limiting distribution. Since for arbitrary $\mathbb{R}$-valued random variables $W_1, ..., W_n$, it holds that
this can easily be done via the results obtained in Proposition (ref) and an R package that is able to evaluate the CDF of a multivariate normal distribution, for example mvtnorm mvtnorm\footnote{If $a$ and $b$ are large (see Appendix (ref) for precise values), a very high-dimensional normal CDF has to be evaluated. At least in case 2, it is then advisable to use a computationally less challenging independence test, of which plenty exist.}. However, we need to estimate the matrices $A^1$ and $\Sigma_U^1$ (case 1) or $A^2$ and $\Sigma_U^2$ (case 2). For the $A$-matrices, we just replace the several $\tau$ and $\nu$ by their empirical counterparts and possibly reduce the size of the matrix if not all values with positive probability are present in the sample. For the $\Sigma_U$-matrices, we estimate $\sigma_{v(i)w(j)}$ with $v, w\in\{\tau, \nu\}$ and $i,j\in \{1, ..., k!\}$ or $i,j\in \{1, ..., k!\cdot l!\}$ in the spirit of pohle2024inference, where $k,l$ are defined as in Definition (ref).
As a means to evaluate the finite sample performance of the introduced confidence intervals and the independence test, this section presents simulated coverage rates as well as size and power values for a collection of DGPs. Additionally, it shows simulated mean bias values to showcase the finite sample performance of the proposed estimator. Each simulation contains $MC=1,000$ pairs of iid observations, features several degrees of dependence and different sample sizes.
The first two DGPs are inspired by classical linear regression in which nominal explanatory variables can be accounted for by including binary indicators. $X$ is defined via $\mathbb{P}(X=A)=\mathbb{P}(X=B)=\mathbb{P}(X=C)=1/3$ and $Y$ via
The second pair of DGPs is based on the multinomial logit model with $X\sim \mathcal{N}(0,1), t(1)$ and
For both DGPs, $\alpha = 0$ leads to independent $X$ and $Y$ and the larger $\alpha$ gets in absolute value, the stronger the dependence\footnote{In Section (ref), I admit that I do not have any concept of dependence ordering in the nominal case. Therefore, the word “stronger” shall be understood in a heuristic sense.}.
Table (ref) shows two DGPs in which both variables have a nominal scale. The first has one uniform and one skewed marginal distribution while the second features two uniform distributions. Again, $\alpha = 0$ corresponds to independence and the farther $\alpha$ is away from 0, the stronger the dependence\textsuperscript{(ref)}.
This subsection presents the simulation results for the confidence intervals and the estimator by means of empirical coverage rate and mean bias graphs and the simulation results for the independence test by means of p-value histograms and power tables.
For the empirical coverage rates, we desire a value close to the nominal value of $90\%$. Although Assumption (ref) is clearly not satisfied for the nominal-continuous DGPs in the independence case $\gamma^*=0$, the coverage rates in Figure (ref) still appear acceptable, especially for stronger dependencies. The nominal-nominal DGPs are deliberately designed such that Assumption (ref) is never satisfied (note that $y_2/y_3$ and $x_1/x_2$ occur with the same marginal and joint probabilities irrespective of the choice of $\alpha$) and this feature is reflected by low coverage probabilities, especially for $n=50$. For larger sample sizes, the empirical coverage rate increases almost to the nominal level. Interestingly, the mean bias plots in Figure (ref) show a similar behavior as the empirical coverage rate plots. The bias is large and positive under independence and decreases both with the sample size and the degree of dependence. That is because the bias arises also mainly due to violations of Assumpotion (ref). Intuitively, the bias increases with the intensity of permutation switching of the estimator across the Monte Carlo replications. If Assumption (ref) was satisfied robustly, i.e.\ in the sense that also in finite samples, the estimator (almost) always chooses the same numbering, there would be no bias.
Regarding the independence test, recall that under the null hypothesis $(\alpha = 0)$, the histograms shall converge to a continuous uniform distribution on $[0,1]$. For $\alpha \ne 0$, we desire a high rejection rate of the null hypothesis, i.e.\ high power.
Figure (ref) displays p-value histograms for all the considered DGPs under independence where the test has been conducted based on $\gamma^*$. They all show convergence towards the desired distribution, albeit at different rates. Figure (ref) in turn shows the results if we change our testing procedure to traditional approaches (global F-Test and $\chi^2$ test\footnote{The G-Test shows worse results than the $\chi^2$ test and is therefore omitted.}). The uniformity of our p-value histograms no longer holds for those DGPs in which the continuous variable follows a Cauchy distribution. This illustrates the major strength of the new testing approach: Due to the rank-based nature of the new measure, it does not rely on any moment conditions.
Comparing the power values of the different tests in Table (ref), we find major advantages of the $\gamma^*$ independence test in the two Cauchy cases. However, there are also large differences in one DGP for which no assumption that the traditional tests make is violated. The DGP defined in Table (ref) delivers a remarkable increase in power if we switch from the $\chi^2$ independence test to the $\gamma^*$ independence test. The reason for that may be the slow convergence of the $\chi^2$ test statistic to its limiting distribution if some cells in the contingency table have very small probabilities bruce2015. For the remaining DGPs, the traditional tests have usually more power.
This section illustrates the two major innovations of this paper in two data examples.
Firstly, the proposed coefficient is the first coefficient which is able to summarize the dependence within a bivariate random vector if one marginal distribution is continuously distributed and the other has a nominal scaling. A self-evident example for such a situation is the fundamental observation of economics: People in some countries have higher incomes than people in others.
Figure (ref) displays the dependence between the variables country and income measured by the new coefficient. Each colored point is located at a border triangle and represents the value of the coefficient for those three countries. That is, the variable country always takes exactly three values and the income data is chosen accordingly. The same information with the corresponding confidence intervals is displayed in Figure (ref), where the x-axis shows the border triangle (see the country abbreviations in Table (ref)). It is clearly visible that border triangles featuring countries like France or Poland that have large sample sizes exhibit narrower confidence intervals than other border triangles, for example the one between Hungary, Austria and Slovenia.
With respect to the sizes of the coefficients, those comparisons in which two countries with a similar mean income (see Table (ref)) such as France and Germany are a part of exhibit rather small values. In contrast to that, the Republic of Serbia, Hungary and Romania have a clear income order which is reflected by a larger coefficient. However, other comparisons like Slovakia, the Czech Republic and Austria also have an unambiguous income order in terms of the mean but still yield a small coefficient estimate. The reason for this behavior is of course that even heavily differing means can be explained by few outliers, a differing number of zeros and other factors that have little influence on a rank-based measure like $\widehat{\gamma}^*$.
\nocite{LIS2025}
Secondly, the proposed coefficient satisfies the property of attainability. As described in Section (ref), the traditional measures only satisfy this property if the marginal distributions have a certain structure. In the balanced setting of a $3\times 3$ contingency table, this amounts to the requirement that the marginals are identical in terms of the probability mass. That means that the values of the variables can differ but for each value of the first variable there needs to exist a value of the second variable with the same marginal probability mass.
An example for a comparison in which this requirement is particularly heavily violated is the one between the variables country and religion. Population data on those variables is available in the world religion database WRD2025. Since we only consider the three monotheistic world religions Christianity, Islam and Judaism and perform the analysis for each border triangle separately, we are indeed in the setting of $3\times 3$ contingency tables. However, the marginals in those tables are typically very different. For example, European countries usually have Christian population majorities and very little Muslims and even fewer Jews. As a consequence, the marginal distribution of the variable religion is skewed towards Christianity but the marginal distribution of the variable country may be more balanced.
The differing assessments of dependence between the new coefficient and the old coefficients, where Cramér's V serves as a representative for the traditional measures, are now nicely illustrated in Figure (ref). The maps for the remaining measures are located in Appendix (ref) (Figure (ref)), but display qualitatively similar correlation patterns to the one for Cramér's V. Two observations are immediate: On the one hand, the new coefficient generally yields larger values. On the other hand, the difference varies heavily across the comparisons. One particularly good example for the strength of the new coefficient are the border triangles Germany, Austria, Switzerland and Germany, Poland, Czechia. Cramér's V indicates a rather weak dependence for both triangles, although the value for the latter triangle is a little larger. The newly proposed coefficient however yields almost the same value as Cramér's V for the German-speaking triangle and a value close to one for the other triangle. Indeed, the dependence in both comparisons is very different. Germany, Austria and Switzerland have very similar social structures and religious groups of similar (relative) size. Thus, the dependence between country and religion is weak in this example. But in contrast to Germany, Poland and Czechia have enforced very restrictive migration policies in the past, which results in small religious minorities, especially on the Muslim side. Because both countries also barely have Jews, the Muslims and the Jews are in this example almost fully concentrated in Germany. The reason why Cramér's V yields such a small value is therefore not the weak dependence but the extremely skewed marginal distribution of the variable religion which is not matched by an equally skewed marginal distribution of the variable country. An additional reason is the structurally different definition of perfect dependence that I defend in this paper which can lead to large differences between traditional measures and $\gamma^*$ even if the marginal probabilities were such that the traditional measures are attainable (see Table (ref)).
Another nice example is the border triangle Bulgaria, Greece and Turkey. While Bulgaria and Greece have Christian majorities, Turkey is a primarily Muslim country. This situation of an obvious statistical dependence which consists of large religious blocks concentrated in different countries is nicely captured by both coefficients.
This paper has proposed a new concept of perfect dependence between two random variables of which at least one has a nominal scale that is attainable irrespective of the involved marginal distributions. In addition to that, it introduces the notion of proper dependence measures to the nominal case by defining a set of desirable properties which each such measure has to fulfill. All existing dependence measures fail to satisfy some of the axioms, especially since the existence axiom requires the ability to deal with a nominal-continuous combination of marginal distributions and the attainability axiom requires compliance with the previously defined concept of perfect dependence. As a result, new measures are necessary which are also proposed in this paper. For one of those, I develop asymptotic theory that can subsequently be exploited for confidence intervals and independence testing. Simulations show good finite sample performance for both methods of statistical inference and two applications illustrate the superiority of the new coefficient relative to existing ones.
Future research could continue the quest for a proper dependence measure also in the case in which only one nominal random variable is present. Also, it would be interesting to extend the notion of margin-freeness to nominal random variables and include this feature into the set of desirable properties geenens2020copula.