Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
75,666 characters · 17 sections · 33 citation commands
Comparing latent inequality with ordinal data
Our results help compare latent distributions when only ordinal data are available. For example, we want to learn about the latent health distribution, but individuals only report ordinal categories like “poor” or “good,” which each include a range of latent values between the corresponding thresholds. Other ordinal examples include mental health, bond ratings, political indices, happiness, consumer confidence, and public school ratings.
We consider two types of inequality: within-group and between-group. “Within-group inequality” means dispersion, often quantified by interquantile ranges. For example, to study whether income inequality in the U.S.\ has increased over time, a common dispersion measure is the 90--10 interquantile range, meaning the income distribution's $90$th percentile minus its $10$th percentile. Interquantile ranges can similarly provide evidence of polarization in contexts like politics. “Between-group inequality” means whether one group is better off or worse off than another. For example, “racial inequality” in health means one racial group tends to have better health than another.
We show how certain pairs of ordinal distributions provide evidence of either latent between-group inequality or differences in latent within-group inequality. As in most ordered choice models, we assume the ordinal variable's value depends on the latent variable's value relative to a set of thresholds. For within-group inequality: if the ordinal CDFs cross, then certain latent interquantile ranges are larger for one distribution, even if one group's thresholds are shifted by a constant from the other group's thresholds, providing some evidence of larger dispersion. For between-group inequality: assuming each threshold for the second group is weakly below the corresponding threshold for the first group, the first group having a lower ordinal CDF implies certain quantiles are higher in its latent distribution. This can be interpreted in terms of latent “restricted stochastic dominance” in the sense of Atkinson1987. All our results are robust to arbitrary increasing transformations of the latent random variables because of the equivariance property of quantiles, i.e., the quantile of the transformation equals the transformation of the quantile.
These positive results complement the negative results about comparing latent means from BondLang2019. In their well-published and already well-cited paper, they stress the near impossibility of comparing latent means in our continuous non-parametric framework, essentially because the mean does not share the equivariance property of quantiles. \Citet{BondLang2019} provide conditions in Section II(A) that imply latent first-order stochastic dominance, but they note such conditions never hold in practice. In the same framework, instead of negative results about comparing latent means, we show there is still much that can be compared between two latent distributions, specifically latent quantiles, latent restricted stochastic dominance, and latent interquantile ranges.
Our results can also be interpreted in terms of partial identification. A pair of ordinal CDFs does not uniquely identify a pair of latent CDFs; there are an infinite number of latent CDF pairs (along with thresholds) in the identified set. Still, sometimes all such latent CDF pairs possess a particular property, such as restricted stochastic dominance or certain interquantile ranges being larger. Although we consider two CDFs that separately are not even partially identified due to the unknown thresholds, our within-group inequality approach is similar in spirit to that of Stoye2010, who derives bounds for dispersion parameters of a single distribution based on the “most compressed” and “most dispersed” CDFs in the identified set.
Distinct from the ordinal inequality literature (e.g., as surveyed by Jenkins2019,Jenkins2020 or SilberYalonetzky2021), we allow a continuous latent distribution without imposing a parametric model. A continuous distribution is more realistic than a discrete distribution because it allows latent differences within the same ordinal category: one person with “good” health can be healthier than another, one A-rated bond can have higher credit worthiness than another, one “pretty happy” person may be happier than another, etc. (Our results still apply to discrete or mixed latent distributions, too.) A continuous latent distribution is not considered by most of the proposed ordinal inequality indexes or the median-preserving spread of AllisonFoster2004, which Madden2014 calls “the breakthrough in analyzing inequality with [ordinal] data” (p.\ 206); often no latent interpretation is provided, or else the “latent” distribution merely allows “rescaling” by changing the cardinal value assigned to each category. Similarly for happiness research, BondLang2019 write, “Happiness researchers almost universally assume either that the ordered responses are measured on a discrete interval scale or that each group's latent happiness distribution is normal (i.e., ordered probit) or logistic (i.e., ordered logit)” (p.\ 1634). Such parametric specifications are usually difficult to even think about; for example, what is the shape of the latent happiness distribution? Further, parametric models' results and conclusions are often sensitive to misspecification. For example, BondLang2019 highlight several empirical happiness studies in which “parametric results are reversed using plausible transformations” (p.\ 1629). Still, sometimes imposing a parametric model can allow other assumptions to be relaxed, like allowing random thresholds, which can be valuable and complement our approach.
With covariates, our results can be applied to pointwise comparisons of conditional distributions. This complements the latent median regression model of ChenEtAl2021. For example, if we have three binary covariates, then we can compare the outcome distribution conditional on $(0,0,0)$ to the outcome distribution conditional on $(1,0,0)$, or more generally compare $(0,j,k)$ with $(1,j,k)$. Even with a continuous covariate, we can nonparametrically estimate the ordinal distribution conditional on two distinct covariate values, then use our results to interpret the differences in terms of the latent distributions. Or more simply, the continuous covariate can be discretized, which would also generally increase statistical precision.
For latent between-group inequality inference, we provide an “inner” confidence set for the true set of quantiles at which the first latent distribution is better than the second latent distribution. This confidence set is computed from the data, and it is contained within the true set with at least the nominal probability asymptotically, as used in other economic settings by ArmstrongShen2015 and Kaplan2022consensus. This provides an analogue of a lower confidence bound. That is, there is strong empirical evidence that all quantiles in the confidence set are indeed in the true set, and most likely the true set is even larger.
Our empirical examples show cases in which our results indicate evidence of latent inequality, even though latent means cannot be compared non-parametrically and latent medians are not informative. We interpret estimated differences as well as both frequentist and Bayesian inference.
(ref) contains identification results. (ref) describes the inner confidence set for latent quantile comparison. (ref) provides empirical illustrations. (ref) collects proofs. (ref) describes frequentist and Bayesian inference.
Acronyms used include those for confidence set (CS), cumulative distribution function (CDF), first-order stochastic dominance (SD1), refined moment selection (RMS), and single crossing (SC).
We state and discuss our assumptions in (ref), followed by within-group inequality results in (ref) and between-group inequality results in (ref).
(ref) describes the formal setting and notation. As in many ordered choice models, the ordinal random variables $X$ and $Y$ are derived from corresponding latent random variables $X^*$ and $Y^*$ using non-random thresholds. For example, if $X^*$ is an individual's true latent health, then they report “poor” health (coded $X=1$, meaning the first category) if $X^*\le\gamma_1$, “fair” (coded $X=2$, meaning the second category) if $\gamma_1<X^*\le\gamma_2$, “good” ($X=3$) if $\gamma_2<X^*\le\gamma_3$, “very good” ($X=4$) if $\gamma_3<X^*\le\gamma_4$, and the highest category “excellent” ($X=5=J$) if $\gamma_4<X^*$.
(ref) visualizes (ref), taking $\Delta_{\gamma,1}=\Delta_{\gamma,2}$ as in the below (ref). It shows how the ordinal CDF values are particular points on the latent CDFs. Only $F_X(1)$, $F_X(2)$, $F_Y(1)$, and $F_Y(2)$ are observable; $\gamma_1$, $\gamma_2$, and $\Delta_\gamma$ are unknown. Continuing the health example, $F_X(1)$ is the first population's proportion reporting “poor” health, which is done when $X^*\le\gamma_1$; $F_Y(1)$ is the second population's proportion reporting “poor” health, but they report “poor” when $Y^*\le\gamma_1+\Delta_\gamma$.
To learn about the latent distributions, there must be some restriction on the threshold differences $\Delta_{\gamma,j}$. Otherwise, it is impossible to distinguish whether a difference in the ordinal distribution is due to a change in the latent distribution or a change in the thresholds. For example, in (ref), there is a larger proportion of $Y$ than $X$ reporting poor health, even though there is a smaller proportion of $Y^*\le\gamma_1$ than $X^*\le\gamma_1$, because of the threshold shift $\Delta_\gamma>0$.
(ref) allows the thresholds to differ by any amount as long as it is the same amount for each category. For example, the $Y$ population may have higher thresholds for different health categories as long as they are all higher by the same amount. (ref) is sufficient for our results on within-group inequality.
(ref) allows the thresholds to differ arbitrarily as long as each $Y$ threshold is weakly lower than the corresponding $X$ threshold. This is sufficient for our results on between-group inequality results for evidence that $X^*$ is better than $Y^*$ because all threshold differences work in the opposite direction, making $Y$ appear better than it would if all $\Delta_{\gamma,j}=0$.
As with most assumptions, the threshold restrictions in (ref) may be reasonable in some applications but not others. For example, comparing self-reported ordinal health to the objective McMaster Health Utility Index Mark 3, LindeboomvanDoorslaer2004 find evidence of a mix of “homogeneous reporting,” “index shift,” and “cut-point shift,” depending on the comparison groups. In terms of our (ref): “homogeneous reporting” means all $\Delta_{\gamma,j}=0$, so both (ref) are satisfied; “index shift” means there is a common non-zero $\Delta_\gamma$, so (ref) is satisfied and (ref) may be satisfied if $\Delta_\gamma<0$; and cut-point shift means (ref) is violated, but possibly (ref) holds. Specifically, some differences across age and sex appear to violate our (ref), but “for language, income and education, we find very few violations of the homogeneous reporting hypothesis, and in the few cases where it is violated, this appears almost invariably due to index rather than cut-point shift” (p.\ 1096). Their analysis is done conditional on a few binary variables; for example, their Table 2 tests for differences across income within the eight subgroups defined by male/female, young/old, and high/low education. Their statistical testing assumes latent normality (pp.\ 1090--1091), but misspecification would tend to make their test more likely to falsely reject the null of homogeneous reporting or index shift.
More generally, (ref) seem especially plausible when comparing the same group to itself in a different time period, as in our empirical examples in (ref). Especially if the time periods are not far apart, we may expect all $\Delta_{\gamma,j}\approx0$, so both (ref) are reasonable.
Further, (ref) is often expected when the distribution of $X$ is better than that of $Y$. For example, if $Y$ and $X$ are older and younger adults' health, respectively, LindeboomvanDoorslaer2004 find that (ref) holds: “given similar objective health limitations\ldots older adults are somewhat more inclined to self-report good health” (p.\ 1096). That is, because older adults are generally in worse health, they lower their thresholds accordingly. This seems natural in many situations: if $Y^*$ values tend to be lower than $X^*$, then individuals may report $Y$ based on lower thresholds than $X$. Although this satisfies (ref), such lower thresholds are still not helpful because they partially offset the latent changes. That is, our findings are conservative in the sense of not detecting as large of a difference between $X^*$ and $Y^*$ as we would have with identical thresholds $\Delta_{\gamma,j}=0$, but we are protected against spurious findings.
For within-group inequality, we characterize certain pairs of ordinal CDFs that imply one latent distribution has greater dispersion than another, as quantified by latent interquantile ranges. Throughout, we maintain (ref).
(ref) illustrates the intuition for how an ordinal CDF crossing implies a relationship between certain latent interquantile ranges. The black line shows the latent CDF of $X^*$, with the black squares showing $F_X^*(\gamma_1)=F_X(1)$ and $F_X^*(\gamma_2)=F_X(2)$. The green line shows the latent CDF of $Y^*$, with the green triangles showing $F_Y^*(\gamma_1)=F_Y(1)$ and $F_Y^*(\gamma_2)=F_Y(2)$, with $\Delta_\gamma=0$ for simplicity. If instead $\Delta_\gamma\ne0$, then the difference between thresholds still remains $(\gamma_2+\Delta_\gamma)-(\gamma_1+\Delta_\gamma)=\gamma_2-\gamma_1$, so the following intuition is unchanged.
In (ref), the difference between the $F_Y(2)$-quantile and $F_Y(1)$-quantile of $Y^*$ is
Because $F_X(1)<F_Y(1)$ and $F_X(2)>F_Y(2)$, the corresponding interquantile range for $X^*$ must be smaller. That is, the black-with-squares line $F_X^*(\cdot)$ reaches the value $F_Y(1)$ to the right of $\gamma_1$, but reaches $F_Y(2)$ to the left of $\gamma_2$, so
Further, consider any quantile indices $\tau_1$ and $\tau_2$ satisfying $F_X(1)<\tau_1\le F_Y(1)$ and $F_Y(2)<\tau_2\le F_X(2)$. As seen in (ref), the $Y^*$ interquantile range is larger than $\gamma_2-\gamma_1$ whereas the $X^*$ interquantile range is smaller: \[ Q_Y^*(\tau_2)-Q_Y^*(\tau_1) >\gamma_2-\gamma_1 >Q_X^*(\tau_2)-Q_X^*(\tau_1) . \]
(ref) formalize and generalize these arguments.
(ref) can be applied multiple times if the ordinal CDFs cross multiple times. For example, if $F_X(1)<F_Y(1)$, $F_X(2)>F_Y(2)$, and $F_X(3)<F_Y(3)$, then it can be applied with $(j,k)$ as $(1,2)$ and $(2,3)$. This indicates evidence of larger dispersion of $X^*$ in one part of the distribution, but larger dispersion of $Y^*$ in another. In contrast, if the ordinal CDFs have only a single crossing, then there is evidence of only one latent distribution having larger dispersion, as in (ref).
(ref) interprets the median-preserving spread of AllisonFoster2004 in terms of continuous latent distributions. Analogous to a mean-preserving spread, a median-preserving spread says two ordinal distributions share the same median category but one is a “spread out” version of the other, i.e., can be constructed by moving probability mass away from the median. \footnote{There is a typo in (3c) of their definition: $X_k\ge Y_k$ should be $X_k\le Y_k$.} Because the median-preserving spread implies a single ordinal CDF crossing, our (ref) applies. When the median differs, the median-preserving spread does not apply, but (ref) still applies if there is a single ordinal CDF crossing.
Within-group inequality with ordinal data is a topic of ongoing interest. Of the 382 Google Scholar citations of AllisonFoster2004, over 100 have come since 2018, and their approach is only one among many for assessing inequality with ordinal data. Recent empirical papers assessing within-group inequality from ordinal data include the study of durable good consumption in Asian countries by DeutschEtAl2020, “Health polarization and inequalities across Europe: an empirical approach” by PascualEtAl2018, the study of time trends in U.S.\ happiness inequality by DuttaFoster2013 and StevensonWolfers2008a, and the cross-country comparisons of subjective well-being and education by BalestraRuiz2015.
For between-group inequality (better/worse), we compare latent quantiles. \Citet{BondLang2019} show that comparisons by latent means or latent first-order stochastic dominance are essentially impossible; quantiles provide a tractable alternative that still provide evidence of between-group inequality. Throughout, we maintain (ref).
Having a larger latent $\tau$-quantile is one piece of evidence of being “better.” Besides the natural intuition, this can also be interpreted in terms of quantile utility maximization (e.g., Manski1988; Rostek2010; deCastroGalvao2019). For example, given strictly increasing utility function $u(\cdot)$, a $\tau$-quantile utility maximizer strictly prefers $X^*$ over $Y^*$ if the $\tau$-quantile of $u(X^*)$ is strictly greater than the $\tau$-quantile of $u(Y^*)$, which is true if and only if $Q_X^*(\tau)>Q_Y^*(\tau)$. That is, learning $Q_X^*(\tau)>Q_Y^*(\tau)$ is equivalent to learning that $X^*$ is preferred by all $\tau$-quantile utility maximizers, regardless of utility function. Of course, if $X^*$ is preferred for some $\tau$ but $Y^*$ is preferred for other $\tau$, then people may reasonably disagree about which is better.
The quantile intuition extends the known conclusion that latent medians can be compared if all $\Delta_{\gamma,j}=0$ and the median category differs. That is, if the median category of ordinal $X$ is above that of $Y$, then the latent median of $X^*$ is above that of $Y^*$. Specifically, if $\gamma_j$ is the threshold between the two categories, then the median of $Y^*$ is below $\gamma_j$ whereas the median of $X^*$ is above. More generally, $\Delta_{\gamma,j}\le0$ is sufficient: if the $Y$ thresholds are even lower than the $X$ thresholds, then the latent median of $Y^*$ is even lower than it appears. That is, the median of $Y^*$ is now below $\gamma_j+\Delta_{\gamma,j}$, which in turn is below $\gamma_j$, which is below the median of $X^*$.
(ref) formally states the result for quantiles.
Besides using our identification results to interpret estimated ordinal distributions, we propose inner confidence sets for the latent sets defined in our results. Specifically, interest is in the true set $\mathcal{T}_X$ in (ref), and in the set $\mathcal{T}_1\times\mathcal{T}_2\equiv\{(\tau_1,\tau_2):\tau_1\in\mathcal{T}_1,\tau_2\in\mathcal{T}_2\}$ in (ref) and a generalization of (ref).
We first develop intuition through the special case in (ref). Then we describe the general between-group inequality method with formal theoretical results in (ref). Similarly, (ref) have general methods and theoretical results for within-group inequality.
An inner confidence set (CS) $\hat{\mathcal{S}}$ should be contained within the true set $\mathcal{S}$ with high asymptotic probability:
This general idea is used in other economic settings by ArmstrongShen2015 and Kaplan2022consensus.
From our identification results, the sets we define are in turn subsets of the ultimate sets of interest. Specifically, $\mathcal{T}_X$ in (ref) is a subset of $\{\tau : Q_X^*(\tau)>Q_Y^*(\tau)\}$. Similarly, $\mathcal{T}_1\times\mathcal{T}_2$ is a subset of the set of all $(\tau_1,\tau_2)$ for which $Q_X^*(\tau_2)-Q_X^*(\tau_1)<Q_Y^*(\tau_2)-Q_Y^*(\tau_1)$. Consequently, our inner CS for $\mathcal{T}_X$ or $\mathcal{T}_1\times\mathcal{T}_2$ is also valid for the larger full set. In contrast, an outer CS for $\mathcal{T}_X$ or $\mathcal{T}_1\times\mathcal{T}_2$ is generally not valid for the corresponding larger set, hence our focus on an inner CS.
To develop intuition, consider $J=2$, so the unknown ordinal parameters are $F_X(1)$ and $F_Y(1)$. The true set includes all $\tau$ values above $F_X(1)$ and below $F_Y(1)$:
This set is of interest because under the conditions of (ref), the latent $\tau$-quantile of $X^*$ is higher than that of $Y^*$ for all $\tau\in\mathcal{T}_X$.
The goal of the inner CS $\hat{\mathcal{T}}_X$ is to be contained within the true $\mathcal{T}_X$ with asymptotic probability at least $1-\alpha$: $\operatorname{P}(\hat{\mathcal{T}}_X \subseteq \mathcal{T}_X)\ge1-\alpha+o(1)$. That is, with high probability, the true set is even larger than the inner CS.
Consider the inner CS
where $ \hat{C}^U_{X(1)}$ is an upper confidence limit for $F_X(1)$, and $ \hat{C}^L_{Y(1)}$ is a lower confidence limit for $F_Y(1)$, both with confidence level $\sqrt{1-\alpha}$. That is,
For example, if $\sqrt{n}(\hat{F}_X(1)-F_X(1)) \xrightarrow{d} \mathrm{N}(0,\sigma^2)$, and $\hat\sigma^2\xrightarrow{p}\sigma^2$, then $\hat{C}^U_{X(1)}=\hat{F}_X(1)+z_{\sqrt{1-\alpha}}\hat\sigma/\sqrt{n}$ satisfies (ref), where $z_p$ denotes the $p$-quantile of the standard normal distribution.
This inner CS $\hat{\mathcal{T}}_X$ satisfies (ref) with asymptotic confidence level $1-\alpha$. The argument has two steps. First, event $\hat{\mathcal{T}}_X\subseteq \mathcal{T}_X$ is implied by the combination of events $\hat{C}^U_{X(1)}\ge F_X(1)$ and $\hat{C}^L_{Y(1)}\le F_Y(1)$; that is, “coverage” of the inner CS is implied by coverage of both confidence limits. Thus, if the latter combination of events occurs with at least $1-\alpha$ asymptotic probability, then so does the former event. Second, if the two data samples are independent, then the joint coverage probability of the two confidence limits equals the product of the marginal coverage probabilities. Formally,
More generally, we want to learn about the full set $\mathcal{T}_X$ from (ref). Specifically, as in (ref), we want a procedure to compute inner CS $\hat{\mathcal{T}}_X$ from data such that it is contained within the true $\mathcal{T}_X$ with asymptotic probability at least $1-\alpha$: $\operatorname{P}(\hat{\mathcal{T}}_X\subseteq\mathcal{T}_X)\ge1-\alpha+o(1)$.
Extending (ref), we propose the general inner CS
where the $ \hat{C}^U_{X(j)}$ are joint upper confidence limits for the $F_X(j)$, and the $ \hat{C}^L_{Y(j)}$ are joint lower confidence limits for the $F_Y(j)$, both with confidence level $\sqrt{1-\alpha}$. That is,
Theoretically, the justification follows the same two-step argument from (ref). First, event $\hat{\mathcal{T}}_X\subseteq \mathcal{T}_X$ is implied by all $2(J-1)$ confidence limits containing the respective true values, so the probability of the latter event is a lower bound for $\operatorname{P}(\hat{\mathcal{T}}_X\subseteq \mathcal{T}_X)$. Second, if the two data samples are independent, then using (ref), all $2(J-1)$ confidence limits jointly cover with probability $(\sqrt{1-\alpha}+o(1))^2=1-\alpha+o(1)$. This is formalized in the proof of (ref).
To implement this idea, we choose the individual confidence limits to share the same pointwise coverage probability. That is, for some $\tilde\alpha$ and all $j=1,\ldots,J-1$,
In principle, the pointwise coverage probability could instead be distributed differently to reflect prior interests or beliefs, but that would both complicate the implementation and introduce opportunity for manipulation, which our implementation avoids.
Our formal results use the following assumptions.
(ref) is a high-level assumption that can hold across a variety of settings, including certain clustered sampling and time series settings. To give an example of a primitive assumption, (ref) states the sufficiency of iid sampling for (ref).
(ref) describes how to construct our inner CS. It uses the following definitions and distributions. Define minimum and maximum $t$-statistics as
Given (ref), each $\hat{t}_X(j)$ or $\hat{t}_Y(j)$ is asymptotically standard normal. Further, the vector $(\hat{t}_X(1),\ldots,\hat{t}_X(J-1))'$ is asymptotically normal with mean zero and covariance matrix equal to the correlation matrix of $\boldsymbol{\mathbf{W}}$ from (ref); the distribution of the minimum of this normal vector is thus the asymptotic distribution of $\hat{T}_X$. Similarly, the vector $(\hat{t}_Y(1),\ldots,\hat{t}_Y(J-1))'$ is asymptotically normal with mean zero and covariance matrix equal to the correlation matrix of $\boldsymbol{\mathbf{M}}$ from (ref), and the distribution of its maximum is the asymptotic distribution of $\hat{T}_Y$.
The following technical details show the above more formally; many readers may wish to skip them. Define diagonal matrices $\mkern3mu\underline{\mkern-3mu \boldsymbol{\mathbf{D}}\mkern-3mu}\mkern3mu_X$, $\mkern3mu\underline{\mkern-3mu \boldsymbol{\mathbf{D}}\mkern-3mu}\mkern3mu_Y$, $\hat{\mkern3mu\underline{\mkern-3mu \boldsymbol{\mathbf{D}}\mkern-3mu}\mkern3mu}_X$, and $\hat{\mkern3mu\underline{\mkern-3mu \boldsymbol{\mathbf{D}}\mkern-3mu}\mkern3mu}_Y$ with row $j$, column $j$ entries
so
By the continuous mapping theorem, $\hat{T}_X\xrightarrow{d}\min(\mkern3mu\underline{\mkern-3mu \boldsymbol{\mathbf{D}}\mkern-3mu}\mkern3mu_X\boldsymbol{\mathbf{W}})$ and $\hat{T}_Y\xrightarrow{d}\max(\mkern3mu\underline{\mkern-3mu \boldsymbol{\mathbf{D}}\mkern-3mu}\mkern3mu_Y\boldsymbol{\mathbf{M}})$. Applying the continuous mapping theorem again, with $\Phi(\cdot)$ the standard normal CDF,
In practice, the critical values $\tilde{\alpha}$ and $\tilde{\beta}$ in (ref) can be simulated using (ref), by taking many random draws from the asymptotic distributions and computing the corresponding quantiles.
(ref) formally states the asymptotic validity of the inner CS in (ref).
We first consider a confidence set corresponding to (ref) for a predetermined pair of categories $j<k$. The population object of interest is the set $\mathcal{T}_{Xj} \times \mathcal{T}_{Yk}$, using notation from (ref) and similarly defining $\mathcal{T}_{Yk}$ as the interval $(F_Y(k),F_X(k)]$ if $F_Y(k)<F_X(k)$ (and the empty set otherwise). Given the assumptions of (ref), $Q_X^*(\tau_2)-Q_X^*(\tau_1) < Q_Y^*(\tau_2)-Q_Y^*(\tau_1)$ for all $(\tau_1,\tau_2)\in\mathcal{T}_{Xj}\times\mathcal{T}_{Yk}$.
This is almost the same as in (ref), except the $X$ and $Y$ are switched at category $k$, so the confidence limits should be switched, too. To be rigorous, we give the details here. Let
with the $t$-statistics defined in (ref). Analogous to (ref), given (ref),
the quantiles of which can be simulated like before.
To consider all possible category pairs, the details are more complicated, but the intuition remains the same. The strategy again is to first construct joint confidence limits for the ordinal CDFs, although this time they are two-sided. Then, among all pairs of ordinal CDFs within the confidence limits, we find the pair that generates the smallest set of $(\tau_1,\tau_2)$ with corresponding interquantile range smaller for $X^*$ than for $Y^*$. This set is contained within the corresponding set for any other pair within the confidence limits, which includes the true pair of ordinal CDFs with probability $1-\alpha$; thus, this set is an inner confidence set. This approach is valid but may be conservative in some cases; it would be valuable to refine its precision in future work.
To be explicit, we construct an inner CS for the population set
where for each $j=1,\ldots,J-1$,
If there is a single crossing, then this $\mathcal{T}$ matches (ref)'s $\mathcal{T}_1\times\mathcal{T}_2$, but this $\mathcal{T}$ is well-defined even without a single crossing. Like before, the goal is to construct an inner CS $\hat{\mathcal{T}}$ such that $\operatorname{P}(\hat{\mathcal{T}}\subseteq\mathcal{T})\ge1-\alpha+o(1)$.
To derive two-sided confidence limits, consider the maximum absolute $t$-statistics,
where $\hat{t}_X(j)$ and $\hat{t}_Y(j)$ are from (ref). Analogous to (ref), given (ref),
the quantiles of which can be simulated like before.
The following empirical examples illustrate our theoretical results, showing how our results help interpret ordinal data in terms of latent relationships. All empirical results can be replicated with the files provided on the first author's website. \footnote{\url{https://kaplandm.github.io}} Code is in R (R.core), with help from package \lstinline{quadprog} \Citep{R.quadprog} for our RMS implementation. For between-group inequality, we assume $\Delta_\gamma=0$ throughout.
For estimation, we use provided sampling weights to compute appropriately weighted empirical ordinal CDFs, and then we interpret the differences in terms of latent quantiles using (ref).
For inference, we use sampling weights when provided but (though not ideal) otherwise treat sampling as iid. We use $\alpha=0.05$ unless otherwise noted.
For inference on ordinal first-order stochastic dominance (SD1), we use both frequentist and Bayesian methods. The Bayesian results use the posterior probability of ordinal SD1 from a Dirichlet--multinomial model with uniform prior. For the null of SD1, we use the refined moment selection (RMS) test of AndrewsBarwick2012 as described in (ref). For the null of non-SD1, we use the intersection--union test as described in (ref). Given some level like $\alpha=0.05$, we say there is “statistically significant” evidence in favor of $X\mathrel{\mathrm{SD}_{1}}Y$ if the posterior probability is above $1-\alpha$ (Bayesian) or the intersection--union test rejects $H_0\colon X\mathrel{\mathit{non}\mathrm{SD}_{1}}Y$ at level $\alpha$ (frequentist), and evidence against $X\mathrel{\mathrm{SD}_{1}}Y$ if the posterior probability is below $\alpha$ (Bayesian) or RMS rejects $H_0\colon X\mathrel{\mathrm{SD}_{1}}Y$ at level $\alpha$ (frequentist).
We study ordinal measures of mental health and general health from the popular NHIS data, available through IPUMS \Citep{IPUMSNHIS2019}. We use the 2006 and 2008 waves because the latter is in the Great Recession while the former is not. Besides the health variables described below, we also use the appropriate sampling weights (\lstinline{PERWEIGHT} for studying general health; \lstinline{SAMPWEIGHT} for mental health because it is from the supplemental survey) and the provided measures of poverty (\lstinline{POORYN}), education (\lstinline{EDUC}), sex (\lstinline{SEX}), and race (\lstinline{RACENEW}).
For mental health, we compare non-recession and recession distributions separately for men in poverty and men not in poverty. The goal is to assess whether the Great Recession is associated with worse mental health, and whether the association is stronger for those in poverty. Although we cannot determine the 2006 poverty status of individuals observed in 2008 because these are repeated cross-sections rather than panel data, the proportion in poverty is very similar in both years ($12.0\%$, $11.7\%$). More potentially problematic is the significant proportions of the samples ($23\%$, $11\%$) for which the poverty measure is unavailable; we do not attempt to assess possible sample selection bias.
The mental health variable is based on the Kessler-6 scale for nonspecific psychological distress introduced by KesslerEtAl2002, which is the sum of variables \lstinline{AEFFORT}, \lstinline{AHOPELESS}, \lstinline{ANERVOUS}, \lstinline{ARESTLESS}, \lstinline{ASAD}, and \lstinline{AWORTHLESS} that measure frequencies of various feelings over the past $30$ days. We code the worst mental health as 1 and the best as 25, i.e., we subtract the raw K6 score from 25. As a rough guide, values of 1--12 help predict serious mental illness KesslerEtAl2003.
Our interpretations from (ref) in terms of latent quantiles are more helpful than considering the latent mean or median. The latent means cannot be compared non-parametrically, as noted by BondLang2019, and the median is always a very high value indicating good mental health.
(ref) shows evidence of worse mental health during the Great Recession for men in poverty, which can be interpreted by applying (ref) to the estimated CDFs. The graphs show the ordinal weighted empirical CDFs for mental health. For men in poverty ((ref)), the CDF for 2006 generally lies below that for 2008, indicating worse mental health in 2008 during the Great Recession. By (ref), the latent mental health distribution during the recession is estimated to be worse over a broad set of quantile indices $\mathcal{T}_X$ that includes $[0.03,0.30]$ as well as $[0.33,0.42]$ and other values above $0.42$ and below $0.03$. As usual, we cannot hope to learn about changes within the lowest category, but in this case the lowest category is a small fraction of the population, well less than $1\%$. Overall, the 2008 latent mental health distribution is estimated to be worse than 2006 over most of the lower part of the distribution.
(ref) also shows the estimated 2006 and 2008 mental health ordinal CDFs for men not in poverty ((ref)). These appear nearly identical for most of the distribution, especially the lower half. The only categories at which the estimated CDFs differ by at least $0.01$ are $22\le j\le 24$, and the only category where $\hat{F}_X(j)/\hat{F}_Y(j)<0.95$ is $j=23$. That is, the changes are mostly within the part of the distribution corresponding to good mental health.
Because there are many categories ($J=25$) and the sample sizes are not very large for men in poverty ($1082$ in year 2006, $1095$ in 2008), the statistical significance is mixed. Considering all $j=1,2,\ldots,25$, (ref) produces an empty $90\%$ inner CS. However, restricting attention to only certain categories yields a non-empty inner CS. For example, if the categories are combined into groups of five (1--5, 6--10, 11--15, 16--20, 21--25), then $\hat{\mathcal{T}}_X=[0.120, 0.121]$: small but not completely empty, and suggesting the strongest evidence is for a change in the lower part of the distribution. Alternatively, if we combine categories 1--12 (which make up a small fraction of the population) and combine categories 19--25 (which correspond to pretty good mental health), but keep separate $j\in\{13,\ldots,18\}$, then the $90\%$ inner CS is
Again, these ranges are not large, but they reflect reasonably strong evidence of worsening mental health for men in poverty during the Great Recession, specifically in lower quantiles of the distribution.
In all, although we only have ordinal data, we can still see that the mental health costs of the Great Recession are concentrated on those in poverty, especially in lower parts of the mental health distribution. Specifically, (ref) lets us interpret the estimated decline in mental health in terms of a broad range of the latent distribution; accounting for statistical uncertainty and being more conservative, (ref) reports a much smaller set of quantiles given a $90\%$ confidence level.
For general health, we compare 2006 to 2008 for men in poverty, as well as comparing across racial and education groups within 2006. The self-reported general health variable (\lstinline{HEALTH}) has values poor, fair, good, very good, and excellent, which we code as 1, 2, 3, 4, and 5, respectively.
(ref) shows the 2006/2008 comparison in the first two rows, using the sample of men in poverty as in the mental health analysis. The first row shows the 2006 (weighted empirical) CDF evaluated at poor (1), fair (2), good (3), and very good (4); the CDF at excellent always equals $1$ by definition. The second row shows the estimated 2008 CDF, which is higher at $1\le j\le 2$ but lower at $3\le j\le 4$. Thus, even in the sample, there is not SD1 in either direction.
This 2006/2008 comparison of men in poverty also provides some evidence that latent health dispersion (within-group inequality) increased during the Great Recession. In the data, the 2006 ordinal CDF crosses the 2008 ordinal CDF once from below, so (ref) applies with $m=2$. For example, the $70$--$16$ interquantile range is estimated to be larger for the latent 2008 distribution because $F_{2006}(2)<0.16<F_{2008}(2)$ and $F_{2008}(4)<0.70<F_{2006}(4)$. (ref) holds even if $\Delta_\gamma<0$, meaning systematically lower thresholds in 2008, which could explain the larger share reporting the very best health category ($j=5$, “excellent”).
(ref) next compares low and high education groups with the 2006 data, with “high” meaning any post-secondary education. Again, the two estimated CDFs are shown; the high education CDF is below the low education CDF at all $1\le j\le4$. That is, the sample shows ordinal SD1, evidence of between-group inequality. More specifically, (ref) says the latent high-education health distribution is estimated to be better at the $\tau$-quantile for (rounding to nearest $0.01$) \[ \tau\in[0.02,0.03]\cup[0.07,0.12]\cup[0.30,0.39]\cup[0.64,0.67] .\] Letting $X$ represent the low education ordinal health distribution and similarly $Y$ for high education, Bayesian and frequentist methods both reject $X\mathrel{\mathrm{SD}_{1}}Y$. Further, there is positive evidence of $Y\mathrel{\mathrm{SD}_{1}}X$: the posterior probability is above $1-\alpha$, and $Y\mathrel{\mathit{non}\mathrm{SD}_{1}}X$ is rejected by the intersection--union test at level $\alpha$ in favor of $Y\mathrel{\mathrm{SD}_{1}}X$.
(ref) shows similar results for the Black/white comparison as for low/high education. There is SD1 in the sample, and SD1 of Black over white ordinal health is rejected by both frequentist and Bayesian analysis, whereas SD1 of white over Black ordinal health is supported by a posterior probability above $1-\alpha$ and the intersection--union test's rejection of non-SD1 in favor of SD1.
A $90\%$ inner CS can also be computed for each comparison using (ref). Unlike in (ref), the sample sizes are very large, so the inner CS is similar to the point estimates. For the low/high education comparison,
That is, there is strong evidence that the high-education latent health distribution is better than the low-education distribution at these quantiles. It is probably better at other quantiles, too, but we do not have strong enough empirical evidence to say so. For the Black/white comparison,
indicating strong evidence that the white latent health distribution is better at these quantiles (and probably more).
We compare continuous latent distributions non-parametrically when only ordinal data are available. Our identification results interpret certain ordinal patterns as evidence of between-group inequality in terms of quantiles, while other ordinal patterns indicate differences in latent within-group inequality in terms of interquantile ranges. We propose an inner confidence set for the former set of latent quantiles. Empirical examples with different ordinal measures of health show how our results provide insight, even when latent means cannot be compared. Our approach can be applied similarly with ordinal measures from education, politics, finance, and other areas. Our results can also extend to comparison of conditional distributions, which can be estimated by semiparametric or non-parametric “distribution regression” even with continuous regressors Frolich2006.
Special thanks to Longhao Zhuo for his involvement in the early stages of this project, including the contributions now published in KaplanZhuo2021 and the R code for the RMS test. Many thanks to co-editor Petra Todd, three anonymous reviewers, Alyssa Carlson, Denis Chetverikov, Yixiao Jiang, Jia Li, Matt Masten, Arnaud Maurel, Zack Miller, Peter Mueser, Shawn Ni, Adam Rosen, Dongchu Sun, and other participants from UConn, Duke, Yale, the 2018 Midwest Econometrics Group, and the 2019 Chinese Economists Society conference for helpful questions, comments, and references. Thanks also to the Cowles Foundation for their hospitality during part of this work.
The authors have no financial interests nor any other conflicts of interest relevant to the content of this study.
\singlespacing