EconBase
← Back to paper

Distributional Change in Ordinal Data with Missing Observations: Minimal Mobility and Partial Identification

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

69,743 characters · 14 sections · 20 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Distributional Change in Ordinal Data with Missing Observations: Minimal Mobility and Partial Identification

abstractEmpirical analyses of ordinal outcomes using repeated cross-sectional data rely on marginal distributions, leaving the joint distribution unobserved and the sources of distributional change unidentified. This paper develops a framework to measure and interpret such changes under limited information. The $L_1$ distance between cumulative distribution functions admits an optimal transport representation as the minimal reallocation of probability mass across ordered categories, which provides a foundation for the analysis. This yields both a scalar measure of discrepancy and a structured characterization of how distributional change must occur, which I term minimal-mobility configurations. To address missing data, I adopt a partial identification approach that delivers sharp bounds on the marginal distributions and, in turn, on both the discrepancy measure and its associated configurations. The resulting framework supports inference using standard resampling methods and provides a transparent basis for assessing sensitivity to nonresponse. An application to Arab Barometer data illustrates the approach.\\ \\ JEL: C14; C18; C21.\\ Keywords: Optimal transport; ordinal data; distributional change; partial identification; missing data; minimal mobility; Wasserstein distance.

Introduction

Empirical analyses often compare distributions of ordinal variables across groups or over time using repeated cross-sectional data, where only marginal distributions are observed. In such settings, the joint distribution linking these marginals is not identified, making it difficult to assess how observed differences across distributions arise. As a result, standard approaches that rely on tracking individual-level transitions are not available, and distributional comparisons must be based solely on information contained in the marginals.

Most empirical work in this setting focuses on comparisons of marginal distributions, for example through differences in category shares, cumulative distribution functions, or stochastic dominance tests (e.g., Jenkins,Tabri2021). While informative, these comparisons do not address how the observed differences can be reconciled in the absence of joint information. Ideally, assessing how one distribution differs from another requires knowledge of the joint distribution linking them, which describes how probability mass is reallocated across categories. When this joint distribution is not observed, many distinct mechanisms can generate the same observed differences. In this paper, I take a complementary perspective and ask: what can be learned about how one distribution differs from another when the joint distribution is unobserved?

I address this question through an extremal perspective, focusing on the least amount of movement across categories required to reconcile the two distributions. I show that this question can be answered using only marginal information in ordinal settings, yielding both a scalar measure of distributional change and a structured representation of how this change must be realized across categories. This distinction is important: while standard methods detect whether distributions differ, they leave open the mechanisms through which these differences arise. The proposed approach instead characterizes the set of feasible minimal reallocations consistent with the data, thereby providing information on the structure of distributional change even in the absence of joint observations.

To operationalize this approach in empirical settings, one must account for the fact that the marginal distributions themselves may be only partially observed due to missing data, a pervasive feature of many datasets. I adopt a worst-case approach and construct sharp bounds on the feasible marginal distributions that are consistent with the observed data. These bounds define an identified set of distributions, which in turn induces an identified set for the discrepancy measure—defined as the $L_1$ distance between the marginal cumulative distribution functions—as well as for the associated representations of how this discrepancy can be realized. The idea of using bounds to address data problems is not new, but gained popularity with the seminal work of Horowitz-Manski and developed in subsequent work (see, e.g., Molinari2020 for a survey of partial identification).

Within this framework, the discrepancy measure admits a representation as the Wasserstein-1 distance between feasible marginal distributions, which characterizes the minimal cost of reallocating probability mass across ordered categories. While the equivalence between the Wasserstein-1 distance and the $L_1$ distance between cumulative distribution functions on the real line is well known, the contribution of this paper is to use this representation to recover structured descriptions of distributional change. In particular, the focus here is not only on the magnitude of discrepancy, but on how this discrepancy must be realized across ordered categories, and on how these objects behave under partial identification due to missing data.

These structured representations of distributional change can be interpreted as optimal transport couplings between the marginal distributions (e.g., villani2009optimal). Referred to here as minimal-mobility configurations, these couplings specify how probability mass is reassigned across categories and therefore provide a representation of distributional change. While the discrepancy measure summarizes the magnitude of this change, the set of optimal couplings characterizes how this change can be realized under the least aggregate reallocation of probability mass. In this sense, these configurations provide a structured benchmark: they describe the minimal movement required to reconcile the observed distributions and delineate the set of reallocation patterns that are consistent with this benchmark. Accordingly, the minimal-mobility configurations are extremal movement structures, distinct from extremal dependence structures characterized by Fr\'echet bounds.

Importantly, these minimal-mobility configurations should not be interpreted as the true data-generating mechanism, but rather as benchmark reallocations that isolate the minimal amount of movement required by the observed marginals. This representation is particularly useful in empirical applications, as it allows one to distinguish between movements that are necessary to account for the observed differences and those that are possible but not identified. It also provides a natural basis for assessing the sensitivity of conclusions to missing data by comparing the range of minimal-mobility configurations implied by different feasible marginal distributions.

The use of optimal transport in econometrics is not new. A growing literature has employed optimal transport as a tool for identification, estimation, and computation in a range of settings, including incomplete models (e.g., Galichon-Henry-2011,Galichon), data combination problems (e.g., dHaultfoeuilleGaillacMaurel2024), treatment effect analysis (e.g., OptimalTreatmentAssignment2025), measurement error (e.g., SchennachStarck2026), matching (e.g., Galichon-Dupuy-Sun), and discrete choice models (e.g., GalichonSalanie2022). See also the recent survey by GalichonHenry2026 for additional references and applications. Across these settings, optimal transport is primarily used to characterize identified sets, derive sharp bounds, or exploit duality to transform otherwise intractable problems into tractable ones. Particularly closely related to the present paper, DaljordPouliotXiaoHu2026 use optimal transport as a reduced-form device to construct sharp lower bounds on the volume of unobserved black market transactions by quantifying the minimal mass that must be reallocated to reconcile observed distributions, in the spirit of partial identification.

The present paper takes a complementary perspective. While optimal transport has also been used as an object of interpretation in structural settings---most notably in the econometrics of matching models, where it is used to recover primitives such as surplus or preferences from observed matches---these approaches rely on a fully specified joint structure. In contrast, I use optimal transport to provide an interpretable representation of distributional differences when the joint distribution is not observed. The optimal transport formulation delivers a direct link between observed discrepancies and statements about the minimal reallocation of probability mass required to reconcile distributions. This shifts the role of optimal transport from a structural identification device to a tool for organizing and interpreting feasible reallocations under incomplete information, yielding benchmark configurations that describe the structure of distributional change implied by the data.

This paper contributes to the literature on distributional analysis and partial identification by providing a new way to interpret differences between ordinal distributions when only marginal information is available. Rather than treating distributional comparisons as purely descriptive, the approach links observed differences to economically meaningful statements about the minimal reallocation of probability mass required to reconcile distributions. While existing methods based on stochastic dominance and summary statistics detect whether distributions differ, they leave open the mechanisms through which these differences arise. In contrast, the proposed approach characterizes the set of feasible minimal reallocations consistent with the data, thereby recovering information about the structure of distributional change that is not accessible from marginal comparisons alone, while remaining agnostic about the unobserved joint distribution.

By combining this representation of distributional change with a worst-case approach to missing data, the paper delivers inference that is robust to nonresponse while preserving a transparent interpretation of the underlying economic objects. Repeated cross-sectional data are a primary source of information in many empirical settings, particularly where panel data are unavailable, such as surveys conducted in parts of the Middle East and North Africa. In these contexts, the absence of joint information makes it difficult to assess how distributional changes arise or how probability mass is reallocated across categories. The empirical usefulness of the framework is illustrated using data from the Arab Barometer, and inference can be implemented using standard resampling procedures (e.g., HorowitzManski2000,Chernozhukov-Hong-Tamer), making the approach readily applicable in practice without problem-specific derivations. The paper also discusses complementary maximal-mobility benchmarks and relates the framework to classical Fr\'echet inequalities Frechet1935,Frechet1951, emphasizing that the resulting bounds characterize extremal movement across categories rather than extremal dependence structures implied by Fr\'echet bounds.

The remainder of the paper is organized as follows. Section (ref) introduces the framework for measuring distributional change in ordinal settings and presents an illustrative example. Section (ref) develops the identification results, characterizing the identified sets for both the discrepancy measure and the associated minimal-mobility configurations. Section (ref) discusses extensions and interpretation, including maximal-mobility benchmarks and the relation to Fr\'echet inequalities. Section (ref) describes the bootstrap procedure of inference on the objects of interest. Section (ref) presents the empirical illustration using data from the Arab Barometer. Section (ref) concludes. Proofs and technical details are collected in the Appendix.

Measuring Distributional Change for Ordinal Variables

I consider the problem of measuring distributional change in ordinal outcomes observed in repeated cross-sectional data, where only marginal distributions are available. The outcome space is given by $\{1,2,\dots,K\}$ with $K\in \mathbb{Z}_+$, where larger values correspond to higher levels of the attribute of interest (e.g., trust in public institutions). Let $\mu$ and $\nu$ denote two probability distributions over this support, that is, elements of the $K$-dimensional probability simplex

equation*[equation* omitted — 142 chars of source]

corresponding to two populations or time periods.

A natural requirement in this setting is that any measure of distributional change should depend only on the ordering of categories, rather than on arbitrary numerical scores. Moreover, it should reflect the magnitude of distributional shifts, assigning larger weight to reallocations across more distant categories.

Motivated by these considerations, I define the measure of distributional change as

equation*[equation* omitted — 89 chars of source]

where $F_\mu(k) \coloneqq \sum_{i \le k} \mu_i$ and $F_\nu(k) \coloneqq \sum_{i \le k} \nu_i$ denote the cumulative distribution functions associated with $\mu$ and $\nu$, respectively. This measure aggregates discrepancies in cumulative population shares across all ordinal thresholds. For each threshold $k$, the quantity $F_\mu(k)$ represents the share of the population with outcomes at or below level $k$, and $|F_\mu(k) - F_\nu(k)|$ captures the difference in these shares across distributions. Summing across thresholds yields an overall measure of distributional change. By construction, $D(\mu,\nu)$ depends only on the ordering of the categories and is invariant to strictly monotone relabelings of the ordinal scale.

Beyond measuring the magnitude of distributional change, it is natural to ask how such differences can be accounted for across categories. In settings where only marginal distributions are observed, there are many possible ways to reallocate probability mass to transform one distribution into another. This raises the question: among all such reallocations, which one involves the least movement across the ordinal scale? The following result shows that $D(\mu,\nu)$ admits a representation that answers this question.

propositionThe measure $D(\mu,\nu)$ admits the representation \begin{equation*} D(\mu,\nu)= \min_{\pi \in \Pi(\mu,\nu)} \sum_{i=1}^K \sum_{j=1}^K |i-j| \, \pi_{ij}, \end{equation*} where $\Pi(\mu,\nu)$ denotes the set of joint distributions on $\{1,\dots,K\}^2$ with marginals $\mu$ and $\nu$.
proofSee Appendix (ref).

This representation shows that $D(\mu,\nu)$ coincides with the Wasserstein-1 distance between $\mu$ and $\nu$ when the cost of moving mass between categories $i$ and $j$ is given by $|i-j|$. It therefore admits a natural interpretation as the minimal number of ordinal threshold crossings per capita required to transform one distribution into another. The result of Proposition (ref) is a special case of the general result on the equivalence between the Wasserstein-1 distance and the $L_1$ distance between cumulative distribution functions on $\mathbb{R}$, put forward by Vallender.

The optimal transport representation also provides additional descriptive structure beyond this scalar measure. Any minimizer, $\pi \in \Pi^{*}(\mu,\nu)\coloneqq\arg\min_{\pi\in\Pi(\mu,\nu)}\sum_{i=1}^{K}\sum_{j=1}^{K}|i-j|\pi_{ij}$, arises as the solution to a linear optimization problem over the set of joint distributions consistent with the observed marginals. Writing $\pi_{ij}$ for the mass reassigned from category $i$ to category $j$, the matrix $\pi$ can be interpreted as a transition table that reallocates mass from $\mu$ to $\nu$.

The optimal couplings are not intended to represent realized transitions between categories, which are not identified from cross-sectional data. Rather, they describe how probability mass can be reallocated across categories in the least costly way to reconcile two distributions. In this sense, they provide a canonical decomposition of distributional differences, grounded in the ordering of the outcome space, without imposing behavioral or structural assumptions.

This interpretation connects the present framework to the literature on economic mobility (e.g., see Shorrocks-Mobility), where transition matrices are used to describe movements across ordered states. In that literature, transition matrices are typically interpreted as reflecting realized movements of individuals. In contrast, the transition structure is not observed here, but is instead induced by a cost-minimization principle. The resulting elements of $\Pi^{*}(\mu,\nu)$ can therefore be interpreted as minimal mobility tables: they describe the least costly ways in which probability mass can be reassigned to reconcile the two distributions. In this sense, the discrepancy $D(\mu,\nu)$ measures the minimal mobility required to reconcile two distributions, while the associated optimal couplings describe how that mobility can be organized across categories in a cost-minimizing way.

It is useful to relate the minimal mobility tables to an underlying joint distribution that is consistent with the observed marginals. Let $\pi \in \Pi(\mu,\nu)$ denote a joint distribution over categories with marginals $\mu$ and $\nu$. Since only marginal distributions are observed, $\pi$ is not identified; rather, the data are consistent with the entire set $\Pi(\mu,\nu)$.

The optimal couplings $\Pi^*(\mu,\nu)$ should therefore not be interpreted as estimates of any particular joint distribution. Instead, they represent a selection from the identified set based on a cost-minimization principle. In particular, they correspond to the joint distributions that minimize total movement across categories among all observationally equivalent reallocations. In this sense, the minimal mobility tables provide a benchmark describing how the observed distributional differences could arise under the least amount of mobility.

This benchmarking interpretation can be made precise. For any joint distribution $\pi \in \Pi(\mu,\nu)$, define the total amount of movement as $\sum_{i,j} |i-j|\,\pi_{ij}$. By construction, this quantity is bounded below by $D(\mu,\nu)$, which depends only on the marginal distributions. Thus, any joint distribution consistent with the marginals must generate at least $D(\mu,\nu)$ units of aggregate movement. The optimal couplings attain this lower bound and therefore represent minimal-mobility configurations.

Illustrative Example

Consider the distributions $\mu=(0.4,0.3,0.2,0.1)$ and $\nu=(0.2,0.3,0.3,0.2)$, defined over four ordered categories. Their discrepancy is $D(\mu,\nu)=\sum_{k=1}^{3}|F_\mu(k)-F_\nu(k)|=0.5$.

An optimal coupling attains this value by reallocating probability mass in a way that minimizes total movement across the ordinal scale. While such couplings often concentrate mass near the diagonal and rely on adjacent-category movements, they need not be unique. In particular, distinct couplings may achieve the same minimal transport cost while exhibiting different patterns of reallocation.

Figure (ref) illustrates this feature. Both couplings shown in the figure reconcile the same pair of marginal distributions and attain the minimal transport cost of $0.5$, yet they differ in how mass is reassigned across categories. The first coupling concentrates mass along adjacent categories, whereas the second involves a longer jump from category 1 to category 3, offset by other reallocations so as to preserve the same total cost.

This comparison highlights two key features of the framework. First, the scalar discrepancy $D(\mu,\nu)$ captures the minimal amount of aggregate movement required to reconcile the distributions. Second, the associated set of optimal couplings provides a structured decomposition of this movement across categories. Importantly, this decomposition is not unique: multiple minimal-mobility configurations may exist, each representing a different way of organizing the same minimal amount of reallocation.

Together, these objects offer a benchmark description of distributional change that separates what must occur, given the data, from what is merely possible. The scalar discrepancy determines the minimal amount of movement required, while the set of optimal couplings characterizes the range of feasible minimal reallocations consistent with this benchmark.

figure[figure omitted — 2,645 chars of source]

Importantly, optimal couplings need not coincide with the true, but unobserved, joint distribution. In particular, they should not be interpreted as the true joint distribution, but rather as benchmark reallocations that isolate the minimal amount of movement required by the observed marginals. Different optimal couplings represent alternative ways of organizing this minimal movement. The role of these couplings is therefore not descriptive but normative: they provide benchmarks that isolate the least amount of movement required to reconcile the observed distributions. In this sense, $D(\mu,\nu)$ delivers a lower bound on distributional change, while the associated optimal couplings characterize how such minimal change can be organized across categories.

Partial Identification

In practice, the marginal distributions $\mu$ and $\nu$ are often not fully observed due to missing data. I consider a setting with repeated independent cross-sections, where the outcome of interest is observed only for a subset of individuals in each sample. In this setting, both the magnitude of mobility and the associated minimal mobility tables may be only partially identified. This section develops bounds for these objects under missing data.

Let $X\sim\mu$ and $Y\sim\nu$ with $Z_{X}$ and $Z_{Y}$ indicating whether $X$ and $Y$ are observed, respectively. The practitioner observes

align*[align* omitted — 209 chars of source]

where “$*$" denotes the missing value code. Additionally, let $p = \Pr(Z_X=1)$ and $q=\Pr(Z_Y=1)$ denote the response probabilities. I assume that $p$ and $q$ are either known or can be consistently estimated from the data.

Focusing on $\mu$, Let $\mu^{obs}$ denote the distribution of $X$ conditional on $Z_X=1$. The relationship between $\mu$, the population of interest, and $\mu^{obs}$ can be written as

equation*[equation* omitted — 85 chars of source]

where $\mu^{mis}$ denotes the distribution of outcomes among non-respondents, which is unobserved. Without additional assumptions on the missing-data mechanism, the distribution $\mu$ is not point-identified. However, it is partially identified. In particular, the set of distributions consistent with the observed data is given by

equation*[equation* omitted — 154 chars of source]

This characterization implies bounds on the cumulative distribution function of $\mu$, $F_\mu$. Let $F_\mu^{obs}(k) = \sum_{i \le k} \mu_i^{obs}$. Then, for each $k=1,\dots,K-1$,

equation*[equation* omitted — 76 chars of source]

These bounds are sharp in the sense of manski2005: for each admissible value of $F_\mu(k)$ within this interval, there exists a distribution $\gamma \in \mathcal{M}_\mu$ that attains it. An identical line of reasoning applies to the setup for the distribution $\nu$, but with $p$ replaced by $q$, yielding the identified set

equation*[equation* omitted — 154 chars of source]

I now turn to the implications of partial identification for the measure of distributional change. Let $\mathcal{M}_\mu$ and $\mathcal{M}_\nu$ denote the identified sets corresponding to two populations or time periods. Since the measure $D(\mu,\nu)$ depends on the unknown distributions, it is itself only partially identified.

The following result provides a complete characterization of the identified set for $D(\mu,\nu)$.

theoremLet $\mathcal{M}_\mu$ and $\mathcal{M}_\nu$ denote the identified sets for the marginal distributions $\mu$ and $\nu$, respectively. Then the identified set for $D(\mu,\nu)$ is the interval \[ [\underline{D},\overline{D}] = \left[ \min_{\gamma\in\mathcal{M}_\mu,\;\eta\in\mathcal{M}_\nu} D(\gamma,\eta), \; \max_{\gamma\in\mathcal{M}_\mu,\;\eta\in\mathcal{M}_\nu} D(\gamma,\eta) \right], \] Equivalently, the endpoints can be computed by finite-dimensional optimization problems over the feasible marginal distributions. In particular, the lower endpoint admits the linear programming representation \begin{align*} D & = \min_{\gamma,\eta,t} \sum_{k=1}^{K-1} t_k\quadsubject to\quad \gamma \in\mathcal{M}_\mu,\quad \eta\in\mathcal{M}_\nu, \quad and \\ & t_k \ge F_\gamma(k)-F_\eta(k), \qquad t_k \ge F_\eta(k)-F_\gamma(k), \qquad k=1,\ldots,K-1. \end{align*} The upper endpoint is obtained by maximizing the same objective over $\gamma\in\mathcal{M}_\mu$ and $\eta\in\mathcal{M}_\nu$, equivalently by evaluating the maximum over the extreme points of the feasible marginal sets.
proofSee Appendix (ref).

Theorem 1 shows that the bounds $\underline{D}$ and $\overline{D}$ can be computed by solving linear programs. The interval $[\underline{D}, \overline{D}]$ provides a robust measure of distributional change between $\mu$ and $\nu$ that accounts for missing data. The lower bound $\underline{D}$ represents the smallest amount of distributional change consistent with the observed data, while the upper bound $\overline{D}$ represents the largest such change. The width of the interval reflects the degree of identification uncertainty induced by missing observations.

Endpoint-Conditioned Optimal Couplings

The identified interval $[\underline{D},\overline{D}]$ characterizes the range of distributional change between $\mu$ and $\nu$ that is consistent with the observed data without any assumptions on the missingness-generating process. It is also of interest to study the corresponding set of optimal transport couplings at each endpoint of this interval.

To this end, define

align*[align* omitted — 289 chars of source]

Thus, $\mathcal{A}_{L}$ contains the pairs of marginal distributions that attain the smallest distributional change consistent with the data, whereas $\mathcal{A}_{U}$ contains those that attain the largest such change. Additionally, for each $(\gamma,\eta)$, let \[ \Pi^{*}(\gamma,\eta) = \arg\min_{\pi\in\Pi(\gamma,\eta)} \sum_{i=1}^{K}\sum_{j=1}^{K}|i-j|\,\pi_{ij} \] denote the set of optimal transport couplings between $\gamma$ and $\eta$. I then define the endpoint-conditioned optimal coupling sets

align*[align* omitted — 204 chars of source]

For each cell $(i,j)$, these sets induce sharp bounds on the amount of mass that can be transported from category $i$ to category $j$ under an optimal coupling associated with either endpoint of the identified set. Finally, let

align*[align* omitted — 331 chars of source]

The following result shows that these quantities are well-defined and can be computed through finite-dimensional optimization problems.

theoremLet \begin{align*} \mathcal{C}_{L} & \coloneqq \Bigl\{ (\pi,\gamma,\eta)\,:\, \gamma\in\mathcal{M}_{\mu},\; \eta\in\mathcal{M}_{\nu},\; \pi\in\Pi(\gamma,\eta),\;\sum_{r=1}^{K}\sum_{s=1}^{K}|r-s|\,\pi_{rs} = D \Bigr\},\quad and\\ \mathcal{C}_{U} &\coloneqq \Bigl\{ (\pi,\gamma,\eta)\,:\, \gamma\in\mathcal{M}_{\mu},\; \eta\in\mathcal{M}_{\nu},\; \pi\in\Pi(\gamma,\eta),\;\sum_{r=1}^{K}\sum_{s=1}^{K}|r-s|\,\pi_{rs} = \overline{D} \Bigr\}. \end{align*} Then the following statements hold. \begin{enumerate} • $\mathcal{C}_{L}$ and $\mathcal{C}_{U}$ are nonempty and compact. • The sets $\Pi^{*}_{L}$ and $\Pi^{*}_{U}$ are nonempty and compact. • For every $(i,j)\in\{1,\ldots,K\}^2$, the endpoint-conditioned flow bounds are attained and admit the representations \begin{align*} \pi^{\,L}_{ij} = \min_{(\pi,\gamma,\eta)\in\mathcal{C}_{L}} \pi_{ij}, \quad \overline{\pi}^{\,L}_{ij} = \max_{(\pi,\gamma,\eta)\in\mathcal{C}_{L}} \pi_{ij}, \quad \pi^{\,U}_{ij} = \min_{(\pi,\gamma,\eta)\in\mathcal{C}_{U}} \pi_{ij},\quadand\quad \overline{\pi}^{\,U}_{ij} =\max_{(\pi,\gamma,\eta)\in\mathcal{C}_{U}} \pi_{ij}. \end{align*} \end{enumerate}
proofSee Appendix (ref)

Theorem 2 shows that the coupling structure associated with each endpoint of the identified set is itself partially identified. In particular, the interval $[\underline{\pi}^{\,L}_{ij},\overline{\pi}^{\,L}_{ij}]$ describes the range of mass that can be transported from category $i$ to category $j$ among all optimal couplings associated with marginal distributions that attain the lower endpoint $\underline{D}$.

In empirical settings where there exists a population joint distribution $\pi_0$ with marginals $\mu$ and $\nu$, the lower endpoint provides a benchmark for the least amount of mobility required to reconcile the marginals. The actual level of mobility, \[ D_0=\sum_{i=1}^K \sum_{j=1}^K |i-j| \, \pi_{0,ij}, \] satisfies $D_0 \geq D(\mu,\nu) \geq \underline{D}$, where the inequalities follow from Proposition (ref) and Theorem (ref). Thus, $\underline{D}$ represents a lower bound on actual mobility.

The associated lower-endpoint coupling set should not be interpreted as containing the true joint distribution $\pi_0$. In general, $D_0 > \underline{D}$, so that $\pi_0$ need not belong to $\mathcal{C}_L$. Instead, the set $\{[\underline{\pi}^{\,L}_{ij},\overline{\pi}^{\,L}_{ij}],i,j\leq K\}$ characterizes the structure of minimal-mobility benchmark configurations that are consistent with the data.

These bounds admit a direct interpretation. If $\underline{\pi}^{\,L}_{ij} > 0$, then every minimal-mobility configuration must involve a positive flow from category $i$ to category $j$. If $\overline{\pi}^{\,L}_{ij} = 0$, then no minimal-mobility configuration uses this transition. More generally, the width of the interval reflects the degree of flexibility in how minimal reallocation can be organized across categories. In this sense, the lower-endpoint couplings describe not what did occur, but what must occur under the least amount of aggregate movement.

The interval $[\underline{\pi}^{\,U}_{ij},\overline{\pi}^{\,U}_{ij}]$ does not admit a similar benchmarking interpretation. Instead, it serves as a diagnostic tool when considered alongside its lower-endpoint counterpart. Comparing the structure of optimal couplings at the lower and upper endpoints reveals how sensitive the minimal-mobility benchmark is to uncertainty in the marginal distributions. When these structures are similar, the benchmark is robust to missing data. When they differ substantially, the implied structure of distributional change depends critically on the admissible range of marginals.

These endpoint couplings therefore provide extremal representations of how distributional change can be organized, given the uncertainty induced by missing data.

Illustrative Continuation: Endpoint Couplings

To illustrate the endpoint-conditioned coupling sets, I extend the example in Section 2.2 by introducing missing-data structure. Suppose the observed distributions are \[ \mu^{obs}=(0.4,0.3,0.2,0.1), \qquad \nu^{obs}=(0.2,0.3,0.3,0.2), \] and the response probabilities are $p=q=0.95$.

The identified sets for the marginal distributions are constructed as in Section (ref). The endpoints of the discrepancy measure are then given by

align*[align* omitted — 269 chars of source]

These quantities are computed numerically via linear programming. In this example, the identified set for the discrepancy measure is $[\underline D,\overline D]=[0.325,\,0.625]$.

Figure (ref) displays one representative optimal coupling at each endpoint. The lower-endpoint coupling corresponds to the smallest amount of aggregate movement consistent with the data and therefore extends the minimal-mobility benchmark developed in Section 2.2 to a setting with missing data. The upper-endpoint coupling corresponds to the largest discrepancy consistent with the admissible marginals.

figure[figure omitted — 2,731 chars of source]

In this example, both endpoint couplings remain concentrated near the diagonal, although the upper-endpoint coupling involves larger adjacent-category reallocations. This illustrates how endpoint couplings can be used to assess the robustness of the minimal-mobility benchmark to uncertainty in the marginal distributions.

Discussion

The couplings considered in this paper should not be interpreted as representing an underlying joint distribution of outcomes across groups, nor as the outcome of a matching or equilibrium process. In the present setting, no such joint population is observed or identified. Instead, a coupling provides a representation of how probability mass must be minimally reallocated to transform one marginal distribution into another.

Accordingly, the identified set of couplings characterizes the set of all minimal reallocation mechanisms that are consistent with the observed data and the maintained assumptions. This interpretation is particularly useful in applications, as it allows one to distinguish between movements across categories that are required by the data and those that are possible but not identified. In this sense, the framework provides information about the structure of distributional change without imposing a specific structural model of joint outcomes.

The minimal-mobility framework also provides a natural basis for counterfactual analysis when only marginal distributions are observed. While the joint distribution is not identified, the minimal-mobility coupling offers a data-driven reference point that requires the least amount of reallocation of probability mass to reconcile the marginals. This reference can be used to study the sensitivity of counterfactual conclusions. Rather than imposing a specific joint distribution, one can consider a neighborhood of couplings that remain close to the minimal-mobility configuration, in the sense of having transport cost near the minimum. Counterfactual quantities of interest—such as transition probabilities or measures of mobility—can then be evaluated over this neighborhood. This approach provides a transparent and disciplined way to assess how conclusions depend on assumptions about the unobserved joint distribution. It complements the partial identification analysis by replacing worst-case reasoning with a structured notion of local robustness around a benchmark implied by the data.

An important limitation arises when the marginal distributions are identical across the two groups or time periods, in which case the minimal transport cost is zero. In this setting, the data do not require any reallocation of probability mass to reconcile the marginals, and the set of optimal couplings becomes large and uninformative about the structure of transitions. Importantly, this does not imply that no movement has taken place. Substantial shifts across categories may occur while leaving marginal distributions unchanged, for example through offsetting transitions. Rather, the framework indicates that such movements are not identified from marginal information alone. In this sense, the approach characterizes the amount of movement that is necessary to explain observed differences, but cannot detect movements that leave marginals invariant. This limitation reflects the fundamental constraints imposed by marginal data and highlights the distinction between minimal required movement and actual underlying transitions. In this case, the minimal-mobility benchmark remains informative as a statement of what can be inferred from the data, even though it provides no restriction on the structure of underlying transitions.

Maximal-mobility Benchmark

In addition to the minimal-mobility benchmark, it is natural to consider the largest amount of movement across categories that is compatible with the observed marginals. Define \[ M(\mu,\nu) \coloneqq \max_{\pi\in\Pi(\mu,\nu)} \sum_{i=1}^K\sum_{j=1}^K |i-j|\,\pi_{ij}. \] This quantity is well-defined, since it is the value of a linear optimization problem over the compact set $\Pi(\mu,\nu)$. It provides an upper benchmark on feasible mobility, in contrast to $D(\mu,\nu)$, which provides a lower benchmark.

If there exists an underlying joint distribution $\pi_0\in\Pi(\mu,\nu)$, with associated mobility \[ D_0=\sum_{i=1}^K\sum_{j=1}^K |i-j|\,\pi_{0,ij}, \] then \[ D(\mu,\nu)\leq D_0\leq M(\mu,\nu). \] Thus, the pair $(D(\mu,\nu),M(\mu,\nu))$ bounds the range of aggregate mobility consistent with the observed marginals. While $D(\mu,\nu)$ has the interpretation of the least movement required to reconcile the distributions, $M(\mu,\nu)$ characterizes the most extreme reallocation patterns consistent with the same marginal information.

As with the minimal-mobility problem, this is a linear program over the set of couplings, and the set of maximizers \[ \Pi^{\max}(\mu,\nu) = \arg\max_{\pi\in\Pi(\mu,\nu)} \sum_{i,j}|i-j|\pi_{ij} \] can be interpreted as maximal-mobility tables. As in the minimal-mobility case, the set of maximizers need not be a singleton, and different maximal-mobility couplings may exhibit distinct patterns of reallocation. While the minimal-mobility benchmark describes the least amount of movement required to reconcile the marginals, the maximal-mobility benchmark captures the most extreme reallocation patterns consistent with the same data. Together, the two benchmarks bound the range of aggregate mobility compatible with the marginals and can be used to assess the extent to which conclusions depend on assumptions about the underlying joint distribution.

To illustrate, consider the four-category example introduced in Section 2.2 with \[ \mu=(0.4,0.3,0.2,0.1), \qquad \nu=(0.2,0.3,0.3,0.2). \] The minimal-mobility benchmark is $D(\mu,\nu)=0.5$, whereas the maximal-mobility benchmark is $M(\mu,\nu)=2.4$. Figure (ref) displays one representative maximal-mobility coupling (not necessarily unique). In contrast to the minimal-mobility coupling, which concentrates mass near the diagonal, the maximal-mobility coupling reallocates mass across distant categories, subject to the marginal constraints. This contrast makes clear that the same pair of marginals can support very different mobility patterns, even though the minimal benchmark remains the more informative object for identifying what movement is necessarily implied by the data.

figure[figure omitted — 1,617 chars of source]

A corresponding partial identification analysis could in principle be developed for the maximal-mobility benchmark by optimizing over the same identified sets for the marginals. I do not pursue this extension here, as the minimal-mobility benchmark delivers the more informative object, capturing the component of distributional change that is necessarily implied by the data. The maximal-mobility benchmark instead serves as a complementary diagnostic, providing an upper envelope for feasible reallocations and a reference point for sensitivity analysis. This role is closely related to the logic of Fr\'echet-type bounds, which characterize extremal dependence structures consistent with given marginals, but without imposing any notion of distance across categories. Developing a full partial identification analysis for this benchmark is left for future work.

Relation to Fr\'echet Inequalities and Extremal Dependence

The minimal- and maximal-mobility benchmarks are related to the classical Fr\'echet inequalities for the intersection of two events Frechet1935,Frechet1951. Those inequalities bound probabilities of intersections of events using only marginal probabilities and without imposing any assumptions on dependence. For any two elements $A$ and $B$ in the power set of $\{1,2,\ldots,K\}^2$, these inequalities are

align[align omitted — 146 chars of source]

where $\mathbb{P}$ is a probability measure on $\{1,2,\ldots,K\}^2$.

The benchmarks developed here play an analogous role for ordinal mobility. Rather than bounding the probability of joint events, they use only the marginal distributions to bound the smallest and largest aggregate movement across categories that is compatible with the data. In this sense, both approaches derive sharp extremal implications from marginal information alone.

The comparison is nonetheless only partial. Fr\'echet inequalities concern probabilities of intersections, whereas the objects studied here are couplings that optimize an ordinal transport criterion. The minimal-mobility benchmark corresponds to couplings that minimize aggregate movement across ordered categories, while the maximal-mobility benchmark corresponds to those that maximize it within the set of feasible couplings consistent with the marginals. In particular, Fr\'echet bounds are invariant to any notion of distance between categories, whereas optimal transport benchmarks depend explicitly on the ordering and metric structure of the outcome space.

This connection becomes especially clear in the partially identified setting developed in Section (ref). There, Theorem (ref) shows that the endpoint-conditioned coupling bounds are sharp extremal characterizations of admissible flows under missing data. In the same way that Fr\'echet inequalities characterize the range of feasible intersection probabilities consistent with given marginals, the endpoint-conditioned coupling bounds characterize the range of feasible category-to-category reallocations consistent with the observed data and the transport criterion. The analogy is not exact, but both objects summarize what can be learned about an unobserved joint structure from marginal information alone.

The distinction between optimal transport benchmarks and Fr\'echet-type bounds can be illustrated using the example in Section (ref). For the marginals $\mu=(0.4,0.3,0.2,0.1)$ and $\nu=(0.2,0.3,0.3,0.2)$, the Fr\'echet inequalities imply that any feasible coupling $\pi$ must satisfy \[ \max\{0,\mu_i+\nu_j-1\} \leq \pi_{ij} \leq \min\{\mu_i,\nu_j\} \] for all $(i,j)$. While both the minimal-mobility coupling $\pi^*$ and the maximal-mobility coupling $\pi^{\max}$, reported in Figures (ref) and (ref), respectively, satisfy these bounds, they do not, in general, attain them. For instance, for $(i,j)=(1,2)$, the upper Fr\'echet bound is $\min\{\mu_1,\nu_2\}=0.3$, whereas the optimal coupling assigns $\pi^*_{12}=0.2<0.3$. Similarly, for $(i,j)=(3,1)$, the upper bound is $\min\{\mu_3,\nu_1\}=0.2$, yet $\pi^*_{31}=0$. These strict inequalities reflect the fact that optimal transport distributes mass across categories to satisfy a global optimality criterion rather than concentrating mass to attain pointwise bounds. A similar observation applies to the maximal-mobility coupling, which reallocates mass toward distant categories but still does not generally saturate the Fr\'echet inequalities.

This example highlights a fundamental distinction: Fr\'echet bounds characterize the set of feasible joint distributions through pointwise constraints on individual cells, whereas optimal transport selects particular couplings within this set based on a global optimality criterion. Couplings that attain the Fr\'echet bounds concentrate mass as much as possible on specific cells, subject to the marginal constraints, whereas optimal transport couplings typically spread mass across multiple cells to minimize (or maximize) aggregate movement. Accordingly, the minimal- and maximal-mobility configurations are extremal movement structures that are fundamentally distinct from extremal dependence structures characterized by Fr\'echet bounds.

Seen from this perspective, the framework developed here also fits naturally within the broader partial identification literature. That literature studies how economically meaningful objects can be bounded sharply when the data and maintained assumptions do not point identify them. In the present setting, the objects of interest are not treatment effects or structural parameters, but measures and configurations of ordinal mobility. The contribution of the paper is to show that optimal transport provides a tractable way to characterize such objects, and to do so in a form that retains a direct empirical interpretation.

Importantly, the optimal transport couplings considered here need not coincide with the extremal configurations that attain the bounds in ((ref)), since the transport objective imposes an ordinal structure that is absent from the latter. Nevertheless, for any coupling $\pi\in\Pi(\mu,\nu)$, the Fr\'echet inequalities apply with $\mathbb{P}=\pi$.

Inference

The objects of interest in this framework are partially identified, as both the marginal distributions and the associated optimal transport representations are only known to lie within identified sets. Inference therefore proceeds using standard methods for partially identified models. The approach taken here is to construct sample analogs of the identified sets by replacing population quantities with their empirical counterparts, and to conduct inference on functionals of these sets using bootstrap methods that account for sampling variability. Several bootstrap and subsampling procedures are available for inference in partially identified models. For concreteness, I describe an implementation based on the approach of HorowitzManski2000, while noting that alternative procedures could be used

Estimation

Let $\{(Y_i,Z_{Yi})\}_{i=1}^n$ and $\{(X_j,Z_{Xj})\}_{j=1}^m$ denote two independent random samples, where $Y_i\sim\mu$ and $X_j\sim\nu$ take values in $\{1,\dots,K\}$, and $Z_{Yi}$ and $Z_{Xj}$ indicate whether outcomes are observed. Let $\hat{p}_Y$ and $\hat{p}_X$ denote the empirical response probabilities, and let $\hat{\mu}^{obs}$ and $\hat{\nu}^{obs}$ denote the empirical distributions among observed units.

Replacing population quantities with their empirical counterparts yields estimated identified sets for the marginals, \[ \widehat{\mathcal{M}}_\mu = \left\{ \gamma \in \Delta^K : \hat{p}_Y \hat{\mu}^{obs}_k \le \gamma_k \le \hat{p}_Y \hat{\mu}^{obs}_k + (1-\hat{p}_Y), \;\; k=1,\dots,K \right\}, \] and similarly for $\widehat{\mathcal{M}}_\nu$. These sets induce plug-in estimators of the endpoints of the identified set for the discrepancy measure, denoted $\hat{\underline{D}}$ and $\hat{\overline{D}}$, as well as estimators of the endpoint-conditioned coupling bounds obtained by solving the corresponding linear programs.

Confidence Sets

Because the estimators are functions of empirical distributions, they are subject to sampling variability. To account for this uncertainty, I employ a bootstrap procedure that resamples the observed data $\{O^Y_i\}_{i=1}^n$ and $\{O^X_j\}_{j=1}^m$ with replacement (independently across samples) and recomputes all objects of interest for each replication.

{\bf Confidence set for $[\underline{D},\overline{D}]$.} Let $\hat{\underline{D}}$ and $\hat{\overline{D}}$ denote the plug-in estimators of the lower and upper bounds. A confidence region for the identified set is constructed as \[ \left[\hat{\underline{D}}-c_{1-\alpha},\;\hat{\overline{D}}+c_{1-\alpha}\right], \] where the critical value $c_{1-\alpha}$ is obtained from the bootstrap distribution of the maximal deviation of the bound estimators. Under standard regularity conditions bickel1981, this procedure yields asymptotically valid coverage of the identified set.

{\bf Confidence sets for endpoint-conditioned couplings.} The endpoint-conditioned coupling bounds define a finite-dimensional parameter vector, indexed by $(i,j)$ and the endpoint. Inference for these objects proceeds analogously by applying the bootstrap to the corresponding estimators. Confidence intervals for each component, as well as simultaneous confidence regions, are constructed using the bootstrap distribution of maximal deviations. Details are provided in Appendix (ref).

A practical advantage of this approach is that inference can be implemented using standard resampling procedures without requiring problem-specific asymptotic derivations. The bootstrap propagates sampling variability through both the estimation of the marginal distributions and the optimization steps defining the identified sets and coupling bounds, thereby accounting for errors due to sampling and missing data.

Empirical Illustration

I illustrate the framework using the Arab Barometer, focusing on question Q700B on favorability toward the United States in Waves 7 and 8 for Iraq and Morocco. The question is

quote“Please tell me if you have a very favorable, somewhat favorable, somewhat unfavorable, or very unfavorable opinion of The United States”

Responses are recorded on a four-point ordinal scale: \[ 1 = \text{very favorable}, \; 2 = \text{somewhat favorable}, \; 3 = \text{somewhat unfavorable}, \; 4 = \text{very unfavorable}. \]

A central question in this application is how much of the population must change their reported attitudes to reconcile the distributions across waves. Standard comparisons of marginal distributions cannot answer this question, as they do not account for how responses are reallocated across categories. The proposed framework provides a lower bound on this reallocation and characterizes how it must occur.

The analysis is based on the publicly available Arab Barometer datasets, which consist of completed interviews. As a result, the missing observations considered in this illustration correspond to item nonresponse, specifically “don’t know” and “refused to answer” responses to the survey question. Unit nonresponse—individuals who were not interviewed—is not observed in the data. Sample sizes are substantial in both countries: for Iraq they are $1299$ and $1190$ in Waves 7 and 8, respectively, and for Morocco they are $1227$ and $1152$. Reported response rates (AAPOR Response Rate 1) vary across countries and waves: they are $77\%$ and $58\%$ for Iraq and $38\%$ and $69\%$ for Morocco in Waves 7 and 8, respectively. Item nonresponse rates for this question are low: for Iraq they are $0.85\%$ and $0.25\%$, and for Morocco $1.39\%$ and $0.26\%$ across the two waves. The framework developed in the paper applies more generally to settings with missing data, but the empirical illustration focuses on item nonresponse due to data availability. Accordingly, the results pertain to distributional differences within the responding population.

figure[figure omitted — 374 chars of source]

Wave 7 is treated as the source distribution and Wave 8 as the target. I set $\alpha=0.05$ and $B=499$. The confidence set for the identified set $[\underline{D},\overline{D}]$ is $[0.125,0.316]$ for Iraq and $[0.167,0.347]$ for Morocco. Normalizing by the maximum possible discrepancy of $3$ yields intervals $[0.042,0.105]$ and $[0.056,0.116]$, respectively.

{\bf Iraq.} The cumulative distributional bounds for Iraq, reported in the left panel of Figure (ref), suggest that the two waves differ, but the discrepancy remains moderate. The normalized interval $[0.042,0.105]$ indicates that the minimum required reallocation corresponds to approximately $4\%$--$11\%$ of the maximum possible movement. Thus, at least this amount of normalized movement is required to reconcile the two waves. This indicates that the observed differences cannot be explained solely by negligible perturbations in the marginal distribution and instead require a nontrivial reshuffling of responses.

The lower bound provides a benchmark for the least amount of change required to reconcile the two waves. The corresponding lower-endpoint coupling shows how this minimal change can be organized across categories. The top heatmaps in the top panel of Figure (ref) report the confidence set for the lower-endpoint coupling and its width. The heatmaps are concentrated along the diagonal, with limited off-diagonal movement. Most mass remains within the same or adjacent categories, and the bulk of reallocation occurs through one-step transitions ($|i-j|=1$), with little evidence of larger jumps. Thus, even the minimal change can be accounted for by relatively local shifts in responses.

This configuration serves as a benchmark: alternative explanations of distributional change must involve at least this level of movement, and typically more complex reallocations. In particular, explanations based on large cross-category shifts would imply movement exceeding the minimal benchmark. This benchmark also restricts the set of admissible explanations within the identified set. In particular, explanations based on large-scale polarization---where a substantial share of the population moves across distant categories---would require a large amount of mass to be transported over long distances. Such patterns would generate a transport cost exceeding the lower endpoint of the identified set. Since the lower endpoint is achieved by reallocations that are concentrated on the diagonal and adjacent categories, explanations based on large-scale polarization are not consistent with minimal-mobility accounts of the data.

The confidence set for the upper-endpoint coupling, given by the bottom heatmaps of the top panel in Figure (ref), provides a diagnostic for sensitivity to missing data. While the interval for the discrepancy measure reflects uncertainty due to missing data, the stability of the endpoint-conditioned couplings indicates that conclusions about how minimal distributional change occurs are robust to item nonresponse.

figure[figure omitted — 453 chars of source]

{\bf Morocco.} The cumulative distribution function bounds in the right panel of Figure (ref) point to somewhat larger distributional change than in Iraq. The normalized confidence set is given by $[0.056,0.116]$. This interval is shifted to the right relative to that of Iraq, indicating that the minimal required reallocation is larger in Morocco.

The confidence set of the lower-endpoint coupling, reported as the top heatmaps in the bottom panel of Figure (ref), again provides the benchmark configuration. Compared to Iraq, it exhibits somewhat more pronounced off-diagonal movement, with a larger share of mass reassigned across adjacent categories. While transitions remain concentrated near the diagonal, their intensity is higher, indicating a somewhat more substantial reshuffling of responses.

As in Iraq, this configuration provides a benchmark against which alternative explanations can be assessed. Explanations involving large jumps across distant categories would imply movement exceeding the minimal benchmark, while explanations based on localized changes are consistent with the observed structure. This suggests that distributional change in Morocco is not only larger in magnitude but also involves somewhat more substantial reshuffling across categories, consistent with a broader shift in attitudes rather than purely localized adjustments.

The comparison with the upper-endpoint coupling, reported as the bottom heatmaps in the bottom panel of Figure (ref), shows that this structure is stable: the qualitative pattern of reallocations remains similar across endpoints. Thus, while uncertainty affects the exact magnitude of minimal required movement, the structure of the benchmark is robust.

{\bf Summary.} The empirical results deliver three main findings. First, a nontrivial amount of movement is required to reconcile the observed distributions: approximately $4\%$--$11\%$ of the maximum possible movement for Iraq and $6\%$--$12\%$ for Morocco. Second, this change is primarily organized through local transitions across adjacent categories, indicating gradual shifts rather than large-scale polarization. Third, while the magnitude of change is partially identified, the structure of minimal-mobility couplings is stable across feasible marginal distributions, suggesting that conclusions about how minimal distributional change occurs are robust to item nonresponse. Taken together, these findings show that the observed differences reflect systematic reallocation patterns that cannot be inferred from marginal comparisons alone.

Conclusion

This paper studies distributional change when only marginal information is available, as in settings with repeated cross-sections. I propose a framework based on optimal transport that measures the smallest amount of reallocation of probability mass required to reconcile two distributions and characterizes how this minimal change is organized across categories.

The analysis delivers two key objects: a scalar measure of distributional change, interpreted as a lower bound on the mobility required to reconcile the marginals, and a set of minimal-mobility couplings that describe the structure of this change. In the presence of missing data, the framework extends naturally to partial identification, yielding bounds on both the magnitude and the structure of distributional change, together with inference procedures that account for sampling uncertainty.

The empirical illustration shows how these objects can be used in practice. The lower endpoint of the discrepancy measure provides a benchmark for the least amount of change required by the data, while the associated couplings describe how this change must be realized across categories. Comparing these benchmark configurations across the identified set provides a diagnostic for sensitivity to missing data.

The analysis also points to natural extensions. In particular, one can define complementary benchmarks based on maximal mobility, which characterize the largest amount of movement compatible with the marginals. More broadly, the framework is related to classical Fr\'echet-type bounds, in that both approaches use marginal information to characterize feasible joint structures. The key distinction is that the present approach imposes an ordinal transport structure, allowing one to organize feasible reallocations according to meaningful notions of distance.

From a practical perspective, the framework can be implemented using a simple sequence of steps. First, estimate the marginal distributions and construct their identified sets using observed data and bounds implied by missingness. Second, compute the endpoints of the discrepancy measure by solving the corresponding optimal transport problems over these sets. Third, recover the associated minimal-mobility configurations (optimal couplings) at these endpoints. Together, these objects provide a transparent decomposition of distributional change into components that are necessarily implied by the data and those that remain unidentified. This makes it possible, in applied work, to move beyond comparisons of marginals and to characterize the minimal structure of distributional change that is robust to missing data and consistent with observed marginals.