Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
63,965 characters · 11 sections · 29 citation commands
Compositional Difference-in-Differences
Many causal questions in Difference-in-Differences (DiD) settings involve vectors of categorical quantities that together form a well-defined total. These arise naturally when individual-level categorical outcomes are aggregated. For instance, in evaluating a minimum wage policy, individuals may be classified as unemployed, part-time employed, or full-time employed, and the compositional vector counts how many fall into each group. In other cases, the data consist directly of quantities across categories, with no underlying individual identifiers, such as the number of votes received by each political party in an election. Compositional vectors also appear when a single quantity is decomposed into meaningful parts, as in electricity generation by source (gas, coal, nuclear, renewable), GDP by sector (agriculture, industry, services), or a budget allocated to each category (for example, food, rent, healthcare, others). In all such cases, the outcome is multidimensional, and both the shares and the total matter for the analysis.
In many empirical settings, researchers focus on studying treatment effect on the intensive margin, captured by the shares of outcomes allocated across categories. When the outcome is categorical, they typically convert unordered qualitative outcomes into binary indicators use parallel trends on each of these binary outcomes. Because shares encode only relative information and are constrained between zero and one, using linear DiD in this case can be problematic, as noted by DiD_Review2011, athey2006identification, wooldridge2023simple, ahn2025event). A range of alternative causal inference methods, such as the geodesic difference-in-differences zhou2025geodesic, and transition-independence approaches ahn2025event, are better suited since they cannot generate invalid counterfactuals. These methods allow researchers to trace how the distribution across categories evolves when a policy changes, while respecting the intrinsic constraints of the compositional shares(such as non-negativity and the unit-sum property). For example, they can be used to assess how a voting reform reshapes the distribution of votes across parties or how an educational intervention changes the allocation of students across majors. However, they have some important limitations.
The first is that treatment often affects not only the relative composition of outcomes but also their total level, the extensive margin. In such cases, focusing solely on shares provides an incomplete picture. A voting reform may reshape the distribution of votes while also changing total turnout. A minimum-wage increase may alter the probabilities of being in different employment categories and affect the overall number of workers. An education reform may shift students across majors while simultaneously changing the total number of graduates. Likewise, a carbon policy may influence the shares of energy sources in electricity production and the total electricity generated. Those methods cannot jointly and consistently recover treatment effects on both shares and totals. Yet, many empirical questions require a framework that coherently incorporates absolute and relative changes simultaneously. The second problem is that those approaches lack economic grounding, as they cannot be supported by behavioral or structural relationships that typically govern how categorical margins evolve. This makes it harder to intuitively evaluate their meaning in practice.
This paper introduces Compositional Difference-in-Differences (CoDiD), a framework for analyzing vectors of category-specific quantities (compositional vectors) in a DiD setting while jointly capturing effects on both the intensive and extensive margins. The key identifying condition is a parallel growth assumption, which delivers point identification of the entire counterfactual compositional vector, and therefore of both counterfactual shares and totals. Parallel growth is simply a parallel trends assumption applied to the log-quantities of each category: in the absence of treatment, every category’s absolute size would expand or contract at the same proportional rate in the treated and control groups. To justify this assumption, I model the observed category quantities as aggregates of individuals’ discrete choices within a standard random utility framework. In this setting, parallel growth is equivalent to assuming parallel trends in the latent expected utilities associated with each category. Intuitively, absent treatment, both groups would experience the same evolution in how attractive each option is on average. For instance, if the expected utility of voting for the Democratic candidate rises in the control group, it would have risen by the same amount in the treated group had treatment not occurred. This interpretation clarifies the causal parameter: the treatment effect captures the extent to which the policy shifts the underlying desirability of each category.
The implied evolution of the shares corresponds to parallel trends in the log-odds transformation of category shares. In the same random utility model, log-odds reflect differences in latent utilities and therefore capture relative preferences. Under this interpretation, parallel growth means that, absent treatment, the evolution of relative preferences between any pair of categories would be the same in both groups. Put differently, the way individuals substitute between categories over time in the control group would mirror the untreated evolution in the treated group. Finally, from a geometric perspective, each share vector corresponds to a point in the simplex, and its evolution traces a path within this constrained space. The parallel growth assumption implies that, absent treatment, the treated group’s counterfactual path would be parallel to the control group’s path under Aitchison geometry, the standard geometric framework for compositional data (aitchison1982statistical; egozcue2003isometric).
In settings with multiple time periods, the parallel growth assumption can be evaluated visually by plotting the log-transformed quantities over time. If the pre-treatment trends do not appear approximately parallel, the credibility of the assumption weakens. To address such cases, I propose a relaxation in the spirit of ban2022robust. The idea is to bound the counterfactual log differences between groups by restricting their changes to lie within the convex hull of observed pre-treatment changes. This procedure yields bounds, rather than point estimates, for both the counterfactual categorical quantities and the resulting counterfactual distributions.
As an application, I examine the impact of the early voting program on the 2008 U.S. presidential elections. These programs, allowing voters to cast ballots before election day, either in person or by mail, have been the subject of intense partisan debate. I revisit this question by analyzing both total turnout and the compositional shift in votes for parties. I consider Maryland and New Jersey, which introduced early voting before 2008 in the treated group, while Pennsylvania and New York serve as the control group. As expected, the treatment increased overall turnout ( increase of $4.4\%$) of total number of voters. The ATT shows that the Republican party lost 1.11 percentage points, 0.9 in favor of democrats, and 0.3 in favor of alternative parties.
This paper makes some methodological contributions at the intersection of difference-in-differences and compositional data analysis.
First, it contributes to the extensive DiD literature. Classical DiD focuses on scalar outcomes, estimating treatment effects on the mean, with surveys and applications summarized in DiD_Review2011, Review2, de2023two, roth2023s, baker2025difference. Recent works have extended DiD to more complex settings: CallawaySant2020 develops methods for staggered treatment adoption, arkhangelsky2021synthetic combines DiD with synthetic controls, and manski2018right, RamRoth2020, ban2022robust relax the parallel trends assumption. A second wave of research extends DiD to entire outcome distributions. Early contributions include quantile DiD meyer1990workers and the changes-in-changes (CiC) approach athey2006identification, which identifies counterfactual distributions non-parametrically. Subsequent work has generalized these ideas to multivariate outcomes torous2021optimal, settings where identifying assumptions apply to cumulative distribution functions HavnesMogstad2015, RothSantanna2021, copula-based approaches CallawayLiOka2018, CallawayLi2019, ghanem2023evaluating, and methods using characteristic functions bonhommesauder2011. However, these methods do not directly address categorical outcome distributions, which are shares. graves2022difference proposes a DiD for categorical outcomes while relying on a linear parallel trends for proportions, which I explained earlier, is not appropriate in general. zhou2025geodesic extends the difference-in-differences framework to non-Euclidean data, including shares, by employing Fréchet means and geodesic transport. While their approach is mathematically elegant, the identifying assumptions are primarily geometric and offer limited economic interpretability. In contrast, my framework delivers both counterfactual quantities and shares that retain economic intuition while incorporating a geometric perspective specifically designed for categorical outcomes. Another related paper is UDID2024_Epi, who propose a general DiD framework applicable to count data in the binomial setting. My analysis differs by focusing on the multinomial case, allowing for richer categorical structures and more general forms of distributional change. Second, the paper contributes to the compositional data analysis (CoDA) literature. CoDA, pioneered by aitchison1982statistical, Aitchison1990, Aitchison1992, Aitchison2002, provides tools for analyzing data constrained to the simplex. Subsequent developments egozcue2003isometric, BillheimerEtAl2001, BarceloVidalEtAl2001 refine transformations and models that respect the geometric structure of compositional data. More recently, arnold2020causal connected CoDA to causal inference. I also combine compositional data methods with econometric identification strategies to analyze causal effects when outcomes are shares, providing a bridge between DiD and compositional statistics.
The paper is organized as follows. Section (ref) introduces the canonical $2 \times 2$ case, presenting the main identifying assumption along with its economic and geometric interpretations. Section (ref) generalizes the framework to settings with multiple time periods. Section (ref) illustrates the methodology through the empirical application. Finally, Section (ref) concludes.
In this section, I focus on the canonical, two-group, two-time period case in the DiD framework, and I clearly expose the assumptions and their implications. I will start with the analytical framework, later discuss the main identifying assumption, and after all the economic and geometric interpretations.
In this section, I introduce the basic framework of my method. Consider some aggregated unit for which we observe the following vector of categorical quantities: \[ q = (q_1, q_2, \dots, q_p). \] I consider a setting with two groups, \( g \in \{0,1\} \), where \( g=1 \) is the treated group and \( g=0 \) is the control group. These groups are observed in two time periods, \( t \in \{0,1\} \), with \( t=0 \) being the pre-treatment period and \( t=1 \) the post-treatment period. For any category \( k \), I define \( q^{0}_{k,g,t} \) as the potential untreated quantity (the quantity in \( k \) if the group $g$ were untreated at time $t$). I also have \( q^{1}_{k,g,t} \): the potential treated quantity (the quantity in \(k \) if the group $g$ were treated at time $t$). For the untreated potential outcomes, I aggregate across categories to define the total quantity for a group-period: \[ S^{0}_{g,t} = \sum_{k=1}^{p} q^{0}_{k,g,t}.\]
The share (or probability) for category \( k \) is then: \[\pi^{0}_{k,g,t}= \frac{q^{0}_{k,g,t}}{S^{0}_{g,t}}. \] The vector of quantities and the entire shares are represented as: \[ q^{0}_{g,t} = \left[q^{0}_{1,g,t}, \dots, q^{0}_{p,g,t}\right], \quad \pi^{0}_{g,t} = \left[\pi^{0}_{1,g,t}, \dots, \pi^{0}_{p,g,t}\right], \quad \text{where} \quad \sum_{k=1}^{p} \pi^{0}_{k,g,t} = 1. \] All the same definitions apply to the treated potential quantities, \( q^{1}_{k,g,t} \), and their associated totals and shares.
From the data, I can identify the vectors $q^{0}_{0,0}$ for the control group and $q^{0}_{1,0}$ for the treated group in the pre-treatment period. In the post-treatment period, I can identify from data the untreated vector of quantities for the control group, $q^{0}_{0,1}$, and the treated counterpart for the treated group, $q^{1}_{1,1}$. The counterfactual of interest is the vector of quantities that I would have observed for the treated group in the absence of treatment $q^{0}_{1,1}$. Identifying this missing object is the core challenge of my analysis. Once this is identified, I can consider the treatment effect parameters.
\paragraph{Treatment effect parameters:}
I introduce treatment effect parameters to capture the impact of the intervention. The first is the growth treatment effect on the treated (GTT), which quantifies the causal effect on the absolute size of a category. For each category \( k \) (where \( k = 1, \ldots, q \)), the GTT is defined as the proportional change in its quantity:
This parameter has an intuitive interpretation: $GTT = 0$ implies the treatment had no effect on the absolute. \(GTT > 0 \) indicates that the treatment caused category \(k\) to grow in absolute terms. For example, a GTT of $0.15$ means the category's size increased by $15\%$ due to the treatment. \(GTT < 0\) signifies that the treatment caused the category to shrink. A GTT of $-0.10$, for instance, corresponds to a $10\%$ decline.
When comparing the shares, there are several ways to do it, depending on the type of effects that we would like to emphasize. To quantify the absolute change in shares, we can focus on the Average treatment effect on the treated (ATT), defined as the simple difference in category proportions (or shares) between the treated potential outcome distribution and its untreated benchmark.
This representation is intuitive, easy to interpret, and directly reflects the absolute reallocation of probability mass across categories induced by the treatment.
Sometimes, it might be interesting to see which category has gained relatively the most as a result of treatment. The compositional treatment effect on the treated (CTT) captures how treatment redistributes weight across categories proportionally and also tells us which category gained relatively the most. To define it, I consider the compositional difference operator commonly used in compositional data analysis to compare shares. This operator quantifies how each category expands or contracts relative to all others. For intuition, fix a category $k$ and consider its proportional change: \[ r_k = \frac{ \pi_{1,1}^{1} (c_k)}{ \pi_{1,1}^{0}(c_k) }, \quad
\] Dividing by the total change gives a vector showing each category’s relative gain or loss. I obtain the compositional difference operation $(\ominus)$ used in compositional data analysis for the control group.
Each component then reflects how much a category’s importance has shifted relative to the others. Components larger than $1/p$ indicate categories that gain relative importance, while components smaller than $1/p$ indicate categories that lose. If all ratios equal one and the operator yields the uniform composition $(1/p,\ldots,1/p)$, we therefore have:
A useful way to think about CTT is through a neutral individual who starts indifferent, assigning equal probability to all $p$ categories $(1/p)$. After treatment, this person updates their weights according to CTT. Components greater than $1/P$ indicate categories that gain relative importance, while smaller components indicate categories that lose relative importance. For example, in a three-party voting scenario, a neutral voter starts with $(1/3, 1/3, 1/3)$. After a campaign (treatment), CTT might imply $(0.5, 0.3, 0.2)$. This shows that support for the first party increased, while the other two declined, with the second party still stronger than the third.
One can also consider the treatment effects on some functionals of the distribution. Consider, for example, a policy evaluated using a scalar index that is a nonlinear functional of compositional shares, such as the Herfindahl–Hirschman Index (HHI): \[ \text{HHI} = \sum_{k=1}^K \pi(c_k)^2 \] A common approach would be to apply DiD directly to the index. However, this strategy extrapolates the index directly, ignoring that it is derived from shares that must satisfy some mathematical and economic constraints. A better approach is to get the counterfactual shares first and later recompute the counterfactual index. Consider a general function $H$ on the shares, a causal analysis on $H(\pi)$ would be the functional treatment effect on the treated (FTT), defined for a given choice of $H$
In the next paragraph, I introduce the identifying assumption and its implications.
This section develops the identification strategy and clarifies its implications for the shares. I start with the following assumption.
Assumption (ref) guarantees that all category-specific quantities are strictly positive in every group and period in the absence of treatment. This ensures that both observed and counterfactual compositional vectors are well defined and comparable across groups and time. In particular, it rules out settings where a category disappears entirely for one group or in one period, which would break the alignment of supports required for the DiD analysis. Assumptions of this form are standard in compositional settings.
I now turn to the core identification idea. The logic mirrors the standard DiD framework: absent treatment, the treated group would have evolved in parallel with the control group. Figure (ref) illustrates this intuition; the untreated control group traces out the counterfactual trajectory that the treated group would have followed had the intervention not occurred.
In many empirical settings, the sizes of the groups defined in each category can vary dramatically in scale. Therefore, a standard linear parallel trends assumption on the raw quantities can be problematic in this context, as it would impose the same absolute change on both large and small groups, which is often unrealistic. To address this issue of scale and to formulate a more plausible identifying assumption, I instead consider parallel trends on the log transformation of the quantities. This approach focuses on proportional, rather than absolute, changes. For a given vector, define the functions $(\log) $ and $(\exp) $ which apply respectively the logarithm and the exponential functions component-wise to the vectors.
Assumption (ref) is the core identifying condition for the model. It states that, in the absence of treatment, the proportional growth (or decline) of each category in the treated group would have been the same as that of the control group. This reflects the intuitive idea that underlying trends affect all categories proportionally to their size. This may appear as a more realistic and flexible assumption than one based on raw quantities in many economic contexts, especially when the sizes of categories differ significantly. It captures the idea that trends often affect all units proportionally to their size. It is also possible to use discrete covariates to make the assumption more believable. I discuss this possibility in appendix (ref). From this assumption, the counterfactual quantity vector for the treated group can be easily recovered, and the corresponding counterfactual probability mass can be obtained by normalization, as we can see in Theorem (ref). Taken together, assumptions (ref) and (ref) impose sufficient structure to point-identify the counterfactual distributions. The following theorem (ref) formalizes this result
This identification result is constructive and provides a straightforward way to implement the method.Before I present the economic and geometric implication of parallel growths, I show how this assumption restricts the evolution of each margin.
The parallel growths assumption that we impose directly on categorical quantities also restrict the evolution of shares and totals in a certain way. To see how it restricts the evolution of the shares, define the multinomial logistic transform. \(\pi = [\pi_1, \dots, \pi_p] \in \mathcal{S}^{q-1}\). The log-odds (or multinomial logit) transformation \(\ell: \mathcal{S}^{p-1} \to \mathbb{R}^{p-1}\) maps probabilities to real-valued indices: \[ \ell(\pi) = \Bigg[ \log \left( \frac{\pi_1}{\pi_p} \right), \dots, \log \left( \frac{\pi_{p-1}}{\pi_p} \right) \Bigg]. \] Each component of this vector is the log-odds of category $k$ compared to the baseline category $p$. In fact, the ratio $\pi_{k,g,t}^{0} / \pi_{p,g,t}^{0}$ is the odds of category $k$ relative to the reference category $p$, and taking the logarithm gives the log-odds. The choice of the baseline category can be arbitrary. Its inverse is given by the softmax function: for any \(y = (y_1, \dots, y_{p-1}) \in \mathbb{R}^{p-1}\),
The parallel growths assumption implies a parallel trends assumption on the log-odds transformation of the probability distribution, as stated in this proposition.
From that, I can also identify the counterfactual distribution with an alternative formula:
We are also interested in examining how it restricts the evolution of the totals. Re-expressing the parallel growths assumption in terms of observable totals and category shares yields a transparent decomposition:
Here, $S_{1,0}^{0}$ is the treated group’s observed pre-treatment total, $S_{0,1}^{0}/S_{0,0}^{0}$ is the aggregate growth rate in the control group, and \[ \lambda \equiv \sum_{k=1}^{p} \left( \frac{\pi_{k,0,1}^{0} }{\pi_{k,0,0}^{0} } \right) \pi_{k,1,0}^{0} \] is the composition adjustment factor, a reallocation multiplier that captures how the treated group’s initial allocation across categories aligns with the control group’s internal reallocation dynamics. The factor $\lambda$ is essential because it accounts for a fundamental source of heterogeneity that direct DiD on log-total ignores: even when two groups experience identical category-level growth rates, they will exhibit different aggregate outcomes if their initial compositions differ. Specifically: If the treated group is initially concentrated in categories that gained relative share in the control group ($\pi_{k,0,1}^{0} > \pi_{k,0,0}^{0}$), then $\lambda > 1$, implying a counterfactual total larger than what aggregate scaling alone would suggest. If it is concentrated in categories that lost relative share, then $\lambda < 1$, implying a counterfactual decline, even if the control group’s total is unchanged. Only when $\pi_{1,0}^{0} = \pi_{0,0}^{0}$ does $\lambda = 1$, reducing Equation (ref) to the standard DiD in log-total prediction.
To see why the composition adjustment factor matters, consider evaluating an early voting policy. Suppose the control state shows no change in total turnout across elections, but its vote composition shifts: Democratic vote share rises while Republican share falls. Now assume the treated state, before the policy, has an electorate that is disproportionately Republican, the group losing relative share in the control state. Even with flat aggregate turnout in the control state, the treated state’s counterfactual turnout would have declined absent the policy, because its composition is weighted toward the shrinking group. In CoDiD, this appears as a composition adjustment factor. Standard DiD, which imposes parallel trends in total turnout, would wrongly assume no counterfactual change and thus underestimate the policy’s effect.
In this section, I provide economic justifications for parallel growths using the standard random utility framework mcfadden1972conditional,mcfadden1974measurement, mcfadden1977modelling.\\
Consider a standard Random Utility Model, in which a decision-maker selects one alternative from a finite set of mutually exclusive choices $\{1, 2, \dots, p\}$. The utility that the individual derives from alternative $k$ in the absence of treatment is modeled as: \[ U^{0}_{k,g,t} = V^{0}_{k,g,t} + \varepsilon^{0}_{k,g,t}, \] where $U^{0}_{k,g, t}$ denotes the latent utility of category $k$ in group $g$ at time $t$ in the absence of treatment, and $V^{0}_{k,g,t}$ is the systematic (deterministic) component of utility (can be seen as the measure of the total attractiveness of the category) and $\varepsilon^{0}_{k,g,t}$ are the random unobserved shocks. The individual is assumed to choose the alternative that yields the highest latent utility: \[ \text{Choice} = \arg\max_{k} U^{0}_{k,g,t}. \] I decompose the total attractiveness into two parts: \[ V^{0}_{k,g,t} = \underbrace{\mu^{0}_{k,g,t}}_{\text{relative attractiveness}} - \underbrace{\log\left(\sum_{k} e^{\mu^{0}_{k,g,t}}\right)+ \log\left(S^{0}_{g,t}\right)}_{\text{absolute attractiveness}}. \] This decomposition shows that \( V^{0}_{k,g,t} \) contains two additive components: the first is the relative desirability of the category (within-choice structure) \( \mu^{0}_{k,g,t} \). This term reflects how desirable option \( k \) is relative to all other alternatives at the same \( (g,t) \). The second is absolute desirability (aggregate intensity) \( - \log\Big(\sum_{j} e^{\mu^{0}_{j,g,t}}\Big)+ \log(S^{0}_{g,t})\). This term captures the overall size or participation level (e.g., how many people are making a choice at all). It scales all utilities equally, affecting the total mass of choices but not their relative shares. When it rises (say, due to higher purchasing power, stronger participation, or favorable aggregate conditions), all utilities shift up equally, people are more likely to choose something, but the proportions among categories might remain the same unless the first component also changes. Together, they form a utility concept that blends relative preferences and absolute participation. This decomposition is useful to see how both components can be affected. We finally have: \[ U^{0}_{k,g,t} = \mu^{0}_{k,g,t} - \log\left(\sum_{k} e^{\mu^{0}_{k,g,t}}\right) + \log\left(S^{0}_{g,t}\right) + \varepsilon^{0}_{k,g,t}. \] In the traditional RUM, we assume that the random terms $(\varepsilon^{0}_{1,g,t}, \dots, \varepsilon^{0}_{p,g,t}$ all follow standard type-1 extreme value distribution and are independent. Because each marginal error term remains Gumbel-distributed with mean \( \gamma \approx 0.5772\) (the Euler–Mascheroni constant), the expected utility of alternative \(k\) is: \[ \mathbb{E}[U^{0}_{k,g,t}] = \mu^{0}_{k,g,t} - \log\left(\sum_{k} e^{\mu^{0}_{k,g,t}}\right) + \log\left(S^{0}_{g,t}\right) + \gamma. \] \(\mathbb{E}[U^{0}_{k,g,t}]\) captures the average latent total desirability in the population for category $k$. The first implication of the RUM is given by the following proposition. Define the following vector of expected utilities: \[\mathbb{E}[U^{0}_{g,t}] = \Big[ \mathbb{E}[U^{0}_{1,1,1}], \cdots, \mathbb{E}[U^{0}_{p,1,1}] \Big]. \]
This equivalence provides a clear economic interpretation of the identifying assumption. It states that, in the absence of treatment, both groups would have experienced the same evolution in how valuable each option is on average. For example, in a voting context, suppose the control group experiences a rise in the expected utility of voting Democrat (e.g., due to national political trends). Assumption (ref) means that the treated group would have experienced the same change in expected utility had it not been exposed to the policy. The same logic applies to all categories. Thus, the parallel growths assumption is not merely a mathematical condition on outcomes; it can also be seen as an economic hypothesis about the common evolution of underlying category-specific attractiveness in the absence of treatment. This provides a micro-founded justification for CoDiD: it estimates how a policy shifts the attractiveness of alternatives as measured by expected utility.
The log-odds of choosing one category over another provide a direct empirical analogue to differences in the underlying expected utilities. They summarize how much more desirable one option is relative to another, capturing the relative strength or simply the preferences in the population. For instance, in a voting context, the log-odds of choosing Democrat over Republican quantify the population’s relative preference for Democrats. When the log-odds rise, it means that the expected utility, or average attractiveness, of voting Democrat has increased relative to that of voting Republican. The parallel growths assumption implies that, in the absence of treatment, these relative preferences would have evolved in parallel across groups for any pair of categories. That is, any secular change, such as a national political shift increasing the attractiveness of Democrats relative to Republicans, would have affected both treated and control groups identically had the treatment not occurred. The treatment effect is then identified as the deviation from this common trajectory.
To summarize, I have shown that assuming parallel trends in log-quantity ratios is not ad hoc; it is equivalent to assuming parallel trends in expected utilities, a natural behavioral assumption under the Random Utility Model. Because utility differences (not levels) drive choice, this assumption is economically meaningful. This foundation justifies the CoDiD approach as a method for estimating how a policy changes the relative attractiveness of each alternative compared to others. In Appendix (ref), I extend this result to more general Random Utility Models that allow for dependence among the error terms, thereby relaxing the IIA assumption implicitly assumed in the standard multinomial RUM.
In this section, I provide some additional geometric justifications for parallel growths. I start by interpreting the restriction on the shares using the compositional difference operator. \\
Interpretation using the compositional difference interpretation: To further support the relevance of the assumption in this context, I provide an alternative interpretation that does not depend on the random utility framework or its underlying assumptions. Suppose our goal is to understand how the treatment redistributes probability mass across categories, that is, how each category gains or loses relative shares compared to the others as a consequence of treatment. An intuitive way of building the counterfactual would be to assume that the relative growth/decline in shares of each category in the control group reflects what would have happened in the treated group without treatment. The compositional difference operator $ (\ominus) $ defined in equation (ref) provides the precise mathematical language to formalize this intuition. I now show that the parallel growths assumption can be interpreted using that intuition. Specifically, assumption (ref) implies that, in the absence of treatment, the relative growth or decline of category shares would have been similar across the treated and control groups. This interpretation is formally established in the following proposition.
The equality states that these two compositional changes are identical. In essence, the underlying forces that caused shares in the control group to be reallocated in a particular way would have produced the exact same pattern of reallocation within the treated group. This provides a clear, distributional interpretation of the parallel growths assumption.
Geometric interpretation: Parallel growths also imply parallel trajectories of shares in the probability simplex. shares lie in the simplex $\mathcal{S}^{p-1}$, where movement can be meaningfully defined using the geometry introduced by aitchison1982statistical. Unlike the Euclidean case, this geometry accounts for the simplex’s constraints and forms the basis of compositional data analysis, redefining operations such as addition, scaling, and distance.
The compositional difference between two distributions $\pi_1$ and $\pi_2$ can be redefined as: \[ \pi_1 \ominus \pi_2 = \pi_1 \oplus (-1 \odot \pi_2). \] This formulation highlights that $\ominus$ works like a genuine vector difference in this vector space: it represents the perturbation needed to move from $\pi_2$ to $\pi_1$. In other words, it captures the direction and magnitude of the adjustment between the two points. Viewing $\ominus$ this way makes the operator’s mathematical structure natural. Under this structure, parallel growths implies that the movements of the two PMFs correspond to trajectories that remain at a constant separation without intersecting, having the same direction. Figure (ref) illustrates this idea: in the Aitchison geometry of the simplex, the trajectories of two distributions evolve in parallel within the 2-dimensional simplex $\mathcal{S}^2$, representing all valid probability distributions over three categories.
In figure (ref), each point in the triangle represents a distribution over three categories. As time evolves, distributions trace curves within the simplex. Although these paths may look non-linear in Euclidean space, in the simplex geometry, they are parallel, maintaining a consistent direction without intersecting. Here, I say that the implied evolution of the distributions is consistent with the nature of their space. Because of that, the compositional treatment effect on the treated can be written as:
This expression mirrors the structure of the standard average treatment effect on the treated (ATT) in difference-in-differences, but it operates in the probability simplex rather than in the real line. In the classical DiD, differences in means capture changes in average outcomes across time and groups. In contrast, the compositional DiD captures differences in entire categorical distributions: each component of the treated group’s pre-treatment shares is rescaled by the corresponding ratio of control group shares, and the resulting vector is normalized to yield a valid probability distribution. Thus, (CTT) extends the ATT from the scalar outcome space to the compositional space, preserving the relative structure of category shares while identifying how treatment shifts the allocation of probability mass across categories. This also helps formalize the geometric implication of the assumption.
To sum up, all these perspectives reinforce the idea that imposing the parallel growths assumption on the quantities leads to desirable and interpretable restrictions on the evolution of category shares and totals. It links the economic intuition of stable evolution of preferences with a geometric structure of parallel movements in the probability simplex.
When multiple pre-treatment periods are available, the plausibility of the parallel growths assumption can be evaluated by examining the pre-treatment trajectories of the log-quantities. If the estimated pre-treatment trends appear non-parallel, this raises concerns about the validity of the assumption. Rather than discarding the design altogether, one can relax the assumption in a way that delivers informative bounds on the treatment effect. In particular, I adopt a relaxation strategy closely related to the approach in ban2022robust. They relax the parallel trends assumption by allowing the counterfactual difference between treated and control units to lie within the convex hull of the differences observed in the two most recent pre-treatment periods ($t = 0$ and $t = -1$). By taking the convex hull of the last two differences, they are saying: the counterfactual evolution cannot jump outside the range spanned by recent dynamics. If the differences between groups have been changing slowly, then the most recent differences provide a plausible envelope for what would have happened in the absence of treatment. This is weaker than requiring parallelism and rules out wild deviations inconsistent with observed history. Analogously, in our setting, I extend this idea by applying the same type of convex-hull relaxation. I consider a more general case, where one can either only use the two most recent periods, all the past periods, or all the past periods with higher weights to the most recent ones. I keep the same notation as before, but now, we have that $t \in \{-T_1,...,0, 1\}$. First, define the pre-treatment log-differences between groups for each category \(k = 1, \dots, p\) and take their minimum and maximum. \[ d^{\min}_k = \min_{t \leq 0} \Bigg(\log\left(q^0_{k,1,t}\right) - \log\left(q^0_{k,0,t}\right) \Bigg), \quad d^{\max}_k = \max_{t \leq 0} \Bigg(\log\left(q^0_{k,1,t}\right) - \log\left(q^0_{k,0,t}\right) \Bigg) \]
This assumption relaxes the strict parallel growth requirement by allowing the post-treatment log-change to vary within the historical range observed in pre-treatment periods. It ensures that the counterfactual outcome for each category remains plausible, respecting the variability seen in the pre-treatment trends. When \(d^{\min}_k = d^{\max}_k\), the pre-treatment trends are exactly parallel, and the standard parallel growth assumption is recovered. Practically, the assumption can be implemented in three steps: compute pre-treatment differences, find their minimum and maximum, and constrain the post-treatment log-change to lie within this interval. This approach yields bounds on counterfactual quantities rather than point estimates, providing a transparent measure of uncertainty when strict parallel growth is questionable. It preserves the interpretability of the DiD framework, as the treated group's counterfactual is still anchored to the evolution of the control group, but now allows for more flexible deviations observed historically. It is also possible to add weights for each time period to determine how much each pre-treatment period contributes to this range. In practice, more recent pre-treatment periods can receive greater importance. The identified set is the collection of all quantities that are consistent with the data and the model assumptions.
Theorem (ref) establishes sharp bounds on both the growth treatment effect on the treated and the compositional treatment effect on the treated by allowing for deviations from parallelism between treated and control groups. The CTT bounds describe a feasible set within the simplex of compositions, capturing all possible redistributions consistent with those counterfactuals. The choice of \(\omega_t\) governs the degree of relaxation: uniform weights yield a fully agnostic specification based solely on pre-treatment extremes, while time-increasing weights change the bounds by prioritizing near-treatment dynamics.
Early voting allows registered voters to cast their ballots in person or by mail before Election Day. I evaluate the effect of this policy on voter decisions using the compositional difference-in-differences (CoDiD) framework. A credible analysis requires comparing states with similar political, social, and demographic characteristics so that observed differences in voter participation can be attributed to early voting rather than to preexisting trends. The treatment group includes Maryland and New Jersey, which introduced early voting before 2008, while Pennsylvania and New York serve as the control group. These states are geographically close, share comparable demographic and political profiles, and consistently supported Democratic presidential candidates from 1992 to 2004, indicating similar partisan alignment and voting behavior during the pre-treatment period. Table (ref) reports their racial and ethnic composition around 2007 to further contextualize the comparison.
This table shows that the two groups have almost similar racial compositions. Combined with their geographic proximity and political alignment over the study period, this makes them well-suited for causal analysis. The outcome of interest is the voter's support across three categories: Democrats, Republicans, and Others. I consider that from the 1992 elections to 2008. These data are publicly available and consistently reported across states. I start by plotting the log-counts to assess the plausibility of parallel growths.
Figure (ref) displays the evolution of log transformations of each party vote counts across 5 presidential elections. The dotted red line denotes the pre-treatment elections, while the purple line indicates the post-treatment election. In the pre-treatment period, both states exhibit broadly parallel trajectories in log-counts, suggesting comparable underlying voter dynamics before the policy change. This visual similarity reinforces the plausibility of the identification strategy. In this empirical illustration, neither parallel trends on raw proportions nor parallel trends on quantity data provide an appropriate comparison, as evidenced by figure (ref) and figure (ref)
My method appears to be the most appropriate in this context, as it relies on the visual assessment of parallel trends between the treatment and control groups. The observed similarity in pre-treatment trajectories provides reassuring evidence that the identifying assumptions are reasonable, thereby strengthening the credibility of the causal analysis.
Empirical findings I begin by examining how the treatment affected the total number of votes cast for each party. Table (ref) reports the estimated Growth Treatment Effects (GTT), that is, the proportional increase in votes due to the treatment, along with 95% confidence intervals obtained via the bootstrap procedure described in appendix (ref).
Overall, the policy increased turnout by around $4.4\%$, demonstrating that early voting can mobilize more voters. This finding is consistent with the existing literature on the effects of early voting. All three categories experienced a statistically significant increase in votes, although the magnitude of the effect varies considerably. The Democratic Party saw a modest 5.5% increase, while the Republican Party experienced a smaller rise of about 2.1%. In contrast, support for Other parties, which includes third-party and independent candidates, grew by nearly 37.5%, representing a much larger relative gain. Importantly, because “Other” parties started from a much smaller base, even a modest absolute increase translates into a large percentage change. Focusing on the two major parties, the Democratic Party benefited more from the policy than the Republican Party. This aligns with the results in berry2025selective, which also found that such policies tend to favor Democrats over Republicans. However, I exercise caution in interpreting these results, as the states I selected were already Democratic-leaning during this period. The next table (ref) gives us what happened with the shares.
Table (ref) reports the observed distribution of voter support in the 2008 elections and the estimated counterfactual distribution implied by Theorem (ref). The results suggest that the early voting policy slightly increased support for the Democratic Party relative to what we would expect in the absence of treatment: the share of Democratic support is estimated at $0.5915$ but would have been $0.5823$ without treatment. However, Republican support would have been higher in the absence of treatment ($0.4076$ counterfactual vs. $0.3959$ observed), indicating a shift away from Republicans toward Democrats, suggesting that the main absolute redistribution of probability mass occurred between the two major parties. This is consistent with some of the results obtained in berry2025selective. Interestingly, support for third parties also increased, suggesting that the treatment may have mobilized some voters outside the two major parties as well. Compared to the major parties, the outside option experienced the largest increase relative to its initial share. The next table ((ref)) presents the results for the compositional treatment effects (CTT).
Each entry represents the post-treatment probability of selecting a given party, under a counterfactual scenario where, before treatment, all parties were equally likely (i.e., each had a baseline probability of \(1/3\)). After the treatment: the probability of voting for “Other” rises to 38.7%, up from the baseline of 33.3%. The probability of voting Republican falls to 30.0%. The probability of voting Democratic also declines slightly, to 31.4%, but remains higher than that of Republicans. Thus, while both major parties lose relative appeal, Republicans lose more ground than Democrats, and third-party candidates gain the most in relative terms. This finding is interesting because most prior work focuses exclusively on the two major parties. My analysis reveals that policy interventions can meaningfully affect the electoral relevance of minor parties, a dimension often overlooked in the literature. It is worth emphasizing that this shift is purely relative: the overall ranking of parties by actual vote share in the treated states remains unchanged (Democrats first, Republicans second, Others third). Hence, the treatment does not overturn the existing political hierarchy, but it does narrow the gap between major and minor parties in terms of voter consideration.
This paper introduces Compositional Difference-in-Differences (CoDiD), a causal inference framework for analyzing how treatments and policies affect entire vectors of categorical quantities and their distributions. CoDiD is particularly suited to settings with discrete, unordered outcomes, such as employment status (employed, unemployed, out of the labor force), voting choices, or health categories, where the policy-relevant question is not just whether an average has changed, but how the composition of all categories and the total quantity have evolved. The framework simultaneously captures absolute changes in totals and relative shifts in shares, allowing researchers and policymakers to identify which categories expanded or contracted and by how much, in both scale and composition. Identification relies on a parallel growth assumption, a multiplicative analog of parallel trends, which posits that, absent treatment, category ratios evolve proportionally over time. This assumption has a natural behavioral interpretation in random utility models, mapping estimated effects to shifts in underlying preferences. CoDiD is practical and flexible, extending to settings with relaxed identifying assumptions. Future work may incorporate continuous covariates to adjust for confounding and estimate heterogeneous treatment effects across subpopulations as well as staggered treatment adoption.