Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
47,587 characters · 22 sections · 29 citation commands
Compositional Synthetic Controls
\ifblind
\else
\fi
This paper develops a synthetic control estimator for compositional outcomes, vectors of shares generated by an underlying categorical process. Derived from a random utility model with interactive fixed effects on relative systematic utilities, the estimator maps compositions to log-odds, where the standard convex hull condition identifies the counterfactual as a convex combination of donor log-odds. Equivalently, it recovers the Fr\'{e}chet barycenter under the Aitchison metric, the canonical geometry of the simplex (the non-linear space of shares) using a single set of weights across all categories. I also developed a placebo inference procedure based on the Aitchison distance. An application to Pennsylvania's electricity generation mix following the Alternative Energy Portfolio Standard uncovers a large and persistent compositional shift: natural gas exceeds its counterfactual by nearly 60 percentage points by 2022, while renewables lose relative ground.
\noindentKeywords: Synthetic control, categorical outcomes, compositional data, Aitchison geometry, discrete choice.
\noindentJEL: C21, C25, C43.
\setcounter{page}{1}
Many policy interventions operate on outcomes that are categorical at the individual level. A job-training program moves workers across employment states, permanent contract, temporary contract, and unemployed. An education reform redirects students among fields of study, sciences, humanities, and social sciences. A transportation subsidy shifts commuters between car, public transit, and bicycle. When such a policy is applied to a single aggregate unit, a state, a region, or a country, the researcher observes the resulting vector of category shares, a composition on the simplex. Alternatively, compositional outcomes may be observed directly at the aggregate level, GDP shares by sector, vote shares by party, electricity generation shares by sources, without access to underlying individual data.
The synthetic control method abadie2003economic, abadie2010synthetic has become a leading approach for causal inference in settings with a single treated unit, a pool of untreated donors, and a well-defined intervention date. However, extending it to outcomes that are compositions, vectors of shares generated by an underlying categorical process, poses three fundamental challenges.
First, multiplicity. Running a separate synthetic control for each category yields a different weight vector per category and fitted shares that need not sum to one. No single counterfactual unit exists. Because the categories arise from a common choice process, individuals select among the same alternatives, a single weight vector that exploits the shared latent structure is the principled response tian2023synthetic, sun2025using.
Second, geometric mismatch. Shares live on the simplex, where ratios and relative proportions, not arithmetic differences, characterize economic structure. The Euclidean distance underlying standard synthetic controls is blind to this. Consider three hypothetical regions allocating GDP across services, industry, and agriculture. Region A allocates $(82.5\%, 15.0\%, 2.5\%)$, Region B $(72.5\%, 25.0\%, 2.5\%)$, and Region C $(92.5\%, 5.0\%, 2.5\%)$. The Euclidean distance from A to B and from A to C is identical, approximately $0.14$, suggesting equally good matches. But the services-to-industry ratio is 5.5 in A, 2.9 in B, and 18.5 in C. In terms of economic structure, B is far closer to A. Euclidean distance entirely misses this.
Third, scale dependence. The Euclidean objective penalizes absolute deviations, not proportional ones. As a result, large relative errors in small-share categories can have little influence on the optimization, whereas modest relative errors in dominant categories may drive the fit. Standard synthetic control methods therefore, tend to overlook small-share categories, even when their relative evolution is of primary substantive interest, such as emerging renewable energy technologies, minority employment groups, or rare modes of transportation.
This paper develops Compositional Synthetic Controls (CSC), an estimator that resolves all three problems with a single set of weights derived from the structural model that generates the categorical data. The argument has four steps.
Step 1: Micro-foundation. Under a random utility model of individual choice, aggregate shares are determined by relative systematic utilities, and the log-odds of each category against a baseline equal these utilities exactly. The log-odds of the observed shares are the structural preference parameters.
Step 2: Factor structure. I impose an interactive fixed effects model on the relative systematic utilities, the same structure that motivates the standard synthetic control estimator, but applied to the primitive objects governing choice rather than to the aggregate outcome directly.
Step 3: Identification. Under a convex hull condition on the factor loadings, the treated unit's counterfactual log-odds are identified as a convex combination of donor log-odds (Proposition 1). The corresponding counterfactual composition is the unique Fr\'{e}chet barycenter of the donor compositions under the Aitchison metric.
Step 4: Estimation. The weights are estimated by minimizing the pre-treatment Aitchison distance. By stacking the log-odds equations across all pre-treatment periods, the estimation reduces to a convex quadratic program, following the multi-outcome stacking principle of sun2025using. The resulting counterfactual is automatically a valid composition, positive shares summing to one.
The structural derivation from discrete choice yields two advantages that are absent from alternatives. First, the Aitchison metric treats all pairwise log-ratios symmetrically: a doubling of a 2% share contributes as much to the objective as a doubling of a 30% share. I formalize this in Proposition (ref), which shows that the Euclidean SCM objective is a share-weighted distortion of the Aitchison objective. Second, the donor weights carry a direct behavioral interpretation as measures of preference similarity: a donor earns weight to the extent that its population's relative preferences track those of the treated unit before the intervention. The treatment effect can therefore be interpreted as the change in the relative attractiveness of each category induced by the policy.
I illustrate the method with an application to Pennsylvania's electricity generation mix following the Alternative Energy Portfolio Standard (AEPS, Act 213 of 2004). The estimated effects reveal a sustained compositional shift: the natural-gas share exceeds its counterfactual by nearly 60 percentage points by 2022, while renewables, despite growing in absolute terms, lose ground to gas on the relative scale. A componentwise analysis would miss this interplay; the compositional counterfactual captures it coherently.
This paper connects three literatures. In the synthetic control literature, the foundational contributions are abadie2003economic and abadie2010synthetic, with extensions to regularization doudchenko2016balancing, abadie2017penalized, matrix completion athey2018matrix, and design-based inference chernozhukov2018exact. The most closely related methodological work is gunsilius2023distributional, who extends synthetic controls to distributional outcomes using optimal transport, and gunsilius2024tangential, who proposes tangential interpolation for manifold-valued outcomes. CSC differs in three respects: the geometry is not imposed abstractly but derived from a structural model of choice; the Aitchison metric has an explicit behavioral interpretation as log-odds matching; and the estimator reduces to a standard convex quadratic program rather than requiring manifold-specific optimization. tian2023synthetic and sun2025using argue for a common weight vector when components share latent factors, which is the structural rationale this paper formalizes. The present paper adopts the stacking principle of sun2025using: because the $p - 1$ log-odds equations share the same unit-level factor loadings, they can be stacked into a single regression, expanding the effective pre-treatment sample from $T_0$ to $(p-1)T_0$ and enabling consistency through the number of categories as well as the number of pre-treatment periods.
In compositional data analysis, the foundational framework is due to aitchison1982statistical, with the Hilbert space structure developed by egozcue2003isometric. CSC embeds the Aitchison geometry within a treatment-effects framework grounded in a structural model, providing both identification conditions and a behavioral interpretation that the statistical CoDA literature does not offer.
In the geodesic and manifold literature, zhu2023geodesic and kurisu2024geodesic build synthetic controls for general unique geodesic spaces. This paper trades generality for structure: specializing to the simplex equipped with the Aitchison metric allows the geometry to carry economic content, the distance between compositions equals the distance between the preference structures that generated them.
Section (ref) defines the data structure and estimands. Section (ref) develops the model, states the main identification result (Theorem 1), and presents the estimator and implementation algorithm. Section (ref) compares CSC to alternative estimators. Section (ref) discusses inference. Section (ref) reports the Pennsylvania application. Section (ref) concludes. Proofs and secondary results are in the Appendix.
There are $J$ units indexed by $j = 1,\dots,J$, observed over $T$ periods $t = 1,\dots,T$. Unit $j=1$ is treated beginning at a known date $T_0$, with $1 < T_0 < T$. Units $j = 2,\dots,J$ form the donor pool and are unaffected throughout. For each unit and period we observe the aggregate shares across $p$ mutually exclusive and exhaustive categories,
where $\mathcal{S}^{p-1} = \{\pi \in \mathbb{R}_{+}^{p} : \sum_{k=1}^{p} \pi_{k} = 1\}$ is the probability simplex. These shares may be computed from individual-level categorical data, as when a census records each person's employment sector and the researcher aggregates to state-level shares, or observed directly as compositional measures at the aggregate level. In the potential-outcomes notation, $\pi_{j,t}^{0}$ and $\pi_{j,t}^{1}$ denote the share vectors absent and under treatment, respectively. For donor units, $\pi_{j,t}^{0} = \pi_{j,t}^{1}$ for all $t$. For the treated unit, $\pi_{1,t}^{0} = \pi_{1,t}^{1}$ for $t \leq T_0$. The object of interest is the counterfactual $\pi_{1,t}^{0}$ for $t > T_0$.
Assumption (ref) ensures that log-ratios are well-defined. It can be relaxed in practice by adding a small positive constant to zero shares, a standard device in compositional data analysis aitchison1982statistical. With a counterfactual $\hat{\pi}_{1,t}^{0}$ in hand, two families of estimands are available.
\paragraph{Share-level ATT.} The level change in each category's share,
Since both vectors sum to one, $\sum_{k} \mathrm{ATT}_{k,t} = 0$: what one category gains, the others jointly lose. This is the compositional analogue of the standard ATT, expressed in percentage points.
\paragraph{Log-ratio treatment effect (LRTE).} The change in relative odds of category $k$ versus category $\ell$,
A value $\mathrm{LRTE}_{k\ell,t} > 0$ means treatment raised category $k$ relative to category $\ell$. The LRTE isolates the treatment-induced redistribution of mass between any pair of categories and is the natural estimand under the simplex geometry. A scalar summary of total compositional displacement is the Aitchison distance between observed and counterfactual compositions,
This section develops the method from its behavioral foundations. I begin with the model of individual choice that generates aggregate shares, introduce a factor structure on the primitive parameters of that model, state the main identification result, and present the estimator with a detailed implementation algorithm.
Following mcfadden1972conditional, I consider the following model. The aggregate shares $\pi_{j,t}$ summarize the choices of individuals in unit $j$ at time $t$. Suppose that in each unit and period there is a large population of decision-makers, each selecting one of the $p$ mutually exclusive alternatives. Individual $i$ in unit $j$ at time $t$ faces the random utility
where $V_{k,j,t}^{0}$ is the systematic utility of alternative $k$, common to all individuals in the unit, and $\varepsilon_{ik,j,t}$ is an idiosyncratic taste shock. When the shocks are independently and identically distributed across individuals and alternatives following the type-I extreme value distribution, the probability that a randomly drawn individual selects alternative $k$ is given by the multinomial logit formula,
The systematic utilities are identified only up to a common additive constant. Choosing category $p$ as the baseline, define the relative systematic utility of alternative $k$,
From the logit formula, the log-odds of category $k$ against the baseline exactly equal the relative systematic utility:
This is the key structural equation: the observable log-odds are the preference parameters that govern choice. Define the log-odds map $\ell\colon \mathrm{int}(\mathcal{S}^{p-1}) \to \mathbb{R}^{p-1}$ by
with inverse $\ell^{-1}(y)_k = e^{y_k}/(1 + \sum_{m=1}^{p-1} e^{y_m})$ for $k < p$ and $\ell^{-1}(y)_p = 1/(1 + \sum_{m=1}^{p-1} e^{y_m})$. Equation (ref) states that $\ell(\pi_{j,t}^{0}) = \tilde{V}_{j,t}^{0} = (\tilde{V}_{1,j,t}^{0}, \dots, \tilde{V}_{p-1,j,t}^{0})$.
The standard synthetic control method is motivated by a factor model for the outcome variable. The idea, formalized by abadie2010synthetic, is that the outcome is generated by an interactive fixed effects structure in which unobserved unit-specific factor loadings interact with time-varying factors, and the condition for identification is that the treated unit's loadings lie in the convex hull of the donor loadings. I adopt exactly this structure, but applied to the primitive objects that generate compositional data: the relative systematic utilities.
Two features of this specification deserve emphasis. First, the factor loadings $\mu_j$ are common across all $p-1$ relative utility equations within a unit. They represent persistent features of the unit's preference environment, demographics, industry composition, regulatory history, that shape the relative attractiveness of all alternatives simultaneously. A state with a high loading on a “green preferences” factor will display higher relative utilities for renewables over fossil fuels, for public transit over driving, and for other environmentally favorable alternatives, consistently across categories and over time. The factors $\lambda_{k,t}$ are common across units but vary by category and time, allowing the same latent trait to have different implications for different alternatives in different periods.
Second, the factor model is imposed on relative utilities, not on shares. Because the log-odds map is nonlinear, a factor model on $\tilde{V}$ does not imply a factor model on $\pi$, and conversely. The factor structure on utilities is the economically primitive specification: it models the latent determinants of choice, not their nonlinear aggregation into probabilities.
Stacking the $p-1$ equations, let $\bm{\delta}_t = (\delta_{1,t},\dots,\delta_{p-1,t})' \in \mathbb{R}^{p-1}$, let $\Theta_t$ be the $(p-1) \times r$ matrix with rows $\theta_{k,t}'$, let $\Lambda_t$ be the $(p-1) \times F$ matrix with rows $\lambda_{k,t}'$, and let $\bm{\varepsilon}_{j,t} = (\varepsilon_{1,j,t},\dots,\varepsilon_{p-1,j,t})'$. Then
The factor model provides the primitive justification for the synthetic control approach. The following condition is the compositional analogue of the standard convex hull requirement.
Assumption (ref) requires that the treated unit's factor loadings lie in the convex hull of the donor units' loadings. It states that the treated unit's persistent preference structure, the unobserved features shaping the relative attractiveness of all alternatives, can be expressed as a mixture of the donors' preference structures. This is the same condition that motivates the standard synthetic control estimator, transposed to the space of preference parameters.
The main result of the paper is the following.
The weights $w^{\ast}$ are estimated by minimizing the pre-treatment discrepancy between the treated unit's log-odds and the synthetic combination. The CSC weights solve
This is a convex quadratic program in $J-1$ unknowns with a simplex constraint, solvable with standard routines. By the isometry between $(\mathcal{S}^{p-1}, \delta_A)$ and $(\mathbb{R}^{p-1}, \|\cdot\|_2)$, the objective is equivalently $\sum_{t=1}^{T_0} \delta_A^{2}(\pi_{1,t},\, \bigoplus_{j \geq 2} w_j \odot \pi_{j,t})$.
\paragraph{Stacked representation.} Because $\ell(\pi_{j,t}) \in \mathbb{R}^{p-1}$, the squared norm in (ref) decomposes across the $p - 1$ log-odds coordinates:
This is a single least-squares problem with $N \coloneqq (p-1) T_0$ scalar observations, each of the form
As emphasized by sun2025using in the context of multi-outcome synthetic controls, the common factor loadings $\mu_j$ shared across all $p - 1$ equations within a unit mean that each log-odds coordinate provides independent information about the same unit-level parameters. The stacking aggregates this information: rather than estimating $w^{\ast}$ from $T_0$ vector-valued observations, one estimates it from $(p-1)T_0$ scalar observations drawn from the same underlying factor model.
Given the estimated weights, the counterfactual shares for any post-treatment period $t > T_0$ are
with $\hat{\pi}_{p,1,t}^{0} = 1 - \sum_{k=1}^{p-1} \hat{\pi}_{k,1,t}^{0}$.
The proof, given in Appendix (ref), applies standard extremum estimator arguments to the quadratic objective on the compact weight simplex. The key observation is that the effective sample size for the stacked regression (ref) is $N = (p-1)T_0$, not $T_0$ alone. Consistency can therefore be achieved in two regimes: the standard long-panel regime where $T_0 \to \infty$ with $p$ fixed, and the many-categories regime where $p \to \infty$ with $T_0$ fixed. In the application below, $p = 3$ and $T_0 = 14$, so $N = 28$. The uniqueness condition rules out collinear donor configurations and is standard in the synthetic control literature.
Algorithm (ref) provides a step-by-step implementation of CSC. The algorithm requires only the share data and a standard quadratic programming solver; the R function solve.QP from the quadprog package or the Python function scipy.optimize.minimize with the SLSQP method are both suitable. A replication package implementing Algorithm (ref) in R with the Pennsylvania data is available online.
To clarify what CSC achieves, I contrast it with two natural alternatives.
\paragraph{Separate SCM.} Running a standard synthetic control independently for each category,
yields $p$ distinct weight vectors. The fitted shares $\sum_{j} w_j^{(k)} \pi_{k,j,t}$ generally do not sum to one, so no single counterfactual exists. Moreover, the method ignores the joint structure: the same latent factors that determine relative utilities for category $k$ also determine those for category $\ell$, and separate estimation discards this information.
\paragraph{Euclidean SCM.} A single weight vector can be obtained by minimizing Euclidean distance on raw shares:
The fitted shares sum to one because the simplex is convex under arithmetic means. However, the Euclidean metric introduces a systematic scale dependence that I now formalize.
The proof is in Appendix (ref). This matters whenever the policy question involves small-share categories. The Euclidean estimator matches the dominant categories at the expense of precisely those alternatives the policy targets.
I propose a placebo permutation test following abadie2010synthetic, adapted to the Aitchison metric (Step 6 of Algorithm (ref)). Each donor unit $j = 2,\dots,J$ is treated in turn as if it had received the intervention at $T_0$. For each placebo, the CSC weights are re-estimated using the remaining donors, and the post-to-pre root mean squared prediction error ratio is computed:
where
and $\mathrm{RMSPE}_{\mathrm{post}}^{(j)}$ is defined analogously over $t = T_0+1,\dots,T$. Using the Aitchison distance ensures consistency with the estimation objective.
To avoid comparing the treated unit to donors with poor pre-treatment fit, I retain only those units whose pre-treatment RMSPE does not exceed a multiple $m$ (typically $m = 5$) of the treated unit's pre-treatment RMSPE. Let $M$ be the number of retained units, including the treated unit. Under the sharp null of no treatment effect for any unit, the test statistic $R^{(1)}$ is exchangeable with the placebo statistics, so the exact $p$-value is
I use annual state-level net generation from the U.S.\ Energy Information Administration for 1990--2023, aggregated into three categories: natural gas; coal and oil; and renewables (conventional hydro, wind, solar, geothermal, and other).
The treated unit is Pennsylvania and the treatment date is $T_0=2003$, the year before the AEPS (Act 213 of 2004), giving fourteen pre-treatment and twenty post-treatment periods. The donor pool of 42 states excludes units with a structural zero in any category over the sample, states that adopted comparable alternative energy portfolio standards during the sample (Ohio in 2008, West Virginia in 2009), and Vermont, whose gas and fossil shares are each below $0.1\%$ throughout and would produce extreme log-odds.
(ref) reports the CSC weights and balance. Nine donors receive positive weight, led by coal-intensive states, Illinois ($0.233$) and Wyoming ($0.208$), with Iowa ($0.168$), Nebraska ($0.149$), Massachusetts ($0.102$), and smaller contributions from Connecticut, New Jersey, Indiana, and Montana. The synthetic control reproduces the pre-treatment mix almost exactly. The Aitchison pre-treatment RMSPE is $0.187$ and the Euclidean share RMSPE is $1.34$ percentage points. (ref) shows the graphical evolution of the shares.
(ref) reports the estimated effects. The natural-gas share rises steadily above the counterfactual, reaching a peak gap of $60$ percentage points in 2022, while coal and oil fall by a corresponding amount; the renewables gap is mildly negative in later years, so Pennsylvania's renewable share grew more slowly than the synthetic control's even as its gas share surged. The gas-versus-fossil log-ratio effect reaches $2.37$ by 2022, the observed gas-to-fossil ratio is about $\exp(2.37)\approx10.7$ times its counterfactual, and the Aitchison distance between observed and counterfactual compositions grows throughout, reaching $2.53$ in 2022.
(ref) reports the placebo distribution and (ref) the placebo gaps. Among the $35$ donors retained under the pre-fit screen, three attain a post-to-pre Aitchison RMSPE ratio at least as large as Pennsylvania's ($R^{(1)}=9.18$): Michigan ($10.46$), Louisiana ($10.22$), and Florida ($10.07$), giving an exact $p$-value of $4/36=0.111$. The nearest placebos are states that underwent their own gas-driven transformations in the 2010s; Pennsylvania still stands apart partly because its shift began in 2004, several years before the shale boom reached most non-adopting states, and went further.
The estimates describe a compositional transition that began at enactment and widened for nearly two decades, concentrated in the substitution of natural gas for coal and oil. Because the synthetic control is built overwhelmingly from coal-intensive states that experienced the same national gas-price decline, the estimated effect is the additional compositional shift in Pennsylvania beyond that common movement. The mildly negative renewables effect is itself informative: the policy's headline renewables target was modest next to the gas-driven transformation of the fossil base, so on the relative scale the simplex makes visible, renewables lost ground to gas even as they grew in absolute terms. A single-share analysis would miss this interplay; the compositional counterfactual captures it coherently.
This paper developed Compositional Synthetic Controls, a method for constructing counterfactual share vectors when the outcome is categorical or compositional and a single aggregate unit receives an intervention. The method rests on a structural model of discrete choice. Under the logit specification, the log-odds of shares equal relative systematic utilities; an interactive fixed effects model on these utilities delivers identification via the standard convex hull condition in log-ratio space (Theorem (ref)). The counterfactual is the Fr\'{e}chet barycenter of the donor compositions under the Aitchison metric: a valid composition by construction, with weights that measure preference similarity. Adopting the multi-outcome stacking principle of sun2025using, the $p - 1$ log-odds equations are pooled into a single regression with effective sample $(p-1)T_0$, ensuring that consistency requires only the product of the number of categories and pre-treatment periods to grow.
CSC has three central advantages over standard Euclidean approaches. It produces a single synthetic unit with a coherent counterfactual composition; it respects the log-ratio structure of the simplex, ensuring that relative movements and small-share categories receive appropriate weight; and the donor weights carry an interpretable behavioral reading as measures of structural similarity in the latent preference parameters. The Pennsylvania application demonstrates these properties in a policy setting that directly targets composition.
Three extensions are natural. First, covariates can be incorporated by augmenting the log-odds objective with predictor balance terms, analogous to the covariate-adjusted SCM of abadie2010synthetic. Second, penalization, ridge or LASSO on the weights, can improve finite-sample performance when the donor pool is large relative to the effective sample $(p-1)T_0$. Third, conformal inference methods chernozhukov2018exact adapted to the Aitchison metric could yield pointwise confidence sets with exact coverage, strengthening inference beyond the permutation approach.
The author reports no conflicts of interest.
The electricity generation data used in Section (ref) are publicly available from the U.S.\ Energy Information Administration. A complete replication package containing all code and data files is available.
The author used generative AI (Claude Opus 4.6) for language editing, LaTeX formatting, and replication code generation. The author reviewed and is responsible for all final content.