Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
72,305 characters · 17 sections · 48 citation commands
A Synthetic Control Approach to Conditional Distributional Treatment Effects
\thispagestyle{empty}
\setcounter{page}{1}
The synthetic control (SC) method, introduced by AbadieGardeazabal2003 and formalised in AbadieEtAl2010, has become one of the most widely used tools for the evaluation of policy interventions in settings with a small number of treated units; Abadie2021 provides a comprehensive review. The classical estimator constructs a weighted average of control unit outcomes that approximates the pre-treatment trajectory of the treated unit, and uses this synthetic unit as the counterfactual in the post-treatment period. DoudchenkoImbens2016 relax the non-negativity constraint on weights, allowing the synthetic control to lie outside the convex hull of the donor pool, and show that this flexibility improves pre-treatment fit. BenMichaelEtAl2021 propose an augmented SC estimator that combines outcome-model bias correction with SC weighting. ChernozhukovEtAl2021 develop an exact, finite-sample valid inference approach for SC based on conformal prediction. FermanPinto2021 study the properties of SC when pre-treatment fit is imperfect. ArkhangelskEtAl2021 develop a synthetic difference-in-differences estimator that accommodates heterogeneous treatment effects. Chen2023 establishes a connection between synthetic control and online learning, showing that the SC estimator can be interpreted as an instance of follow-the-leader, which yields oracle inequalities for counterfactual predictions even under adversarial outcomes. While powerful for estimating average treatment effects at the aggregate level, all these methods provide no information about distributional consequences, a critical limitation for policy questions involving inequality, minimum wages, or tax reforms.
The broader literature on distributional treatment effects includes the counterfactual wage decompositions of MachadoMata2005 and Melly2005 and ChernozhFVMelly2013, the unconditional quantile regressions of FirpoFortinLemieux2009, and the distributional DiD approaches of ChernozhFVMelly2013 and Fernandezval2026. None of these papers considers a synthetic control approach.
A first important extension was proposed by Gunsilius2023, who replaces scalar outcomes with quantile functions, thereby allowing inference on heterogeneous treatment effects across the distribution. His estimator matches unconditional quantile functions of the control units to that of the treatment unit in the pre-treatment period, using a constrained quantile-on-quantile regression. While this distributional synthetic control (DSC) method is a substantial advance, it is inherently unconditional: It cannot accommodate individual-level covariates such as education, experience or demographics that are central to labour economics and related fields.
The present paper fills this gap by proposing a SC estimator for conditional distribution functions. ChenFeng2026 extend the DSC framework to group-level heterogeneity, but their heterogeneity is driven by unobservable group membership rather than observable covariates, leaving the gap addressed here open. Conditioning on covariates is not simply a refinement of the unconditional approach: it changes the identification problem in a fundamental way. When the outcome distribution depends on individual characteristics, as wages depend on education and experience, the counterfactual object of interest is a conditional distribution function, and any identification assumption must be formulated at the level of the model that generates it.
We work within the semiparametric DR framework of ForesiPeracchi1995, where the conditional distribution function of an outcome $Y$ given covariates $X \in \mathbb{R}^p$ is modeled as $F(y\mid x)=\Lambda(x'\theta(y))$ for a known link $\Lambda$ and unknown parameter function $\theta:\mathcal{Y}\to\mathbb{R}^p$. This model has appealing properties in applications: it handles mass points, irregularities, and multimodality naturally, and estimation reduces to a sequence of binary outcome regressions indexed by $y$. The DR literature includes KneibEtAl2023 and Klein2024 on flexible extensions, spady:2025 on semiparametric efficiency, and Wied2024 on instrumental variable estimation in the DR framework; RotheWied2013 derive the asymptotic theory used here; and BiewenErhardt2025 apply distribution regression to minimum wage analysis. To the best of our knowledge, the combination of synthetic control methods with semiparametric distribution regression has not been studied in the existing literature.
Our identification assumption Parallel Trends in Parameters (PTP) is formulated directly in the DR parameter space: the counterfactual parameter function of the treated group evolves from the pre-treatment period to the post-treatment period along the same weighted combination of trends observed in the control groups, requiring the weights to sum to one. The economic rationale is that the DR parameters (e.g. quantile-specific returns to education and experience in a Mincer framework) are driven by a small number of common aggregate shocks (e.g. skill-biased technological change, business-cycle fluctuations, globalisation) whose differential local impact is captured by state-specific factor loadings. PTP formalises the condition that the factor loadings of a treated unit can be approximated by a weighted combination of donor states' loadings, which is the direct functional analogue of the latent-factor rationale underlying classical SC methods AbadieEtAl2010. This approximation is natural in the DR parameter space rather than at the CDF level, because the parameter function is linear in the common factors whereas the CDF is not. This formulation has also two other important advantages over a parallel trends condition on the CDF level. First, it ensures that the counterfactual conditional distribution function remains within the DR model class, so that all estimators have a direct parametric interpretation. Second, it allows us to formulate the weight estimation problem as a finite-dimensional quadratic program (QP) at each value of $y$, or equivalently as a single infinite-dimensional QP over the function space $\ell^{\infty}(\mathcal{Y})^p$. Dropping the traditional non-negativity constraint yields a closed-form weight estimator, avoids interior-solution requirements, and transparently accounts for potentially negative weights that arise when some control units lie outside the convex hull of the treated unit.
One contribution of the paper is the asymptotic analysis of the proposed estimator for the case of fixed number of groups $J$ and increasing individual observations within each group. This framework fits to micro-data applications such as the Current Population Survey (CPS) for US states, which we consider, as the number of states is small relative to the within-group sample size $n$. Statistical precision derives from this sample size, not from the number of time periods $T$. In fact, one pre-period is sufficient for obtaining consistent parameter estimators, whereas more pre-periods increase the precision of the weight estimators. An important result is the joint weak convergence of the pointwise counterfactual CDF differences to a Gaussian vector, where both DR estimation error and weight estimation error contribute at the same $\sqrt{n}$ rate. Based on this, we show asymptotic normality of the integrated squared treatment effect estimator when the true effect is non-zero. For testing the null hypothesis of no treatment effect, we propose a supremum statistic over the pointwise CDF differences, applicable either on the whole support of $Y$ or a subset. Its null distribution is the supremum of a mean-zero Gaussian process. This supremum statistic has a non-degenerate null distribution, which is approximated by a Gaussian process simulation, and power growing to one with $n$. A plug-in confidence interval for the integrated squared effect is reported in a second step, conditional on rejection. We characterize under which types of deviation from PTP the results still hold.
If multiple pre-periods are available, it is possible to perform pre-trend diagnostics. We use them to perform causal inference in an empirical application to the 1992 New Jersey minimum wage increase. In fact, we see that the DR-SC approach reveals a clear heterogeneity pattern that is invisible in aggregate analyses. The minimum wage effect is sharply concentrated in the corresponding corridor for low-education and low-experience workers, the group most directly affected by the policy. For other groups the effects are at most marginal, in particular for high-education/experience workers.
The paper proceeds as follows. Section 2 introduces the framework for general $T_0\geq 1$ pre-treatment and $T_1\geq 1$ post-treatment periods and discusses point as well as confidence estimation of the treatment effect. Moreover, the supremum test and the pre-trend test for the case $T_0\geq 2$ are presented. Section 3 derives the supporting asymptotic theory. Section 4 presents simulation evidence, which highlights the advantages of the conditional over an unconditional approach. Section 5 contains the empirical application, Section 6 concludes.
Consider $J+1$ groups $i=1,\ldots,J+1$ at $T=T_0+T_1$ time points, where $t=1,\ldots,T_0$ are pre-treatment periods and $t=T_0+1,\ldots,T_0+T_1$ are post-treatment periods. Group $i=1$ is the treatment group; groups $i=2,\ldots,J+1$ form the donor pool of potential control groups. In all pre-treatment periods none of the groups is treated. Only group $i=1$ receives the treatment from period $T_0+1$ onward. Within each group-time cell $(i,t)$ we observe an i.i.d.\ sample of $n_{it}$ individuals with outcome $Y\in\mathcal{Y}\subseteq\mathbb{R}$ and a $p$-dimensional covariate vector $X\in\mathcal{X}\subseteq\mathbb{R}^p$. Typical applications in labour economics have $J$ small (e.g., $J+1=10$ to 50 federal states or countries) and $n_{it}$ large (e.g., several thousand individuals per group-time cell). We fix $J$ and let $n:=\sum_{i,t}n_{it}\to\infty$, assuming $n_{it}/n\to r_{it}\in(0,\infty)$.
For each group $i$ and time period $t$, the conditional distribution function of $Y$ given $X=x$ satisfies
where $\Lambda:\mathbb{R}\to(0,1)$ is a known strictly increasing link function (such as the standard normal CDF for the probit link, or the logistic CDF for the logit link) with derivative $\lambda$ and $\theta_{it}:\mathcal{Y}\to\mathbb{R}^p$ is an unknown c\`adl\`ag parameter function. For fixed $y$, $\theta_{it}(y)$ is the coefficient vector in the binary outcome regression of $\mathbf{1}\{Y\leq y\}$ on $X$. Numerical sensitivity analyses in dette2025 confirm that the choice of the link function has a negligible impact on estimation results.
We are interested in the treatment effect on a region of the outcome distribution $\mathcal{Y}_0\subseteq\mathcal{Y}$, measured via the integrated squared discrepancy between the observed and counterfactual conditional distributions:
where $F^{0}_{1,t}(y\mid x)$ is the counterfactual conditional distribution of the treated group in post-treatment period $t$, and
is the pointwise CDF difference. The integrated discrepancy $f_t(x,\mathcal{Y}_0)$ captures both the magnitude and the shape of the treatment effect on the conditional distribution at covariate value $x$ over the region $\mathcal{Y}_0$. It equals zero if and only if $\Delta_t(y\mid x)=0$ for almost all $y\in\mathcal{Y}_0$. Setting $\mathcal{Y}_0=\mathcal{Y}$ recovers the full integrated effect; a proper subset $\mathcal{Y}_0\subsetneq\mathcal{Y}$ focuses attention on a specific region of the distribution, such as the minimum-wage corridor $[\mathrm{MW}_{\mathrm{old}},\mathrm{MW}_{\mathrm{new}}+\varepsilon]$. One may also consider the doubly integrated measure $\int_{\mathcal{X}}f_t(x,\mathcal{Y}_0)dF_{X}(x)$, which averages over the covariate distribution. If not denoted otherwise, $f_t(x,\mathcal{Y}_0) = f_t(x)$ for notational simplicity in the following.
The fundamental identification challenge is that $F_{11}^{0}(y|x)$ is not observed. We identify it via the following assumption, which has to hold on the set of interest $\mathcal{Y}_0$.
Assumption PTP states that the trend in the treated group's DR parameter function from the pre- to the post-treatment period would have equalled a weighted combination of the control groups' trends, absent treatment. This is a direct analogue of the classical parallel trends assumption in the difference-in-differences (DiD) literature, formulated in the parameter space of the DR model rather than on the level of the outcome or its conditional mean. We impose only the adding-up constraint; negative weights arise when some control units lie outside the convex hull of the treated unit. Differently to the case of negative weights in regression-based DiD CallawayLi2019, this is not a problem of model specification, but a particular feature, which yields flexibility. DoudchenkoImbens2016 make a similar argument in the SC context and show that allowing negative weights improves pre-treatment fit.\\
For each cell $(i,t)$ and $y$ on a grid $\mathcal{Y}_m=\{y_1,\ldots,y_m\}\subset\mathcal{Y}_0$, the DR parameter $\theta_{it}(y)$ is estimated by maximizing the log-likelihood of the binary outcome model. For a fixed $y\in\mathcal{Y}$, the estimator solves:
This is a standard probit or logit regression of $\mathbf{1}\{Y\le y\}$ on $X$, estimated separately for each $y$ on a finite grid $\mathcal{Y}_{m}=\{y_{1},...,y_{m}\}\subset\mathcal{Y}_0$ (e.g. the 0.05,0.1,\ldots quantiles of the distribution of $Y$).
The weights $w=(w_{2},...,w_{J+1})^{\prime}$ are estimated by minimizing the pre-treatment discrepancy averaged over all $T_0$ pre-treatment periods:
By averaging over pre-treatment periods and integrating the quadratic discrepancy over the support of the outcome $y$, we move from a point-wise to a distribution-wide weight estimation. This ensures that the weights $w$ are constant across the entire conditional distribution and across time. This global objective function acts as a natural stabilizer, effectively reducing the dimensionality of the optimisation problem and mitigating the risk of over-fitting that typically arises when weighting high-dimensional objects.
In practice, the integral over $\mathcal{Y}_0$ is replaced by a sum over the grid $\mathcal{Y}_{m}$
This is a quadratic program in $w$ with linear constraints and a positive semidefinite objective. Let $\hat G\in\mathbb{R}^{J\times J}$ denote the time-averaged Gram matrix with entries \[ \hat G_{kl}=\frac{1}{T_0\cdot m}\sum_{t=1}^{T_0}\sum_{l'=1}^m \hat\theta_{kt}(y_{l'})'\hat\theta_{lt}(y_{l'}) \] and $\hat c\in\mathbb{R}^J$ the vector with entries \[ \hat c_k=\frac{1}{T_0\cdot m}\sum_{t=1}^{T_0}\sum_{l'=1}^m\hat\theta_{1t}(y_{l'})'\hat\theta_{kt}(y_{l'}). \] Solving the Lagrangian of (ref) yields the closed-form solution (ref), unchanged in form.
The counterfactual DR parameter for post-treatment period $t\in\{T_0+1,\ldots,T_0+T_1\}$ is estimated as:
which exploits the pre-treatment balance enforced by the weight estimator (ref) (see Remark (ref)). The counterfactual conditional distribution function is then $\hat{F}_{1,t}^{0}(y|x)=\Lambda(x^{\prime}\hat{\theta}_{1,t}^{0}(y))$ and the treatment effect estimator for post-period $t$ is:
approximated in practice by a Riemann sum over $\mathcal{Y}_{m}$.
Since $f_t(x)\ge0$ by construction, we report a one-sided confidence interval. When the treatment effect is non-zero, $f_t(x)>0$, the asymptotic normality of $\hat f_t(x)$ (established in Theorem (ref)(c) in Section (ref)) yields, at significance level $\alpha$, the one-sided $(1-\alpha)$ confidence interval
where $u_{1-\alpha}$ is the $(1-\alpha)$-quantile of the $N(0,1)$-distribution and
and $\hat K_t$ is the covariance kernel for post-period $t$:
The generalized residuals (scores) are $$\psi_{i,t,k}(y) = \dfrac{\lambda(X_{i,t,k}'\hat\theta_{i,t}(y))}{\Lambda(X_{i,t,k}'\hat\theta_{i,t}(y))\,(1-\Lambda(X_{i,t,k}'\hat\theta_{i,t}(y)))}\bigl(\mathbf{1}\{Y_{i,t,k} \le y\} - \Lambda(X_{i,t,k}'\hat{\theta}_{i,t}(y))\bigr),$$ the standardized score residual of the binary log-likelihood (ref). The vector $\hat{\tilde\Theta}_{x,t}(y_l)=\bigl(x'\hat\theta_{2,t}(y_l),\ldots,x'\hat\theta_{J+1,t}(y_l)\bigr)'\in\mathbb{R}^J$ collects the scalar projections of the period-$t$ donor parameters onto $x$, and $\hat{V}_w$ is an estimator for $V_w$ from Theorem (ref). Moreover, $$\widehat{\mathcal{I}}_{i,t}(y)=\frac{1}{n_{i,t}}\sum_k \dfrac{\lambda(X_{i,t,k}'\hat\theta_{i,t}(y))^2}{\Lambda(X_{i,t,k}'\hat\theta_{i,t}(y))\,(1-\Lambda(X_{i,t,k}'\hat\theta_{i,t}(y)))} X_{i,t,k}X_{i,t,k}'$$ is the empirical Fisher information at $y$ (the sample analogue of $\mathcal{I}_{i,t}$ in Lemma (ref)). The first term in (ref) captures DR estimation uncertainty for the treated group in period $t$; the second captures uncertainty from the control groups; the third captures weight estimation uncertainty, corresponding to the $\boldsymbol{\Theta}_t(y)'V_w\boldsymbol{\Theta}_t(y')$ component of (ref). We estimate $V_w$ by the following plug-in estimator. With $\hat P=\partial w^*/\partial c$ at $\hat G$ (see (ref)) and $\hat\Theta^{(s)}(y)=(\hat\theta_{2,s}(y),\ldots,\hat\theta_{J+1,s}(y))'\in\mathbb{R}^{J\times p}$, define the cell loadings $\hat B_{i,s}(y)\in\mathbb{R}^{J\times p}$ (under Assumption (ref))
and the plug-in estimator is the score-outer-product sum
with the score residual $\psi$ and Fisher information $\widehat{\mathcal{I}}$ of (ref); consistency follows from Theorem (ref).
The interval (ref) achieves nominal one-sided $(1-\alpha)$ coverage asymptotically when $f_t(x)>0$. When $f_t(x)=0$, $\hat\sigma^2_t(x)$ degenerates to zero and the lower bound collapses to $0$; testing $H_0:f_t(x)=0$ requires a separate procedure, described in Section (ref).
The null hypothesis $H_0:f_t(x,\mathcal{Y}_0)=0$ is equivalent to $H_0:\Delta_t(y|x)=0$ for (Lebesgue-almost) all $y\in\mathcal{Y}_0$. We test this hypothesis via the supremum of the pointwise CDF differences over $\mathcal{Y}_0$.
The test statistic for post-treatment period $t$ and region $\mathcal{Y}_0$ is
where $\hat\delta_{x,t}(y_l):=\hat F_{1,t}(y_l|x)-\hat F^0_{1,t}(y_l|x)$. The full-distribution test uses $\mathcal{Y}_0=\mathcal{Y}_m$; a focused test restricts to a proper subset such as the minimum-wage corridor. The asymptotic theory applies to both with $m$ replaced by $|\{l:y_l\in\mathcal{Y}_0\}|$.
Critical values are obtained by simulation of the Gaussian process $Z_x$ from Remark (ref):
The $p$-value is $\hat p = S^{-1}\sum_s\mathbf{1}\{T^{(s)}\ge T_n(x)\}$. If Stage 1 rejects $H_0$, the confidence interval (ref) provides quantitative precision for $f(x)$.
When $T_0\geq 2$, evidence for the validity of PTP can be given by pre-testing. As usual in difference-in-difference approaches (Lechner2011), the idea is to extrapolate from the time before the treatment (where all groups are non-treated) to the time after, where the counterfactual outcome of the treated group is not observed.
The tests can be formed in two ways: for a specific $x$ with the statistic (ref), or for all $x$ simultaneously via the modification below. In the focused form, each pre-period $t\in\{1,\ldots,T_0-1\}$ is treated as a pseudo-post period and $T_n(x,t,\mathcal{Y})$ from (ref) is computed with weights $\hat w$ from periods $1,\ldots,t-1$ and the counterfactual $\hat\theta^0_{1,t}$ compared with $\hat\theta_{1,t}$; under Assumption (ref) it should not exceed $\hat c_{1-\alpha}$, and exceedance signals a pre-treatment violation. We apply this with $T_0=3$ in Section (ref).
The second form tests PTP for all $x$ at once. If the test does not reject (as is the case in our empirical application), there is direct evidence for a parallel trend structure over the whole parameter space. For strictly increasing $\Lambda$, $\delta_{x,t}(y)=0$ for every $x$ iff $\theta_{1,t}(y)=\theta^0_{1,t}(y)$, so one replaces $|\hat\delta_{x,t}(y_l)|$ in (ref) by the parameter-difference norm:
Critical values follow the same simulation scheme as for $T_n$, drawing the vector-valued process and taking the supremum of its norm.
A rejection of the full-distribution test $T_n(x,t,\mathcal{Y})$ that does not extend to $T_n(x,t,\mathcal{Y}_0)$ provides evidence that any PTP violation is driven by distributional shifts outside $\mathcal{Y}_0$, and is therefore consistent with Assumption (ref) on $\mathcal{Y}_0$ as required for $\hat f_t(x,\mathcal{Y}_0)$ to be a valid causal estimand.
This section establishes the asymptotic foundations for the estimators and the tests introduced in Sections (ref). All proofs are deferred to the Appendix. We maintain the following conditions throughout:
The technical foundation rests on three results (Lemmas (ref), (ref) and Theorem (ref)), stated and proved in Appendix (ref): weak convergence of the DR estimators in $l^\infty(\mathcal{Y}_0)^p$ (Lemma (ref)), asymptotic normality of the Gram-matrix estimator (Lemma (ref)), and asymptotic normality of the weight estimator (Theorem (ref)), from which consistency of the plug-in $\hat V_w$ follows by continuity. We now state the main result. Fix a post-treatment period $t\in\{T_0+1,\ldots,T_0+T_1\}$. Let $\Theta_t(y):=(\theta_{2,t}(y),\dots,\theta_{J+1,t}(y))^{\prime}\in\mathbb{R}^{J \times p}$ collect the period-$t$ DR parameters of all control groups.
If Assumption (ref) fails with residual $r_n(y):=\theta_{1,T_0}(y)-\sum_{i=2}^{J+1}w_i^*\theta_{i,T_0}(y)\neq 0$, the counterfactual estimator $\hat{f}_t(x)$ is generally not a consistent estimator of $f_t(x)$. Instead, it converges to a pseudo-value $\tilde f_t(x)\neq f_t(x)$. The following corollary gives a precise condition under which consistency for $f_t(x)$ is nevertheless restored.
To mimic the empirical application with a Mincer's earnings equation, we consider $J+1=5$ groups and two time periods. For $i=1,\ldots,5$ and $j=0,1$ we have the model:
with i.i.d.\ $X_{1,it},X_{2,it},X_{3,it} \sim N(0,1)$ and $\epsilon_{it}$ a noise distribution. This leads to $F_{it}(y|x) = \Phi\bigl(\theta_{it}(y)'x\bigr)$ with $\theta_{it}(y) = \bigl(y-\beta_{0,it},-\beta_{1,it},-\beta_{2,it},-\beta_{3,it}\bigr)$.
The $J=4$ control groups have parameter vectors $\boldsymbol{\beta}_i = \bar{\boldsymbol{\beta}} + c\,\mathbf{e}_i$ for $i=2,\ldots,5$, where $\bar{\boldsymbol{\beta}}=(1,1,1,1)'$, $c=0.8$, and $\mathbf{e}_i$ is the $i$-th standard basis vector in $\mathbb{R}^4$. Each control group thus differs from the others in exactly one coefficient. Under the null hypothesis, the treated group has $\boldsymbol{\beta}_1 = \frac{1}{J}\sum_{i=2}^{5}\boldsymbol{\beta}_i = (1.2,1.2,1.2,1.2)'$, so that PTP holds exactly with equal true weights $w^*_i=1/4$. This design ensures that the Gram matrix $G^*$ has full rank $J=4$ and condition number $\kappa(G^*)=55$, satisfying Assumption (ref).
Under the alternative hypothesis ($H_1$), the treated group's post-treatment outcome receives a covariate-heterogeneous effect $\Delta\cdot \beta_{1,11}$, i.e.\ a shift only along the $X_1$ direction, with $\Delta\in\{0,0.1,\ldots,0.5\}$. This shifts the conditional distribution at covariate values with $X_1\neq0$, while leaving the unconditional (marginal) distribution almost unchanged: since $E[X_1]=0$, the marginal mean and median are preserved and only the spread changes slightly. The design thus isolates a setting in which conditioning on covariates is essential.
The simulation focuses on the size and power properties of the supremum test (ref). We compare three procedures: the conditional supremum test (ref) with correctly specified errors ($\epsilon_{i,t}\sim N(0,1)$, “Probit”) and with misspecified errors ($\epsilon_{i,t}\sim Lo(0,1)$, “Logit”), and an unconditional distributional synthetic-control test that applies the same idea to the marginal outcome distribution (intercept-only DR, $p=1$), the natural in-framework analogue of an unconditional approach in the spirit of Gunsilius2023. The DR working model uses the probit link throughout. We use a grid of $m=10$ equidistant quantile levels of the outcome between $0.05$ and $0.95$, sample sizes $n_{it} \in \{200, 500, 1000\}$, and $N_{\rm mc}=1000$ Monte Carlo replications. Critical values are obtained by Gaussian process simulation with $S=10{,}000$ draws using the estimated covariance matrix $\hat K$.
The conditional tests are evaluated at $x=(1,1,0,0)'$ (so that $x_1=1$ and the conditional shift equals $\Delta$); the unconditional test targets the marginal CDF. Figure (ref) shows that all three procedures control size at $\Delta=0$, whereas the unconditional test is conservative. Under $H_1$, the conditional tests detect the effect with power increasing in $n$ and $\Delta$ (slightly higher under the correctly specified probit link than under the misspecified logit link), whereas the unconditional test has low power: The covariate-heterogeneous effect is invisible in the marginal distribution. Conditioning on covariates is therefore necessary to detect effects that are heterogeneous across the covariate space.
Figure (ref) reports properties of the one-sided 95% confidence interval (ref) for $f_t(x)$ from the conditional probit estimation, under $H_1$ ($\Delta>0$). Coverage is close to the nominal level across $\Delta$; the small deviations at the smallest effect sizes reflect the positive finite-sample bias of $\hat f$, which is a mean of squared quantities. The mean half-width falls with $n$ and grows with $\Delta$, as expected from the asymptotic theory. In practice the CI is reported only when the supremum test rejects $H_0$, which already selects cases where $f_0$ is not negligible relative to estimation error.
Under $T_0=2$ pre-treatment periods, the $T_0^{-1}$-scaling of $\Sigma_{G,c}$ in Lemma (ref) halves the asymptotic variance of the Gram-matrix estimator, and hence of $V_w$, of the weight estimator relative to $T_0=1$. We verified in an auxiliary simulation (same DGP, $T_0=2$, second pre-period drawn i.i.d.\ with the same parameters) that size remains well controlled at the nominal 5% level and power is modestly improved, consistent with the theoretical efficiency gain.
In April 1992, New Jersey raised its state minimum wage from \$4.25 to \$5.05 per hour (an increase of 19%) while the federal minimum remained unchanged at \$4.25. This policy change has been studied in the literature before. CardKrueger1994 find no negative employment effects using a DiD comparison of fast-food restaurants with neighbouring Pennsylvania, contrasting findings in NeumarkWascher1992. The DR-SC approach taken here offers a complementary perspective: rather than average employment effects, we trace the full conditional wage distribution across different worker types, allowing us to assess where in the distribution the policy had bite.
The distributional consequences of minimum wage increases have been widely studied: DiNardoFortinLemieux1996 document a link between minimum wage erosion and rising wage inequality using CPS data, and AutorManningSmith2016 provide evidence that the minimum wage compresses the lower tail of the wage distribution. ChernozhFVMelly2013 study counterfactual wage distributions using distribution regression. The contribution of the present approach is a clearer causal identification in terms of the workers' characteristics.
The validity of PTP in this application rests on the observation that the three pre-treatment years (1989--1992) span a well-defined common macroeconomic episode: the onset and trough of the 1990--91 recession, followed by an early recovery. This recession depressed low-education wages across all US states, but with differential intensity depending on each state's exposure to interest-rate-sensitive sectors (construction, durable manufacturing) and to the contemporaneous defense-spending contraction BlanchardKatz1992. These cross-state differences in cyclical exposure correspond to the state-specific factor loadings $\Lambda_i(y)$ in the latent-factor model of Remark 2.2: New Jersey's exposure profile (concentrated in finance, insurance, and real estate) is approximated by the weighted combination of donor states with Florida (finance and real estate) receiving the largest positive weight and South Dakota and Missouri (manufacturing-heavy, lower cyclical correlation with New Jersey) receiving negative weights, as reported in Table (ref). The convincing pre-trend tests in Section (ref) confirm that this synthetic New Jersey tracks the observed conditional wage distribution across all three pre-treatment years.
We use data from the CPS Basic Monthly files (flood2025ipums) organised into April--March policy years aligned with the reform date: $t=1$ (Apr 1989--Mar 1990), $t=2$ (Apr 1990--Mar 1991) and $t=3$ (Apr 1991--Mar 1992) are pre-treatment ($T_0=3$), and $t=4$ (Apr 1992--Mar 1993) is post-treatment ($T_1=1$). This periodisation aligns the start of the post-period exactly with the April 1992 increase, so that no partially-treated calendar year has to be discarded. We restrict to workers with a valid hourly wage, at most \$99/hour and non-missing education and experience. The outcome $Y$ is the log nominal hourly wage. The covariate vector $X$ includes a constant, years of education, years of potential experience, and its square, all standardised to zero mean and unit variance using the pooled sample. The standardised experience-squared variable is constructed as the square of standardised experience, $\widetilde{x_3}=\bigl((\mathrm{exper}-\bar x_2)/s_2\bigr)^2$. New Jersey policy-year sample sizes are 3{,}852, 3{,}998 and 4{,}031 (pre-treatment) and 3{,}905 (post-treatment).
The donor pool consists of $J=42$ US states, excluding eight states or districts with their own minimum wage changes during 1989--1993 (Connecticut, Massachusetts, Rhode Island, New Hampshire, Vermont, Alaska, Hawaii, and the District of Columbia).
As evaluation points, we consider the median worker (education = 12, experience = 10) as well as the four combinations of the $10\%$ and $90 \%$ values of education (10, 16) and experience (2, 37).
We estimate the DR parameters on a grid of $m=32$ quantiles of the pooled log-wage distribution (10th--90th percentiles). The total sample size is $n=381{,}953$ across all group-time cells. The Gram matrix $\hat G$ is the time-average over the three pre-treatment periods (see Section 2.2). It has full rank $J=42$, condition number $\kappa(\hat G)=9.6\times10^4$, and is numerically stable without regularisation. Table (ref) reports the five largest weights by magnitude.
Florida receives the largest positive weight, with New York second. The full weight distribution is shown in Figure (ref). 20 of 42 states receive negative weights, reflecting that New Jersey's pre-treatment wage structure is not fully within the convex hull of the donor pool, which the unconstrained formulation handles naturally.
Due to the post-period evidence from (ref), the evaluation point $x_{10}$ (educ=10, exper=2), reflecting low-education young workers, is of particular interest. Figure (ref) shows the pre-treatment fit for this point in all three pre-treatment years. The synthetic control tracks the observed New Jersey CDF closely across all pre-treatment years, with maximum sup-distance $\sup_y|\hat F_{\rm NJ,t}(y|x_{10})- \hat F^0_{\rm NJ,t}(y|x_{10})|\leq 0.037$ in each pre-period.
Figure (ref) reports the full-distribution and focused (MW-corridor) $p$-values of $T_n(x,t,\mathcal{Y})$ for the two pseudo-post pre-periods (Apr89/90$\to$Apr90/91 and $\to$Apr91/92) and the post-treatment period ($\to$Apr92/93), at all five evaluation points. The corridor for the focused test is fixed a priori to $\mathcal{Y}_0=[\log\$4.25,\log\$5.10]$ by the policy (raising the floor from \$4.25 to \$5.05).
Both pseudo-post tests do not reject $H_0$ at the five evaluation points, on the full distribution and on the corridor (all $p>0.35$; at $x_{10}$, $p=0.74$ and $p=0.54$ on $\mathcal{Y}$, and $p=0.81$ and $p=0.70$ on $\mathcal{Y}_0$). The 1990--91 recession is absorbed within the twelve-month windows rather than surfacing at a period transition, and the late-1991 announcement window (contained in the third pre-period) produces no detectable corridor pre-trend. Unlike a calendar-year periodisation, the design therefore requires no reconciliation of a full-distribution pre-trend rejection with the focused test. It seems that the parallel-trends assumption holds directly, both overall and on the corridor.
Beyond the five evaluation points, we also apply the $x$-simultaneous pre-trend test (ref), which checks parallel trends in the full DR parameter vector and hence for all covariate values at once. Neither pseudo-post transition rejects at the $10\%$ level: $\tilde T_n=144.9$ against $\hat c_{0.90}=245.3$ ($p=0.75$) for Apr89/90$\to$Apr90/91, and $\tilde T_n=209.6$ against $\hat c_{0.90}=241.7$ ($p=0.22$) for $\to$Apr91/92. So, it is reasonable to assume that parallel trends in parameters holds in the pre-periods not only at the points of interest but jointly across the covariate space.
Figure (ref) contrasts the full-distribution and focused tests at $x_{10}$. The focused statistic $T_n(x,t,\mathcal{Y}_0)$ serves two distinct roles. In the pre-periods it is the empirical counterpart of the orthogonality condition in Corollary (ref): should a full-distribution pre-trend reject while the corridor test passes, there is evidence that a potential pre-treatment misfit $r_n$ is orthogonal to the treatment effect $\Delta_t$ on $\mathcal{Y}_0$, and $\hat f_t(x,\mathcal{Y}_0)$ that remains consistent and asymptotically normal even where the Span Condition fails on the full support $\mathcal{Y}$. Here both pre-trend tests pass, so this safeguard is not needed and the corridor effect is identified without it. In the post-period the focused test instead provides power: It concentrates the supremum on the interval where the policy operates. At $x_{10}$ the full-distribution test is only marginal ($p=0.053$) while the focused test rejects clearly ($p=0.011$), with the supremum attained inside the corridor.
Figure (ref) displays the pointwise CDF differences $\hat\Delta_t(y\mid x):=\hat F_{1,t}(y\mid x)-\hat F^0_{1,t}(y\mid x)$ with 90% pointwise confidence bands for all five evaluation points. Table (ref) summarises the scalar treatment effect estimates $\hat f_t(x,\mathcal{Y})$ and the supremum test results. The effect is sharply concentrated in the MW corridor for the young low-educated group $x_{10}$, where $\hat\Delta_t$ falls by up to $0.12$ inside $[\$4.25,\$5.10]$. Both the focused test ($p_{\mathcal{Y}_0}=0.011$) and the full-distribution test ($p=0.053$) reject at the $10\%$ level. The high-education/high-experience point $x_{90}$ shows no effect on either test ($p=0.83$, $p_{\mathcal{Y}_0}=0.84$). We report the one-sided $90\%$ lower confidence bound $\mathrm{LCB}=\max\{0,\hat f_t-z_{0.90}\hat{\mathrm{se}}\}$ exactly when the supremum test rejects at $10\%$. This is only the case for $x_{10}$. Figure (ref) summarises the treatment effect estimates visually.
\paragraph{Low-education, low-experience workers ($x_{10}$: educ=10, exper=2).} The clearest effect is at $x_{10}$: $\hat f_t=1.71\times10^{-3}$, with the focused test rejecting on the corridor ($p_{\mathcal{Y}_0}=0.012$) and the full-distribution test rejecting at the $10\%$ level ($T_n=73.0$, $\hat c_{0.90}=64.8$, $p=0.054$); the one-sided $90\%$ lower bound is $0.27\times10^{-3}$ (full support) and $0.97\times10^{-3}$ (corridor). The shape is the canonical minimum-wage spike: $\hat\Delta_t(y)$ falls by up to $0.12$ inside the corridor $[\$4.25,\$5.10]$ and is near zero outside it. Probability mass previously below \$5.10 is shifted up to just above the new floor. The supremum is attained within the corridor, so the focused test is the sharper instrument here.
\paragraph{High-education, low-experience workers ($x_{hl}$: educ=16, exper=2).} The point estimate $\hat f_t=1.55\times10^{-3}$ is sizeable but not significant ($T_n=37.1$, $p=0.686$; corridor $p_{\mathcal{Y}_0}=0.150$), with a wide band reflecting the small high-education--low-experience cell. The pointwise confidence intervals of the CDF differences lie largely above zero, which would be consistent with the wage-compression/ripple pattern of AutorManningSmith2016.
\paragraph{Median worker ($x_0$: educ=12, exper=10).} The median worker shows a small effect that is marginal on the corridor only: $\hat f_t=0.58\times10^{-3}$ ($T_n=26.7$, $p=0.294$; $p_{\mathcal{Y}_0}=0.064$), consistent with a modest direct effect for the share of median workers whose wages were near the old floor.
\paragraph{Low-education, high-experience workers ($x_{lh}$: educ=10, exper=37).} The estimate is small, $\hat f_t=0.61\times10^{-3}$. The full-distribution test does not reject ($T_n=28.6$, $p=0.750$); the focused test is only marginal ($p_{\mathcal{Y}_0}=0.060$, rejecting at the $10\%$ level), and the one-sided lower bound is zero ($\mathrm{LCB}=0$, i.e.\ the interval $[0,\infty)$). There is no firmly resolved effect near the MW threshold for this group, consistent with their wages lying well above \$5.10.
\paragraph{High-education, high-experience workers ($x_{90}$: educ=16, exper=37).} For the high-education/high-experience point, we have no evidence for an effect: $\hat f_t=0.21\times10^{-3}$ with no rejection on either test ($T_n=27.6$, $\hat c_{0.90}=80.2$, $p=0.824$; corridor $p_{\mathcal{Y}_0}=0.846$). The absence of any corridor effect for a group whose wages lie far above the floor is as expected and supports the identifying assumptions.
\paragraph{Overall pattern.} The DR-SC approach reveals a clear heterogeneity pattern. Under the April--March design the minimum-wage effect is sharply and specifically concentrated in the MW corridor for low-education/low-experience young workers ($x_{10}$), the group most directly affected by the policy. For the high-education/high-experience group ($x_{90}$), there is no evidence for an effect, and the remaining groups show at most marginal corridor effects. Relative to a calendar-year periodisation, which we have also considered, less of the response is attributed to broad concurrent shifts, consistent with the cleaner pre-trends documented above.
This complements the aggregate DiD evidence of CardKrueger1994: While their design isolates average employment effects, the DR-SC approach reveals precisely which part of the conditional wage distribution was affected, and for which worker groups.
Regarding the interpretation as causal effect, under PTP, $\hat f_t(x,\mathcal{Y})$ estimates the total causal effect of all NJ-specific developments relative to the synthetic control between the pre- and post-treatment periods. For low-education/low-experience workers, the minimum wage increase is the most plausible dominant cause. For other groups, the effects may reflect a combination of the MW and concurrent NJ-specific developments. The DR-SC method provides a clean characterisation of where in the distribution the effects occur, which is informative about the operative mechanisms even when causal attribution is uncertain.
A variance decomposition of (ref) shows the pointwise bands are dominated by the weight-estimation term $V_w$ ($40$--$60\%$ of the pointwise variance across the five points, against only $\approx15\%$ for the treated cell), reflecting the ill-conditioned Gram matrix ($\kappa(\hat G)\approx10^5$). As a robustness check we re-estimate with the ridge weights of Remark (ref), selecting $\lambda$ by leave-one-period-out cross-validation on the pre-treatment fit.
The chosen $\lambda^*\approx3.8\times10^{-3}$ (of order $n^{-1/2}$) lowers $\kappa$ to $1.1\times10^4$, evens out the weights (the largest falls from $0.52$ to $0.24$), and improves the held-out pre-fit by about $20\%$, indicating that the unregularised weights mildly overfit. The substantive conclusions are unchanged: The effect stays concentrated in the corridor for $x_{10}$, no evidence for an effect for $x_{90}$, and all pre-trend tests still pass. Regularisation narrows the pointwise bands by roughly a third and sharpens the $x_{10}$ inference (its one-sided lower bound moves well above zero). Because $\lambda^*$ introduces a small, deliberate shrinkage of the weights towards equality (in the spirit of augmented synthetic control BenMichaelEtAl2021), we report it as a robustness analysis rather than the main specification.
We have proposed a SC estimator for conditional distribution functions in the semiparametric DR framework, with three main contributions. First, the Parallel Trends in Parameters assumption keeps the counterfactual within the model class, and dropping non-negativity on weights yields a closed-form estimator. Second, both DR estimation error and weight estimation error contribute at the same $\sqrt{n}$ rate to the variance of the counterfactual, and both components are characterised explicitly. Third, a two-stage inference procedure is proposed: a supremum test whose null distribution is approximated by Gaussian process simulation, followed by a plug-in confidence interval for the integrated difference when the null is rejected. The supremum statistic has a valid non-degenerate limit under $H_0$, grows to infinity under $H_1$, and yields $p$-values that fully exploit the large-$n$ precision of the DR estimators.
The New Jersey application illustrates how the proposed framework can uncover heterogeneous distributional patterns across the covariate space. The test statistic $T_n(x,t,\mathcal{Y}_0)$ with point-specific critical values from GP simulation reveals that the estimated effect is concentrated in the minimum-wage corridor for young low-educated workers ($x_{10}$: focused $p_{\mathcal{Y}_0}=0.012$ vs.\ full $p=0.054$), consistent with a direct minimum wage effect, while the other groups show at most marginal corridor effects and the high-education/experience group does not show any noticeable effect ($p_{\mathcal{Y}_0}=0.846$). Pre-trend tests pass on both the full distribution and the corridor. These patterns are invisible in aggregate SC analyses.
Future directions include the theory and applications of uniform confidence bands for $f(x)$ over a covariate set $\mathcal{X}$, enabling formal tests for x-heterogeneity of the treatment effect, and a nonparametric estimation of the link function in the DR model.