EconBase
← Back to paper

A Synthetic Control Approach to Conditional Distributional Treatment Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

72,305 characters · 17 sections · 48 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Synthetic Control Approach to Conditional Distributional Treatment Effects

\thispagestyle{empty}

abstract\begin{singlespace} This paper proposes a synthetic control (SC) framework for the estimation of conditional distributional treatment effects. Identification rests on a parallel trends condition formulated in the parameter space of the semiparametric distribution regression (DR) model, which keeps the counterfactual conditional distribution within the model class. The weights solve a least-squares problem subject to an adding-up constraint, yielding a closed-form estimator. We derive the asymptotic distribution of the counterfactual estimator, with DR estimation error and weight estimation error contributing at the same rate to the asymptotic variance. Moreover, we propose a supremum test for the null of no treatment effect, whose limit is the supremum of a Gaussian process. Simulations illustrate that conditioning on covariates can reveal effects being difficult to detect from the unconditional distribution alone. An application to the 1992 New Jersey minimum wage increase using CPS data finds effects concentrated in the minimum-wage corridor for low-education, low-experience workers. Keywords: Causal inference, functional delta method, distribution regression, minimum wage, parallel trends. JEL Codes: C12, C14, C21, C25, J31, J38. \end{singlespace}

\setcounter{page}{1}

Introduction

The synthetic control (SC) method, introduced by AbadieGardeazabal2003 and formalised in AbadieEtAl2010, has become one of the most widely used tools for the evaluation of policy interventions in settings with a small number of treated units; Abadie2021 provides a comprehensive review. The classical estimator constructs a weighted average of control unit outcomes that approximates the pre-treatment trajectory of the treated unit, and uses this synthetic unit as the counterfactual in the post-treatment period. DoudchenkoImbens2016 relax the non-negativity constraint on weights, allowing the synthetic control to lie outside the convex hull of the donor pool, and show that this flexibility improves pre-treatment fit. BenMichaelEtAl2021 propose an augmented SC estimator that combines outcome-model bias correction with SC weighting. ChernozhukovEtAl2021 develop an exact, finite-sample valid inference approach for SC based on conformal prediction. FermanPinto2021 study the properties of SC when pre-treatment fit is imperfect. ArkhangelskEtAl2021 develop a synthetic difference-in-differences estimator that accommodates heterogeneous treatment effects. Chen2023 establishes a connection between synthetic control and online learning, showing that the SC estimator can be interpreted as an instance of follow-the-leader, which yields oracle inequalities for counterfactual predictions even under adversarial outcomes. While powerful for estimating average treatment effects at the aggregate level, all these methods provide no information about distributional consequences, a critical limitation for policy questions involving inequality, minimum wages, or tax reforms.

The broader literature on distributional treatment effects includes the counterfactual wage decompositions of MachadoMata2005 and Melly2005 and ChernozhFVMelly2013, the unconditional quantile regressions of FirpoFortinLemieux2009, and the distributional DiD approaches of ChernozhFVMelly2013 and Fernandezval2026. None of these papers considers a synthetic control approach.

A first important extension was proposed by Gunsilius2023, who replaces scalar outcomes with quantile functions, thereby allowing inference on heterogeneous treatment effects across the distribution. His estimator matches unconditional quantile functions of the control units to that of the treatment unit in the pre-treatment period, using a constrained quantile-on-quantile regression. While this distributional synthetic control (DSC) method is a substantial advance, it is inherently unconditional: It cannot accommodate individual-level covariates such as education, experience or demographics that are central to labour economics and related fields.

The present paper fills this gap by proposing a SC estimator for conditional distribution functions. ChenFeng2026 extend the DSC framework to group-level heterogeneity, but their heterogeneity is driven by unobservable group membership rather than observable covariates, leaving the gap addressed here open. Conditioning on covariates is not simply a refinement of the unconditional approach: it changes the identification problem in a fundamental way. When the outcome distribution depends on individual characteristics, as wages depend on education and experience, the counterfactual object of interest is a conditional distribution function, and any identification assumption must be formulated at the level of the model that generates it.

We work within the semiparametric DR framework of ForesiPeracchi1995, where the conditional distribution function of an outcome $Y$ given covariates $X \in \mathbb{R}^p$ is modeled as $F(y\mid x)=\Lambda(x'\theta(y))$ for a known link $\Lambda$ and unknown parameter function $\theta:\mathcal{Y}\to\mathbb{R}^p$. This model has appealing properties in applications: it handles mass points, irregularities, and multimodality naturally, and estimation reduces to a sequence of binary outcome regressions indexed by $y$. The DR literature includes KneibEtAl2023 and Klein2024 on flexible extensions, spady:2025 on semiparametric efficiency, and Wied2024 on instrumental variable estimation in the DR framework; RotheWied2013 derive the asymptotic theory used here; and BiewenErhardt2025 apply distribution regression to minimum wage analysis. To the best of our knowledge, the combination of synthetic control methods with semiparametric distribution regression has not been studied in the existing literature.

Our identification assumption Parallel Trends in Parameters (PTP) is formulated directly in the DR parameter space: the counterfactual parameter function of the treated group evolves from the pre-treatment period to the post-treatment period along the same weighted combination of trends observed in the control groups, requiring the weights to sum to one. The economic rationale is that the DR parameters (e.g. quantile-specific returns to education and experience in a Mincer framework) are driven by a small number of common aggregate shocks (e.g. skill-biased technological change, business-cycle fluctuations, globalisation) whose differential local impact is captured by state-specific factor loadings. PTP formalises the condition that the factor loadings of a treated unit can be approximated by a weighted combination of donor states' loadings, which is the direct functional analogue of the latent-factor rationale underlying classical SC methods AbadieEtAl2010. This approximation is natural in the DR parameter space rather than at the CDF level, because the parameter function is linear in the common factors whereas the CDF is not. This formulation has also two other important advantages over a parallel trends condition on the CDF level. First, it ensures that the counterfactual conditional distribution function remains within the DR model class, so that all estimators have a direct parametric interpretation. Second, it allows us to formulate the weight estimation problem as a finite-dimensional quadratic program (QP) at each value of $y$, or equivalently as a single infinite-dimensional QP over the function space $\ell^{\infty}(\mathcal{Y})^p$. Dropping the traditional non-negativity constraint yields a closed-form weight estimator, avoids interior-solution requirements, and transparently accounts for potentially negative weights that arise when some control units lie outside the convex hull of the treated unit.

One contribution of the paper is the asymptotic analysis of the proposed estimator for the case of fixed number of groups $J$ and increasing individual observations within each group. This framework fits to micro-data applications such as the Current Population Survey (CPS) for US states, which we consider, as the number of states is small relative to the within-group sample size $n$. Statistical precision derives from this sample size, not from the number of time periods $T$. In fact, one pre-period is sufficient for obtaining consistent parameter estimators, whereas more pre-periods increase the precision of the weight estimators. An important result is the joint weak convergence of the pointwise counterfactual CDF differences to a Gaussian vector, where both DR estimation error and weight estimation error contribute at the same $\sqrt{n}$ rate. Based on this, we show asymptotic normality of the integrated squared treatment effect estimator when the true effect is non-zero. For testing the null hypothesis of no treatment effect, we propose a supremum statistic over the pointwise CDF differences, applicable either on the whole support of $Y$ or a subset. Its null distribution is the supremum of a mean-zero Gaussian process. This supremum statistic has a non-degenerate null distribution, which is approximated by a Gaussian process simulation, and power growing to one with $n$. A plug-in confidence interval for the integrated squared effect is reported in a second step, conditional on rejection. We characterize under which types of deviation from PTP the results still hold.

If multiple pre-periods are available, it is possible to perform pre-trend diagnostics. We use them to perform causal inference in an empirical application to the 1992 New Jersey minimum wage increase. In fact, we see that the DR-SC approach reveals a clear heterogeneity pattern that is invisible in aggregate analyses. The minimum wage effect is sharply concentrated in the corresponding corridor for low-education and low-experience workers, the group most directly affected by the policy. For other groups the effects are at most marginal, in particular for high-education/experience workers.

The paper proceeds as follows. Section 2 introduces the framework for general $T_0\geq 1$ pre-treatment and $T_1\geq 1$ post-treatment periods and discusses point as well as confidence estimation of the treatment effect. Moreover, the supremum test and the pre-trend test for the case $T_0\geq 2$ are presented. Section 3 derives the supporting asymptotic theory. Section 4 presents simulation evidence, which highlights the advantages of the conditional over an unconditional approach. Section 5 contains the empirical application, Section 6 concludes.

Framework and Estimation

Framework

Consider $J+1$ groups $i=1,\ldots,J+1$ at $T=T_0+T_1$ time points, where $t=1,\ldots,T_0$ are pre-treatment periods and $t=T_0+1,\ldots,T_0+T_1$ are post-treatment periods. Group $i=1$ is the treatment group; groups $i=2,\ldots,J+1$ form the donor pool of potential control groups. In all pre-treatment periods none of the groups is treated. Only group $i=1$ receives the treatment from period $T_0+1$ onward. Within each group-time cell $(i,t)$ we observe an i.i.d.\ sample of $n_{it}$ individuals with outcome $Y\in\mathcal{Y}\subseteq\mathbb{R}$ and a $p$-dimensional covariate vector $X\in\mathcal{X}\subseteq\mathbb{R}^p$. Typical applications in labour economics have $J$ small (e.g., $J+1=10$ to 50 federal states or countries) and $n_{it}$ large (e.g., several thousand individuals per group-time cell). We fix $J$ and let $n:=\sum_{i,t}n_{it}\to\infty$, assuming $n_{it}/n\to r_{it}\in(0,\infty)$.

For each group $i$ and time period $t$, the conditional distribution function of $Y$ given $X=x$ satisfies

equation[equation omitted — 131 chars of source]

where $\Lambda:\mathbb{R}\to(0,1)$ is a known strictly increasing link function (such as the standard normal CDF for the probit link, or the logistic CDF for the logit link) with derivative $\lambda$ and $\theta_{it}:\mathcal{Y}\to\mathbb{R}^p$ is an unknown c\`adl\`ag parameter function. For fixed $y$, $\theta_{it}(y)$ is the coefficient vector in the binary outcome regression of $\mathbf{1}\{Y\leq y\}$ on $X$. Numerical sensitivity analyses in dette2025 confirm that the choice of the link function has a negligible impact on estimation results.

remark[Monotonicity] The model ensures that $y\mapsto F_{it}(y\mid x)$ is a valid distribution function (monotone in $y$, taking values in $(0, 1)$) if $y\mapsto x'\theta_{it}(y)$ is non-decreasing for all $x$ in the support. This monotonicity can be imposed ex post via rearrangement Chernozhukov2010 or isotonic regression Wied2024.

We are interested in the treatment effect on a region of the outcome distribution $\mathcal{Y}_0\subseteq\mathcal{Y}$, measured via the integrated squared discrepancy between the observed and counterfactual conditional distributions:

equation[equation omitted — 213 chars of source]

where $F^{0}_{1,t}(y\mid x)$ is the counterfactual conditional distribution of the treated group in post-treatment period $t$, and

equation[equation omitted — 97 chars of source]

is the pointwise CDF difference. The integrated discrepancy $f_t(x,\mathcal{Y}_0)$ captures both the magnitude and the shape of the treatment effect on the conditional distribution at covariate value $x$ over the region $\mathcal{Y}_0$. It equals zero if and only if $\Delta_t(y\mid x)=0$ for almost all $y\in\mathcal{Y}_0$. Setting $\mathcal{Y}_0=\mathcal{Y}$ recovers the full integrated effect; a proper subset $\mathcal{Y}_0\subsetneq\mathcal{Y}$ focuses attention on a specific region of the distribution, such as the minimum-wage corridor $[\mathrm{MW}_{\mathrm{old}},\mathrm{MW}_{\mathrm{new}}+\varepsilon]$. One may also consider the doubly integrated measure $\int_{\mathcal{X}}f_t(x,\mathcal{Y}_0)dF_{X}(x)$, which averages over the covariate distribution. If not denoted otherwise, $f_t(x,\mathcal{Y}_0) = f_t(x)$ for notational simplicity in the following.

The fundamental identification challenge is that $F_{11}^{0}(y|x)$ is not observed. We identify it via the following assumption, which has to hold on the set of interest $\mathcal{Y}_0$.

assumption[Parallel Trends in Parameters, PTP] There exist weights $w_2,\ldots,w_{J+1}\in\mathbb{R}$ with $\sum_{i=2}^{J+1}w_i=1$ such that for all $y\in\mathcal{Y}_0$ and all post-treatment periods $t=T_0+1,\ldots,T_0+T_1$, \[ \theta^{0}_{1,t}(y)-\theta_{1,T_0}(y) =\sum_{i=2}^{J+1}w_i\bigl[\theta_{i,t}(y)-\theta_{i,T_0}(y)\bigr], \] where $\theta_{1,t}^{0}(y)$ is the DR parameter function of the counterfactual $F_{1,t}^{0}(y\mid \cdot)=\Lambda(\cdot'\theta_{1,t}^{0}(y))$ for each post-treatment period $t$.

Assumption PTP states that the trend in the treated group's DR parameter function from the pre- to the post-treatment period would have equalled a weighted combination of the control groups' trends, absent treatment. This is a direct analogue of the classical parallel trends assumption in the difference-in-differences (DiD) literature, formulated in the parameter space of the DR model rather than on the level of the outcome or its conditional mean. We impose only the adding-up constraint; negative weights arise when some control units lie outside the convex hull of the treated unit. Differently to the case of negative weights in regression-based DiD CallawayLi2019, this is not a problem of model specification, but a particular feature, which yields flexibility. DoudchenkoImbens2016 make a similar argument in the SC context and show that allowing negative weights improves pre-treatment fit.\\

remark[Economic motivation of PTP] Assumption PTP can be motivated by a latent factor model of the distribution regression parameters. Suppose that, absent treatment, the conditional wage distribution in each state is shaped by a small number of common macroeconomic factors such as inflation, technological change, or business-cycle conditions. These forces affect all states, but with different intensities depending on each state’s industry composition, labour market structure, and demographic profile. Formally, assume that the untreated DR parameter function admits the representation \[ \theta^{0}_{it}(y) = \mu_{i}(y) + \Lambda_{i}(y)' F_{t} + u_{it}(y), \] where $F_{t} \in \mathbb{R}^{r}$ is a vector of common factors, $\Lambda_{i}(y)$ is a vector of state-specific factor loadings, $\mu_{i}(y)$ captures time-invariant heterogeneity, and $u_{it}(y)$ is an idiosyncratic disturbance. If the factor loadings of the treated unit can be approximated by a weighted average of the loadings of the donor states, $\Lambda_{1}(y) \approx \sum_{i=2}^{J+1} w_{i}\Lambda_{i}(y)$ with $\sum_{i=2}^{J+1} w_{i} = 1$, then the counterfactual evolution of the treated unit satisfies approximately \[ \theta^{0}_{1,t}(y) - \theta_{1,T_0}(y) \approx \sum_{i=2}^{J+1} w_{i}\bigl(\theta_{i,t}(y) - \theta_{i,T_0}(y)\bigr). \] Thus PTP is the direct functional analogue of the latent-factor rationale underlying classical synthetic control methods, see Abadie2021. This latent-factor structure has empirical support in the labour economics literature. In a Mincer framework, the DR parameters $\theta_{it}(y)$ collect quantile-specific returns to education and experience. The time variation in these returns (documented across US states by KatzMurphy1992 and CardLemieux2001) is well explained by a small number of common demand shocks: Skill-biased technological change raises returns to education uniformly across states but with different intensities depending on the local industry mix, aggregate business-cycle fluctuations compress or expand the lower tail of the wage distribution through their differential impact on low-education employment AutorManningSmith2016 and trade exposure affects manufacturing-intensive states more strongly AutorDornHanson2013. Each of these forces maps naturally to a common factor $F_t$, with state-specific industry composition determining the loading $\Lambda_i(y)$. An advantage of formulating PTP in the DR parameter space (rather than directly at the CDF level) is that the factor structure is linear in $\theta_{it}(y)$, whereas the corresponding CDF $F_{it}(y|x)=\Lambda(x'\theta_{it}(y))$ is a nonlinear. This linearity delivers constant weights $w$ across the full distribution, which is economically natural (the same macroeconomic forces act on all quantiles of the conditional distribution).
remark[Pre-treatment balance] A natural sufficient condition for PTP is that (i) the pre-treatment parameter functions satisfy $\theta_{1,t}(y)=\sum_{i=2}^{J+1}w_i\theta_{i,t}(y)$ for all $y$ and all $t=1,\ldots,T_0$ (perfect pre-treatment balance), and (ii) the control groups would have followed the same weighted trend as the treatment group in the absence of treatment. Under perfect pre-treatment balance, Assumption PTP simplifies to $\theta^{0}_{1,t}(y)=\sum_{i=2}^{J+1} w_i\theta_{i,t}(y)$ for each post-period $t$, which means the counterfactual DR parameter is simply a weighted average of the period-$t$ DR parameters of the control groups. We will use this simplification as the basis for estimation.
remark[Comparison to CDF-level parallel trends] One might alternatively formulate parallel trends directly on the CDF level: $F_{1,t}^{0}(y\mid x)=\sum_{i=2}^{J+1}w_i F_{i,t}(y\mid x)$ for all $y,x$, and each post-period $t$. Under Assumption PTP and the DR model, this holds approximately. A second-order Taylor expansion yields the discrepancy: \[ \Lambda\Bigl(x'\sum_{i=2}^{J+1}w_i\theta_{i,t}(y)\Bigr)-\sum_{i=2}^{J+1}w_i\Lambda\bigl(x'\theta_{i,t}(y)\bigr) \approx-\frac{1}{2}\lambda'\bigl(x'\overline{\theta}_t(y)\bigr)\cdot \mathrm{Var}_{w}\bigl(x'\theta_{i,t}(y)\bigr), \] where $\overline{\theta}_t(y)=\sum_i w_i\theta_{i,t}(y)$ and $\lambda=\Lambda'$. The approximation error is thus proportional to the variance of the DR index $x'\theta_{i,t}(y)$ across control groups, small when control groups are similar and large when they are heterogeneous.

Point Estimation

For each cell $(i,t)$ and $y$ on a grid $\mathcal{Y}_m=\{y_1,\ldots,y_m\}\subset\mathcal{Y}_0$, the DR parameter $\theta_{it}(y)$ is estimated by maximizing the log-likelihood of the binary outcome model. For a fixed $y\in\mathcal{Y}$, the estimator solves:

equation[equation omitted — 268 chars of source]

This is a standard probit or logit regression of $\mathbf{1}\{Y\le y\}$ on $X$, estimated separately for each $y$ on a finite grid $\mathcal{Y}_{m}=\{y_{1},...,y_{m}\}\subset\mathcal{Y}_0$ (e.g. the 0.05,0.1,\ldots quantiles of the distribution of $Y$).

The weights $w=(w_{2},...,w_{J+1})^{\prime}$ are estimated by minimizing the pre-treatment discrepancy averaged over all $T_0$ pre-treatment periods:

equation[equation omitted — 228 chars of source]

By averaging over pre-treatment periods and integrating the quadratic discrepancy over the support of the outcome $y$, we move from a point-wise to a distribution-wide weight estimation. This ensures that the weights $w$ are constant across the entire conditional distribution and across time. This global objective function acts as a natural stabilizer, effectively reducing the dimensionality of the optimisation problem and mitigating the risk of over-fitting that typically arises when weighting high-dimensional objects.

In practice, the integral over $\mathcal{Y}_0$ is replaced by a sum over the grid $\mathcal{Y}_{m}$

equation[equation omitted — 247 chars of source]

This is a quadratic program in $w$ with linear constraints and a positive semidefinite objective. Let $\hat G\in\mathbb{R}^{J\times J}$ denote the time-averaged Gram matrix with entries \[ \hat G_{kl}=\frac{1}{T_0\cdot m}\sum_{t=1}^{T_0}\sum_{l'=1}^m \hat\theta_{kt}(y_{l'})'\hat\theta_{lt}(y_{l'}) \] and $\hat c\in\mathbb{R}^J$ the vector with entries \[ \hat c_k=\frac{1}{T_0\cdot m}\sum_{t=1}^{T_0}\sum_{l'=1}^m\hat\theta_{1t}(y_{l'})'\hat\theta_{kt}(y_{l'}). \] Solving the Lagrangian of (ref) yields the closed-form solution (ref), unchanged in form.

equation[equation omitted — 169 chars of source]
remark[Ridge regularisation] When the condition number $\kappa(\hat G)$ is large, the closed-form solution (ref) may be numerically unstable. A practical remedy is to replace $\hat G$ by the ridge-regularized matrix $\hat G_\lambda = \hat G + \lambda I_J$ for some $\lambda>0$, yielding \begin{equation} \hat w_\lambda = \hat G_\lambda^{-1}\hat c -\hat G_\lambda^{-1}\mathbf{1} \,\frac{\mathbf{1}'\hat G_\lambda^{-1}\hat c-1}{\mathbf{1}'\hat G_\lambda^{-1}\mathbf{1}}. \end{equation} The regularisation introduces a bias of order $O(\lambda\|w^*\|)$ in the weights, which shrinks $\hat w_\lambda$ toward $J^{-1}\mathbf{1}$. For $\lambda\to 0$ the solution (ref) converges to (ref) whenever $\hat G$ is invertible. In practice, $\lambda$ should be chosen small enough that $\lambda/\sigma_{\min}(\hat G)\ll 1$, so that the numerical benefit outweighs the shrinkage bias. The theoretical results of Section (ref) continue to hold for $\hat w_\lambda$ whenever $\lambda = o(n^{-1/2})$.

The counterfactual DR parameter for post-treatment period $t\in\{T_0+1,\ldots,T_0+T_1\}$ is estimated as:

equation[equation omitted — 116 chars of source]

which exploits the pre-treatment balance enforced by the weight estimator (ref) (see Remark (ref)). The counterfactual conditional distribution function is then $\hat{F}_{1,t}^{0}(y|x)=\Lambda(x^{\prime}\hat{\theta}_{1,t}^{0}(y))$ and the treatment effect estimator for post-period $t$ is:

equation*[equation* omitted — 161 chars of source]

approximated in practice by a Riemann sum over $\mathcal{Y}_{m}$.

remark[Direct PTP estimator] An alternative estimator uses Assumption 1 directly, \[ \tilde\theta^{0}_{1,t}(y) \;=\; \hat\theta_{1,T_0}(y) +\sum_{i=2}^{J+1}\hat w_{i} \bigl(\hat\theta_{i,t}(y)-\hat\theta_{i,T_0}(y)\bigr), \] and is consistent under PTP without requiring perfect pre-treatment balance. However, it incurs additional variance relative to $\hat{\theta}^{0}_{1,t}$ in (ref), because pre-period estimation errors in $\hat\theta_{1,T_0}$ and $\hat\theta_{i,T_0}$ propagate into the counterfactual. Under perfect balance the two estimators are asymptotically equivalent, so $\hat{\theta}^{0}_{1,t}$ is preferred on efficiency grounds. Corollary (ref) identifies the weaker condition under which $\hat f_t(x)$ remains a consistent estimator of $f_t(x)$ even when pre-treatment balance is imperfect.

Confidence interval for $f_t(x)$

Since $f_t(x)\ge0$ by construction, we report a one-sided confidence interval. When the treatment effect is non-zero, $f_t(x)>0$, the asymptotic normality of $\hat f_t(x)$ (established in Theorem (ref)(c) in Section (ref)) yields, at significance level $\alpha$, the one-sided $(1-\alpha)$ confidence interval

equation[equation omitted — 153 chars of source]

where $u_{1-\alpha}$ is the $(1-\alpha)$-quantile of the $N(0,1)$-distribution and

eqnarray*[eqnarray* omitted — 294 chars of source]

and $\hat K_t$ is the covariance kernel for post-period $t$:

eqnarray[eqnarray omitted — 863 chars of source]

The generalized residuals (scores) are $$\psi_{i,t,k}(y) = \dfrac{\lambda(X_{i,t,k}'\hat\theta_{i,t}(y))}{\Lambda(X_{i,t,k}'\hat\theta_{i,t}(y))\,(1-\Lambda(X_{i,t,k}'\hat\theta_{i,t}(y)))}\bigl(\mathbf{1}\{Y_{i,t,k} \le y\} - \Lambda(X_{i,t,k}'\hat{\theta}_{i,t}(y))\bigr),$$ the standardized score residual of the binary log-likelihood (ref). The vector $\hat{\tilde\Theta}_{x,t}(y_l)=\bigl(x'\hat\theta_{2,t}(y_l),\ldots,x'\hat\theta_{J+1,t}(y_l)\bigr)'\in\mathbb{R}^J$ collects the scalar projections of the period-$t$ donor parameters onto $x$, and $\hat{V}_w$ is an estimator for $V_w$ from Theorem (ref). Moreover, $$\widehat{\mathcal{I}}_{i,t}(y)=\frac{1}{n_{i,t}}\sum_k \dfrac{\lambda(X_{i,t,k}'\hat\theta_{i,t}(y))^2}{\Lambda(X_{i,t,k}'\hat\theta_{i,t}(y))\,(1-\Lambda(X_{i,t,k}'\hat\theta_{i,t}(y)))} X_{i,t,k}X_{i,t,k}'$$ is the empirical Fisher information at $y$ (the sample analogue of $\mathcal{I}_{i,t}$ in Lemma (ref)). The first term in (ref) captures DR estimation uncertainty for the treated group in period $t$; the second captures uncertainty from the control groups; the third captures weight estimation uncertainty, corresponding to the $\boldsymbol{\Theta}_t(y)'V_w\boldsymbol{\Theta}_t(y')$ component of (ref). We estimate $V_w$ by the following plug-in estimator. With $\hat P=\partial w^*/\partial c$ at $\hat G$ (see (ref)) and $\hat\Theta^{(s)}(y)=(\hat\theta_{2,s}(y),\ldots,\hat\theta_{J+1,s}(y))'\in\mathbb{R}^{J\times p}$, define the cell loadings $\hat B_{i,s}(y)\in\mathbb{R}^{J\times p}$ (under Assumption (ref))

equation[equation omitted — 159 chars of source]

and the plug-in estimator is the score-outer-product sum

equation[equation omitted — 303 chars of source]

with the score residual $\psi$ and Fisher information $\widehat{\mathcal{I}}$ of (ref); consistency follows from Theorem (ref).

The interval (ref) achieves nominal one-sided $(1-\alpha)$ coverage asymptotically when $f_t(x)>0$. When $f_t(x)=0$, $\hat\sigma^2_t(x)$ degenerates to zero and the lower bound collapses to $0$; testing $H_0:f_t(x)=0$ requires a separate procedure, described in Section (ref).

Inference: Supremum Test for $H_0: f(x,\mathcal{Y}_0)=0$

The null hypothesis $H_0:f_t(x,\mathcal{Y}_0)=0$ is equivalent to $H_0:\Delta_t(y|x)=0$ for (Lebesgue-almost) all $y\in\mathcal{Y}_0$. We test this hypothesis via the supremum of the pointwise CDF differences over $\mathcal{Y}_0$.

The test statistic for post-treatment period $t$ and region $\mathcal{Y}_0$ is

equation[equation omitted — 129 chars of source]

where $\hat\delta_{x,t}(y_l):=\hat F_{1,t}(y_l|x)-\hat F^0_{1,t}(y_l|x)$. The full-distribution test uses $\mathcal{Y}_0=\mathcal{Y}_m$; a focused test restricts to a proper subset such as the minimum-wage corridor. The asymptotic theory applies to both with $m$ replaced by $|\{l:y_l\in\mathcal{Y}_0\}|$.

Critical values are obtained by simulation of the Gaussian process $Z_x$ from Remark (ref):

enumerate• On a previously chosen grid $y_l$ estimate the covariance kernel matrix by the sandwich estimator $\hat{K}(y_l, y_{l'})$ from (ref). • Compute the Cholesky factorisation $\hat K = LL'$ of the covariance matrix. • Draw $S$, e.g. $S=10{,}000$, independent realisations $B^{(s)} = L\varepsilon^{(s)}$, $\varepsilon^{(s)}\sim\mathcal{N}(0,I_m)$, and set $T^{(s)} = \sup_l|B^{(s)}_l|$. • Reject $H_0$ at level $\alpha$ if $T_n(x)>\hat c_{1-\alpha}$, the empirical $(1-\alpha)$-quantile of $\{T^{(s)}\}_{s=1}^S$.

The $p$-value is $\hat p = S^{-1}\sum_s\mathbf{1}\{T^{(s)}\ge T_n(x)\}$. If Stage 1 rejects $H_0$, the confidence interval (ref) provides quantitative precision for $f(x)$.

Pre-Trend Test

When $T_0\geq 2$, evidence for the validity of PTP can be given by pre-testing. As usual in difference-in-difference approaches (Lechner2011), the idea is to extrapolate from the time before the treatment (where all groups are non-treated) to the time after, where the counterfactual outcome of the treated group is not observed.

The tests can be formed in two ways: for a specific $x$ with the statistic (ref), or for all $x$ simultaneously via the modification below. In the focused form, each pre-period $t\in\{1,\ldots,T_0-1\}$ is treated as a pseudo-post period and $T_n(x,t,\mathcal{Y})$ from (ref) is computed with weights $\hat w$ from periods $1,\ldots,t-1$ and the counterfactual $\hat\theta^0_{1,t}$ compared with $\hat\theta_{1,t}$; under Assumption (ref) it should not exceed $\hat c_{1-\alpha}$, and exceedance signals a pre-treatment violation. We apply this with $T_0=3$ in Section (ref).

The second form tests PTP for all $x$ at once. If the test does not reject (as is the case in our empirical application), there is direct evidence for a parallel trend structure over the whole parameter space. For strictly increasing $\Lambda$, $\delta_{x,t}(y)=0$ for every $x$ iff $\theta_{1,t}(y)=\theta^0_{1,t}(y)$, so one replaces $|\hat\delta_{x,t}(y_l)|$ in (ref) by the parameter-difference norm:

equation[equation omitted — 159 chars of source]

Critical values follow the same simulation scheme as for $T_n$, drawing the vector-valued process and taking the supremum of its norm.

A rejection of the full-distribution test $T_n(x,t,\mathcal{Y})$ that does not extend to $T_n(x,t,\mathcal{Y}_0)$ provides evidence that any PTP violation is driven by distributional shifts outside $\mathcal{Y}_0$, and is therefore consistent with Assumption (ref) on $\mathcal{Y}_0$ as required for $\hat f_t(x,\mathcal{Y}_0)$ to be a valid causal estimand.

Asymptotic Theory

This section establishes the asymptotic foundations for the estimators and the tests introduced in Sections (ref). All proofs are deferred to the Appendix. We maintain the following conditions throughout:

assumption[Sampling] For each $(i,t)$, $\{(Y_{itk},X_{itk})\}_{k=1}^{n_{it}}$ are i.i.d.; different cells are independent; $n_{it}/n\to r_{it}\in(0,\infty)$.
assumption[DR Regularity] For each $(i,t)$ and $y\in\mathcal{Y}_0$, $\theta_{it}(y)$ uniquely maximizes the expected log-likelihood and the score satisfies conditions for the functional CLT in $\ell^{\infty}(\mathcal{Y}_0)^p$, see ChernozhFVMelly2013, Condition DR.
assumption[Gram Matrix] The population Gram matrix $G^*\in\mathbb{R}^{J\times J}$, defined as the population analogue of $\hat G$ in (ref), \[ G^*_{kl} = \frac{1}{T_0}\sum_{t=1}^{T_0}\int_{\mathcal{Y}_0}\theta_{kt}(y)'\theta_{lt}(y)\,dy, \] is positive definite.
assumption[Pre-Treatment Balance] The pre-treatment DR parameter function of the treated group lies in the closed linear span of the control groups' pre-treatment DR parameter functions in $L^{\infty}(\mathcal{Y}_0)^{p}$, i.e., there exist weights $w^{*}=(w^{*}_{2},\ldots,w^{*}_{J+1})'$ with $\mathbf{1}'w^{*}=1$ such that \[ \theta_{1,t}(y) = \sum_{i=2}^{J+1}w^{*}_{i}\,\theta_{i,t}(y) \qquad \text{for almost all } y\in\mathcal{Y}_0 \text{ and all } t=1,\ldots,T_0. \]
remarkAssumption (ref) is the functional analogue of the classical SC span condition: the treated unit's pre-treatment path lies in the convex (here: affine) hull of the donor pool in the DR parameter space. It implies perfect pre-treatment balance in all pre-periods, so that the weight estimator (ref) is consistent for $w^*$ and $\hat{\theta}^{0}_{1,t}$ is a consistent estimator of $\theta^{0}_{1,t}$ for each post-period $t$. Together with Assumption (ref), it also implies that $w^*$ is the unique minimum-norm solution to the pre-treatment balance equations. This assumption can be relaxed, in particular under an orthogonality condition (Corollary (ref)) and using a different estimator (Remark (ref)). Moreover, it can be tested, see Section (ref).
remarkAssumption (ref) ensures that the control groups are sufficiently distinct to identify the weights uniquely.

The technical foundation rests on three results (Lemmas (ref), (ref) and Theorem (ref)), stated and proved in Appendix (ref): weak convergence of the DR estimators in $l^\infty(\mathcal{Y}_0)^p$ (Lemma (ref)), asymptotic normality of the Gram-matrix estimator (Lemma (ref)), and asymptotic normality of the weight estimator (Theorem (ref)), from which consistency of the plug-in $\hat V_w$ follows by continuity. We now state the main result. Fix a post-treatment period $t\in\{T_0+1,\ldots,T_0+T_1\}$. Let $\Theta_t(y):=(\theta_{2,t}(y),\dots,\theta_{J+1,t}(y))^{\prime}\in\mathbb{R}^{J \times p}$ collect the period-$t$ DR parameters of all control groups.

theorem[Asymptotic distribution of counterfactual and treatment effect] Under Assumptions (ref)--(ref), for a fixed evaluation point $x\in\mathcal{X}$ and post-treatment period $t$: (a) $\sqrt{n}(\hat{\theta}^{0}_{1,t}(\cdot)-\theta^{0}_{1,t}(\cdot))\rightsquigarrow\mathbb{G}^0_t(\cdot)$ in $\ell^{\infty}(\mathcal{Y}_0)^p$, where \begin{equation} \operatorname{Cov}(\mathbb{G}^0_t(y),\mathbb{G}^0_t(y')) =\underbrace{\sum_{i=2}^{J+1}\frac{(w^*_i)^2}{r_{i,t}} \operatorname{Cov}(\mathbb{G}_{i,t}(y),\mathbb{G}_{i,t}(y'))}_{DR estimation error (post-period $t$)} +\underbrace{\boldsymbol{\Theta}_t(y)'V_w\,\boldsymbol{\Theta}_t(y').}_{\substack{weight estimation error\\[2pt] \scriptstyle V_w\text{ aggregates the pre-period }r_{i,s}^{-1}\Phi_{i,s}\text{ of~\eqref{eq:SigmaGc}}}} \end{equation} \textup{(b)} Define the common-scale process $\mathbb{H}_{it}:=r_{it}^{-1/2}\mathbb{G}_{it}$. For $Z_{n,x,t}(y) := \sqrt{n}(\hat\delta_{x,t}(y)-\delta_{x,t}(y))$, it holds $Z_{n,x,t}(\cdot) \rightsquigarrow Z_{x,t}(\cdot) = \lambda(x'\theta_{1,t}(\cdot)) \cdot x' \mathbb{H}_{1,t}(\cdot) - \lambda(x'\theta^{0}_{1,t}(\cdot)) \cdot x' \mathbb{G}^0_t(\cdot)$. \textup{(c)} If $f_t(x)>0$: $\sqrt{n}(\hat{f}_t(x)-f_t(x))\xrightarrow{d}\mathcal{N}(0,\sigma^2_t(x))$, where $\sigma^2_t(x)$ is the variance of $\dot\phi_{(\theta_{1,t}(y),\theta^{0}_{1,t}(y))}(\mathbb{H}_{1,t},\mathbb{G}^0_t)$ and \[ \dot\phi(h,h^0)=2\int_{\mathcal{Y}_0} [\Lambda(x'\theta_{1,t}(y))-\Lambda(x'\theta^{0}_{1,t}(y))] [\lambda(x'\theta_{1,t}(y))x'h-\lambda(x'\theta^{0}_{1,t}(y))x'h^0]\,dy. \]
remarkIn Theorem (ref)(c), the condition $f_t(x)>0$ is essential: under $H_0:\delta(y)=0$, the derivative $\dot\phi$ vanishes and $\sigma^2(x)=0$, so this result does not provide a valid test for $H_0$. It underpins the confidence interval (ref) in Section (ref), which is valid conditional on $f_t(x)>0$ having been established by the supremum test of Section (ref). In fact, Theorem (ref)(b) is the theoretical basis for this test. Moreover, the covariance kernel $K$ of $Z(\cdot)$ is estimated with the sandwich formula for the confidence interval (ref).
remarkThe two-component covariance in (ref) corrects a common implicit error: weight estimation uncertainty and DR estimation uncertainty are both $O_p(n^{-1/2})$ and are both included in the variance of the counterfactual. The two terms are uncorrelated because $\hat w$ depends only on pre-period data while the first term involves post-period data (Assumption (ref)). The terms can be consistently estimated (as in (ref)) by replacing all population quantities with their sample counterparts and estimating the covariance kernels of $\mathbb{G}_{it}$ via the sandwich formula for binary outcome regressions.
remark[Supremum test] Theorem (ref)(b) immediately yields the asymptotic null distribution of (ref): under $H_0:\Delta_t(\cdot|x)=0$ on $\mathcal{Y}_0$, the continuous mapping theorem gives \[ T_n(x,t,\mathcal{Y}_0) \xrightarrow{d} \sup_{y\in\mathcal{Y}_0}|Z_{n,x,t}(y)|, \] where $Z_{n,x,t}$ is the Gaussian process from Theorem (ref)(b). Under $H_1:f_t(x,\mathcal{Y}_0)>0$, $T_n(x,t,\mathcal{Y}_0)\to\infty$ in probability.
remark[Limiting distribution of the all-$x$ pre-trend test] Theorem (ref).(a) yields with the continuous mapping theorem \begin{equation} \tilde T_n(t,\mathcal{Y}_0)\;\xrightarrow{d}\;\sup_{y\in\mathcal{Y}_0}\left|\mathbb{H}_{1,t}(y)-\mathbb{G}^0_t(y)\right|, \end{equation} The covariance kernel of this limit process is obtained from (ref) by dropping the scalar link factors $\lambda(x'\cdot)$ and replacing each projection $x'\hat{\mathcal I}^{-1}X_k$ by $\hat{\mathcal I}^{-1}X_k$, so that the kernel is $\mathbb{R}^{p\times p}$-valued and the weight term becomes the $p\times p$ matrix $\hat\Theta^{(t)}(y_l)'\hat V_w\,\hat\Theta^{(t)}(y_{l'})$.
remark[Confidence interval] Theorem (ref).(c) yields with some algebraic manipulations \begin{equation*} P(f_t(x) \in CI(t,x))\xrightarrow{d} 1-\alpha \end{equation*} for $f_t(x) > 0.$. Because the bound is reported only after the supremum test rejects, its coverage is to be understood conditional on that rejection.

Imperfect Pre-Treatment Balance

If Assumption (ref) fails with residual $r_n(y):=\theta_{1,T_0}(y)-\sum_{i=2}^{J+1}w_i^*\theta_{i,T_0}(y)\neq 0$, the counterfactual estimator $\hat{f}_t(x)$ is generally not a consistent estimator of $f_t(x)$. Instead, it converges to a pseudo-value $\tilde f_t(x)\neq f_t(x)$. The following corollary gives a precise condition under which consistency for $f_t(x)$ is nevertheless restored.

corollary[Robustness to imperfect balance] Suppose Assumption (ref) and Assumptions (ref)--(ref) hold, $f_t(x)>0$, and Assumption (ref) fails with residual sequence $r_n(y):=\theta_{1,T_0}(y)-\sum_{i=2}^{J+1}w_i^*\theta_{i,T_0}(y)$. Under Assumption (ref), the true counterfactual satisfies $\theta^{0}_{1,t}(y)=\sum_i w_i^*\theta_{i,t}(y)+r_n(y)$, so the scaled bias of $\hat{f}_t(x)$ relative to $f_t(x)$ is \begin{eqnarray} \sqrt{n}\bigl[\tilde{f}_t(x)-f_t(x)\bigr] \nonumber &=& \sqrt{n}\Bigl[2\langle r_n,\Delta_t\rangle_{x} + \int_{\mathcal{Y}_0}\left[\lambda(\eta_t(y))^2+\Delta_t(y|x) \lambda'(\eta_t(y))\right](x'r_n(y))^2\,dy \Bigr] \\ &+& O(\|r_n\|^3), \end{eqnarray} where $\eta_t(y) := x'\theta^{0}_{1,t}(y)$ and $\langle r_n,\Delta_t\rangle_{x}:=\int_{\mathcal{Y}_0}\lambda(x'\eta_t(y))\, x'r_n(y)\,\Delta_t(y\mid x)\,dy$. Then $\sqrt{n}(\hat{f}_t(x)-f_t(x))\xrightarrow{d}\mathcal{N}(0,\sigma^2_t(x))$ as in Theorem (ref)(c) if and only if (ref) converges to zero. Two sufficient conditions are: \begin{itemize} • $\langle r_n,\Delta_t\rangle_x=o(n^{-1/2})$ and $\|r_n\|^2=o(n^{-1/2})$; • $\langle r_n,\Delta_t\rangle_x=0$ and $\|r_n\|=o(n^{-1/4})$. \end{itemize} Condition (b) is strictly weaker than requiring $\|r_n\|=o(n^{-1/2})$, which would be needed without orthogonality to control the first-order term alone.
remark[Economic interpretation] The orthogonality condition $\langle r_n,\Delta\rangle_{x}=0$ holds whenever the pre-treatment misfit $r_n(y)$ is concentrated in regions of the outcome space where the treatment effect $\Delta(y\mid x)$ is negligible, and vice versa. In the New Jersey application, $\Delta(y\mid x_{10})$ is sharply concentrated in the minimum-wage corridor $[\log 4.25,\log 5.10]$. Corollary (ref) therefore implies that the estimate for young low-education workers is robust to pre-treatment misfit outside this corridor, regardless of its magnitude.

Simulation Study

To mimic the empirical application with a Mincer's earnings equation, we consider $J+1=5$ groups and two time periods. For $i=1,\ldots,5$ and $j=0,1$ we have the model:

equation*[equation* omitted — 126 chars of source]

with i.i.d.\ $X_{1,it},X_{2,it},X_{3,it} \sim N(0,1)$ and $\epsilon_{it}$ a noise distribution. This leads to $F_{it}(y|x) = \Phi\bigl(\theta_{it}(y)'x\bigr)$ with $\theta_{it}(y) = \bigl(y-\beta_{0,it},-\beta_{1,it},-\beta_{2,it},-\beta_{3,it}\bigr)$.

The $J=4$ control groups have parameter vectors $\boldsymbol{\beta}_i = \bar{\boldsymbol{\beta}} + c\,\mathbf{e}_i$ for $i=2,\ldots,5$, where $\bar{\boldsymbol{\beta}}=(1,1,1,1)'$, $c=0.8$, and $\mathbf{e}_i$ is the $i$-th standard basis vector in $\mathbb{R}^4$. Each control group thus differs from the others in exactly one coefficient. Under the null hypothesis, the treated group has $\boldsymbol{\beta}_1 = \frac{1}{J}\sum_{i=2}^{5}\boldsymbol{\beta}_i = (1.2,1.2,1.2,1.2)'$, so that PTP holds exactly with equal true weights $w^*_i=1/4$. This design ensures that the Gram matrix $G^*$ has full rank $J=4$ and condition number $\kappa(G^*)=55$, satisfying Assumption (ref).

Under the alternative hypothesis ($H_1$), the treated group's post-treatment outcome receives a covariate-heterogeneous effect $\Delta\cdot \beta_{1,11}$, i.e.\ a shift only along the $X_1$ direction, with $\Delta\in\{0,0.1,\ldots,0.5\}$. This shifts the conditional distribution at covariate values with $X_1\neq0$, while leaving the unconditional (marginal) distribution almost unchanged: since $E[X_1]=0$, the marginal mean and median are preserved and only the spread changes slightly. The design thus isolates a setting in which conditioning on covariates is essential.

The simulation focuses on the size and power properties of the supremum test (ref). We compare three procedures: the conditional supremum test (ref) with correctly specified errors ($\epsilon_{i,t}\sim N(0,1)$, “Probit”) and with misspecified errors ($\epsilon_{i,t}\sim Lo(0,1)$, “Logit”), and an unconditional distributional synthetic-control test that applies the same idea to the marginal outcome distribution (intercept-only DR, $p=1$), the natural in-framework analogue of an unconditional approach in the spirit of Gunsilius2023. The DR working model uses the probit link throughout. We use a grid of $m=10$ equidistant quantile levels of the outcome between $0.05$ and $0.95$, sample sizes $n_{it} \in \{200, 500, 1000\}$, and $N_{\rm mc}=1000$ Monte Carlo replications. Critical values are obtained by Gaussian process simulation with $S=10{,}000$ draws using the estimated covariance matrix $\hat K$.

The conditional tests are evaluated at $x=(1,1,0,0)'$ (so that $x_1=1$ and the conditional shift equals $\Delta$); the unconditional test targets the marginal CDF. Figure (ref) shows that all three procedures control size at $\Delta=0$, whereas the unconditional test is conservative. Under $H_1$, the conditional tests detect the effect with power increasing in $n$ and $\Delta$ (slightly higher under the correctly specified probit link than under the misspecified logit link), whereas the unconditional test has low power: The covariate-heterogeneous effect is invisible in the marginal distribution. Conditioning on covariates is therefore necessary to detect effects that are heterogeneous across the covariate space.

figure[figure omitted — 474 chars of source]

Figure (ref) reports properties of the one-sided 95% confidence interval (ref) for $f_t(x)$ from the conditional probit estimation, under $H_1$ ($\Delta>0$). Coverage is close to the nominal level across $\Delta$; the small deviations at the smallest effect sizes reflect the positive finite-sample bias of $\hat f$, which is a mean of squared quantities. The mean half-width falls with $n$ and grows with $\Delta$, as expected from the asymptotic theory. In practice the CI is reported only when the supremum test rejects $H_0$, which already selects cases where $f_0$ is not negligible relative to estimation error.

figure[figure omitted — 362 chars of source]

Under $T_0=2$ pre-treatment periods, the $T_0^{-1}$-scaling of $\Sigma_{G,c}$ in Lemma (ref) halves the asymptotic variance of the Gram-matrix estimator, and hence of $V_w$, of the weight estimator relative to $T_0=1$. We verified in an auxiliary simulation (same DGP, $T_0=2$, second pre-period drawn i.i.d.\ with the same parameters) that size remains well controlled at the nominal 5% level and power is modestly improved, consistent with the theoretical efficiency gain.

Empirical Study: New Jersey Minimum Wage 1992

Background and Data

In April 1992, New Jersey raised its state minimum wage from \$4.25 to \$5.05 per hour (an increase of 19%) while the federal minimum remained unchanged at \$4.25. This policy change has been studied in the literature before. CardKrueger1994 find no negative employment effects using a DiD comparison of fast-food restaurants with neighbouring Pennsylvania, contrasting findings in NeumarkWascher1992. The DR-SC approach taken here offers a complementary perspective: rather than average employment effects, we trace the full conditional wage distribution across different worker types, allowing us to assess where in the distribution the policy had bite.

The distributional consequences of minimum wage increases have been widely studied: DiNardoFortinLemieux1996 document a link between minimum wage erosion and rising wage inequality using CPS data, and AutorManningSmith2016 provide evidence that the minimum wage compresses the lower tail of the wage distribution. ChernozhFVMelly2013 study counterfactual wage distributions using distribution regression. The contribution of the present approach is a clearer causal identification in terms of the workers' characteristics.

The validity of PTP in this application rests on the observation that the three pre-treatment years (1989--1992) span a well-defined common macroeconomic episode: the onset and trough of the 1990--91 recession, followed by an early recovery. This recession depressed low-education wages across all US states, but with differential intensity depending on each state's exposure to interest-rate-sensitive sectors (construction, durable manufacturing) and to the contemporaneous defense-spending contraction BlanchardKatz1992. These cross-state differences in cyclical exposure correspond to the state-specific factor loadings $\Lambda_i(y)$ in the latent-factor model of Remark 2.2: New Jersey's exposure profile (concentrated in finance, insurance, and real estate) is approximated by the weighted combination of donor states with Florida (finance and real estate) receiving the largest positive weight and South Dakota and Missouri (manufacturing-heavy, lower cyclical correlation with New Jersey) receiving negative weights, as reported in Table (ref). The convincing pre-trend tests in Section (ref) confirm that this synthetic New Jersey tracks the observed conditional wage distribution across all three pre-treatment years.

We use data from the CPS Basic Monthly files (flood2025ipums) organised into April--March policy years aligned with the reform date: $t=1$ (Apr 1989--Mar 1990), $t=2$ (Apr 1990--Mar 1991) and $t=3$ (Apr 1991--Mar 1992) are pre-treatment ($T_0=3$), and $t=4$ (Apr 1992--Mar 1993) is post-treatment ($T_1=1$). This periodisation aligns the start of the post-period exactly with the April 1992 increase, so that no partially-treated calendar year has to be discarded. We restrict to workers with a valid hourly wage, at most \$99/hour and non-missing education and experience. The outcome $Y$ is the log nominal hourly wage. The covariate vector $X$ includes a constant, years of education, years of potential experience, and its square, all standardised to zero mean and unit variance using the pooled sample. The standardised experience-squared variable is constructed as the square of standardised experience, $\widetilde{x_3}=\bigl((\mathrm{exper}-\bar x_2)/s_2\bigr)^2$. New Jersey policy-year sample sizes are 3{,}852, 3{,}998 and 4{,}031 (pre-treatment) and 3{,}905 (post-treatment).

The donor pool consists of $J=42$ US states, excluding eight states or districts with their own minimum wage changes during 1989--1993 (Connecticut, Massachusetts, Rhode Island, New Hampshire, Vermont, Alaska, Hawaii, and the District of Columbia).

As evaluation points, we consider the median worker (education = 12, experience = 10) as well as the four combinations of the $10\%$ and $90 \%$ values of education (10, 16) and experience (2, 37).

Weight Estimation and Pre-Treatment Fit

We estimate the DR parameters on a grid of $m=32$ quantiles of the pooled log-wage distribution (10th--90th percentiles). The total sample size is $n=381{,}953$ across all group-time cells. The Gram matrix $\hat G$ is the time-average over the three pre-treatment periods (see Section 2.2). It has full rank $J=42$, condition number $\kappa(\hat G)=9.6\times10^4$, and is numerically stable without regularisation. Table (ref) reports the five largest weights by magnitude.

table[table omitted — 272 chars of source]

Florida receives the largest positive weight, with New York second. The full weight distribution is shown in Figure (ref). 20 of 42 states receive negative weights, reflecting that New Jersey's pre-treatment wage structure is not fully within the convex hull of the donor pool, which the unconstrained formulation handles naturally.

figure[figure omitted — 236 chars of source]

Due to the post-period evidence from (ref), the evaluation point $x_{10}$ (educ=10, exper=2), reflecting low-education young workers, is of particular interest. Figure (ref) shows the pre-treatment fit for this point in all three pre-treatment years. The synthetic control tracks the observed New Jersey CDF closely across all pre-treatment years, with maximum sup-distance $\sup_y|\hat F_{\rm NJ,t}(y|x_{10})- \hat F^0_{\rm NJ,t}(y|x_{10})|\leq 0.037$ in each pre-period.

figure[figure omitted — 357 chars of source]

Pre-Trend Tests

Figure (ref) reports the full-distribution and focused (MW-corridor) $p$-values of $T_n(x,t,\mathcal{Y})$ for the two pseudo-post pre-periods (Apr89/90$\to$Apr90/91 and $\to$Apr91/92) and the post-treatment period ($\to$Apr92/93), at all five evaluation points. The corridor for the focused test is fixed a priori to $\mathcal{Y}_0=[\log\$4.25,\log\$5.10]$ by the policy (raising the floor from \$4.25 to \$5.05).

Both pseudo-post tests do not reject $H_0$ at the five evaluation points, on the full distribution and on the corridor (all $p>0.35$; at $x_{10}$, $p=0.74$ and $p=0.54$ on $\mathcal{Y}$, and $p=0.81$ and $p=0.70$ on $\mathcal{Y}_0$). The 1990--91 recession is absorbed within the twelve-month windows rather than surfacing at a period transition, and the late-1991 announcement window (contained in the third pre-period) produces no detectable corridor pre-trend. Unlike a calendar-year periodisation, the design therefore requires no reconciliation of a full-distribution pre-trend rejection with the focused test. It seems that the parallel-trends assumption holds directly, both overall and on the corridor.

Beyond the five evaluation points, we also apply the $x$-simultaneous pre-trend test (ref), which checks parallel trends in the full DR parameter vector and hence for all covariate values at once. Neither pseudo-post transition rejects at the $10\%$ level: $\tilde T_n=144.9$ against $\hat c_{0.90}=245.3$ ($p=0.75$) for Apr89/90$\to$Apr90/91, and $\tilde T_n=209.6$ against $\hat c_{0.90}=241.7$ ($p=0.22$) for $\to$Apr91/92. So, it is reasonable to assume that parallel trends in parameters holds in the pre-periods not only at the points of interest but jointly across the covariate space.

figure[figure omitted — 473 chars of source]

Figure (ref) contrasts the full-distribution and focused tests at $x_{10}$. The focused statistic $T_n(x,t,\mathcal{Y}_0)$ serves two distinct roles. In the pre-periods it is the empirical counterpart of the orthogonality condition in Corollary (ref): should a full-distribution pre-trend reject while the corridor test passes, there is evidence that a potential pre-treatment misfit $r_n$ is orthogonal to the treatment effect $\Delta_t$ on $\mathcal{Y}_0$, and $\hat f_t(x,\mathcal{Y}_0)$ that remains consistent and asymptotically normal even where the Span Condition fails on the full support $\mathcal{Y}$. Here both pre-trend tests pass, so this safeguard is not needed and the corridor effect is identified without it. In the post-period the focused test instead provides power: It concentrates the supremum on the interval where the policy operates. At $x_{10}$ the full-distribution test is only marginal ($p=0.053$) while the focused test rejects clearly ($p=0.011$), with the supremum attained inside the corridor.

figure[figure omitted — 612 chars of source]

Main Results: Post-Treatment CDFs and Heterogeneity Patterns for the Treatment Effects

Figure (ref) displays the pointwise CDF differences $\hat\Delta_t(y\mid x):=\hat F_{1,t}(y\mid x)-\hat F^0_{1,t}(y\mid x)$ with 90% pointwise confidence bands for all five evaluation points. Table (ref) summarises the scalar treatment effect estimates $\hat f_t(x,\mathcal{Y})$ and the supremum test results. The effect is sharply concentrated in the MW corridor for the young low-educated group $x_{10}$, where $\hat\Delta_t$ falls by up to $0.12$ inside $[\$4.25,\$5.10]$. Both the focused test ($p_{\mathcal{Y}_0}=0.011$) and the full-distribution test ($p=0.053$) reject at the $10\%$ level. The high-education/high-experience point $x_{90}$ shows no effect on either test ($p=0.83$, $p_{\mathcal{Y}_0}=0.84$). We report the one-sided $90\%$ lower confidence bound $\mathrm{LCB}=\max\{0,\hat f_t-z_{0.90}\hat{\mathrm{se}}\}$ exactly when the supremum test rejects at $10\%$. This is only the case for $x_{10}$. Figure (ref) summarises the treatment effect estimates visually.

figure[figure omitted — 301 chars of source]
table[table omitted — 1,258 chars of source]
figure[figure omitted — 326 chars of source]

\paragraph{Low-education, low-experience workers ($x_{10}$: educ=10, exper=2).} The clearest effect is at $x_{10}$: $\hat f_t=1.71\times10^{-3}$, with the focused test rejecting on the corridor ($p_{\mathcal{Y}_0}=0.012$) and the full-distribution test rejecting at the $10\%$ level ($T_n=73.0$, $\hat c_{0.90}=64.8$, $p=0.054$); the one-sided $90\%$ lower bound is $0.27\times10^{-3}$ (full support) and $0.97\times10^{-3}$ (corridor). The shape is the canonical minimum-wage spike: $\hat\Delta_t(y)$ falls by up to $0.12$ inside the corridor $[\$4.25,\$5.10]$ and is near zero outside it. Probability mass previously below \$5.10 is shifted up to just above the new floor. The supremum is attained within the corridor, so the focused test is the sharper instrument here.

\paragraph{High-education, low-experience workers ($x_{hl}$: educ=16, exper=2).} The point estimate $\hat f_t=1.55\times10^{-3}$ is sizeable but not significant ($T_n=37.1$, $p=0.686$; corridor $p_{\mathcal{Y}_0}=0.150$), with a wide band reflecting the small high-education--low-experience cell. The pointwise confidence intervals of the CDF differences lie largely above zero, which would be consistent with the wage-compression/ripple pattern of AutorManningSmith2016.

\paragraph{Median worker ($x_0$: educ=12, exper=10).} The median worker shows a small effect that is marginal on the corridor only: $\hat f_t=0.58\times10^{-3}$ ($T_n=26.7$, $p=0.294$; $p_{\mathcal{Y}_0}=0.064$), consistent with a modest direct effect for the share of median workers whose wages were near the old floor.

\paragraph{Low-education, high-experience workers ($x_{lh}$: educ=10, exper=37).} The estimate is small, $\hat f_t=0.61\times10^{-3}$. The full-distribution test does not reject ($T_n=28.6$, $p=0.750$); the focused test is only marginal ($p_{\mathcal{Y}_0}=0.060$, rejecting at the $10\%$ level), and the one-sided lower bound is zero ($\mathrm{LCB}=0$, i.e.\ the interval $[0,\infty)$). There is no firmly resolved effect near the MW threshold for this group, consistent with their wages lying well above \$5.10.

\paragraph{High-education, high-experience workers ($x_{90}$: educ=16, exper=37).} For the high-education/high-experience point, we have no evidence for an effect: $\hat f_t=0.21\times10^{-3}$ with no rejection on either test ($T_n=27.6$, $\hat c_{0.90}=80.2$, $p=0.824$; corridor $p_{\mathcal{Y}_0}=0.846$). The absence of any corridor effect for a group whose wages lie far above the floor is as expected and supports the identifying assumptions.

\paragraph{Overall pattern.} The DR-SC approach reveals a clear heterogeneity pattern. Under the April--March design the minimum-wage effect is sharply and specifically concentrated in the MW corridor for low-education/low-experience young workers ($x_{10}$), the group most directly affected by the policy. For the high-education/high-experience group ($x_{90}$), there is no evidence for an effect, and the remaining groups show at most marginal corridor effects. Relative to a calendar-year periodisation, which we have also considered, less of the response is attributed to broad concurrent shifts, consistent with the cleaner pre-trends documented above.

This complements the aggregate DiD evidence of CardKrueger1994: While their design isolates average employment effects, the DR-SC approach reveals precisely which part of the conditional wage distribution was affected, and for which worker groups.

Regarding the interpretation as causal effect, under PTP, $\hat f_t(x,\mathcal{Y})$ estimates the total causal effect of all NJ-specific developments relative to the synthetic control between the pre- and post-treatment periods. For low-education/low-experience workers, the minimum wage increase is the most plausible dominant cause. For other groups, the effects may reflect a combination of the MW and concurrent NJ-specific developments. The DR-SC method provides a clean characterisation of where in the distribution the effects occur, which is informative about the operative mechanisms even when causal attribution is uncertain.

Robustness: Ridge-Augmented Weights

A variance decomposition of (ref) shows the pointwise bands are dominated by the weight-estimation term $V_w$ ($40$--$60\%$ of the pointwise variance across the five points, against only $\approx15\%$ for the treated cell), reflecting the ill-conditioned Gram matrix ($\kappa(\hat G)\approx10^5$). As a robustness check we re-estimate with the ridge weights of Remark (ref), selecting $\lambda$ by leave-one-period-out cross-validation on the pre-treatment fit.

The chosen $\lambda^*\approx3.8\times10^{-3}$ (of order $n^{-1/2}$) lowers $\kappa$ to $1.1\times10^4$, evens out the weights (the largest falls from $0.52$ to $0.24$), and improves the held-out pre-fit by about $20\%$, indicating that the unregularised weights mildly overfit. The substantive conclusions are unchanged: The effect stays concentrated in the corridor for $x_{10}$, no evidence for an effect for $x_{90}$, and all pre-trend tests still pass. Regularisation narrows the pointwise bands by roughly a third and sharpens the $x_{10}$ inference (its one-sided lower bound moves well above zero). Because $\lambda^*$ introduces a small, deliberate shrinkage of the weights towards equality (in the spirit of augmented synthetic control BenMichaelEtAl2021), we report it as a robustness analysis rather than the main specification.

Conclusion

We have proposed a SC estimator for conditional distribution functions in the semiparametric DR framework, with three main contributions. First, the Parallel Trends in Parameters assumption keeps the counterfactual within the model class, and dropping non-negativity on weights yields a closed-form estimator. Second, both DR estimation error and weight estimation error contribute at the same $\sqrt{n}$ rate to the variance of the counterfactual, and both components are characterised explicitly. Third, a two-stage inference procedure is proposed: a supremum test whose null distribution is approximated by Gaussian process simulation, followed by a plug-in confidence interval for the integrated difference when the null is rejected. The supremum statistic has a valid non-degenerate limit under $H_0$, grows to infinity under $H_1$, and yields $p$-values that fully exploit the large-$n$ precision of the DR estimators.

The New Jersey application illustrates how the proposed framework can uncover heterogeneous distributional patterns across the covariate space. The test statistic $T_n(x,t,\mathcal{Y}_0)$ with point-specific critical values from GP simulation reveals that the estimated effect is concentrated in the minimum-wage corridor for young low-educated workers ($x_{10}$: focused $p_{\mathcal{Y}_0}=0.012$ vs.\ full $p=0.054$), consistent with a direct minimum wage effect, while the other groups show at most marginal corridor effects and the high-education/experience group does not show any noticeable effect ($p_{\mathcal{Y}_0}=0.846$). Pre-trend tests pass on both the full distribution and the corridor. These patterns are invisible in aggregate SC analyses.

Future directions include the theory and applications of uniform confidence bands for $f(x)$ over a covariate set $\mathcal{X}$, enabling formal tests for x-heterogeneity of the treatment effect, and a nonparametric estimation of the link function in the DR model.