arXiv 23 Apr 2026 · Econometrics
arXiv:2604.21548 · PDF · DOI · OpenAlex · Extracted main text
Treatment effect distributions are not identified without restrictions on the joint distribution of potential outcomes. Existing approaches either impose rank preservation -- a strong assumption -- or derive partial identification bounds that are often wide. We show that a single scalar parameter, rank stickiness, suffices for nonparametric point identification while permitting rank violations. The identified joint distribution -- the coupling that maximizes average rank correlation subject to a relative entropy constraint, which we call the Bregman-Sinkhorn copula -- is uniquely determined by the marginals and rank stickiness. Its conditional distribution is an exponential tilt of the marginal with a Bregman divergence as the exponent, yielding closed-form conditional moments and rank violation probabilities; the copula nests the comonotonic and Gaussian copulas as special cases. The empirical Bregman-Sinkhorn copula converges at the parametric $\sqrt{n}$-rate with a Gaussian process limit, despite the infinite-dimensional parameter space. We apply the framework to estimate the full treatment effect distribution, derive a variance estimator for the average treatment effect tighter than the Fréchet--Hoeffding and Neyman bounds, and extend to observational studies under unconfoundedness.
appendix boundary found by appendix_command · 72% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | J. J. Heckman, J. Smith, and N. Clements (1997) Making the most out of programme evaluations and social experiments: Accounting for heterogeneity in programme impacts | 0.737 | 3 | 2 | 100% |
| 2 | S. Athey and G. W. Imbens (2006) Identification and inference in nonlinear difference-in-differences models | 0.644 | 2 | 2 | 100% |
| 3 | V. Chernozhukov and C. Hansen (2005) An IV model of quantile treatment effects | 0.644 | 2 | 2 | 100% |
| 4 | Y. Brenier (1991) Polar factorization and monotone rearrangement of vector-valued functions | 0.511 | 2 | 2 | 50% |
| 5 | A. W. van der Vaart and J. A. Wellner (1996) Weak convergence and empirical processes: with applications to statistics | 0.511 | 2 | 2 | 50% |
| 6 | N. Deb and T. Liang (2025) No-regret generative modeling via parabolic monge-ampère pde | 0.405 | 1 | 1 | 100% |
| 7 | M. H. Farrell, T. Liang, and S. Misra (2021) Deep neural networks for estimation and inference | 0.405 | 1 | 1 | 100% |
| 8 | T. Liang (2025) Distributional shrinkage i: Universal denoisers in multi-dimensions | 0.405 | 1 | 1 | 100% |
| 9 | T. Liang (2025) Distributional shrinkage ii: Optimal transport denoisers with higher-order scores | 0.405 | 1 | 1 | 100% |
| 10 | S. Athey and G. W. Imbens (2017) The state of applied econometrics: Causality and policy evaluation | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 26 scored citations.