Jiafeng Chen, Jonathan Roth, Jann Spiess
arXiv 31 Dec 2025 · Econometrics
arXiv:2512.25032 · PDF · DOI · OpenAlex · Extracted main text
We consider the extent to which we can learn from a completely randomized experiment whether all individuals have treatment effects that are weakly of the same sign, a condition we call monotonicity. From a classical sampling perspective, it is well-known that monotonicity is not falsifiable. By contrast, we show from the design-based perspective -- in which the units in the population are fixed and only treatment assignment is stochastic -- that the distribution of treatment effects in the finite population (and hence whether monotonicity holds) is formally identified. We argue, however, that the usual definition of identification is unnatural in the design-based setting because it imagines knowing the distribution of outcomes over different treatment assignments for the same units. We thus evaluate the informativeness of the data by the extent to which it enables frequentist testing and Bayesian updating. We show that frequentist tests can have nontrivial power against some alternatives, but power is generically limited. Likewise, we show that there exist (non-degenerate) Bayesian priors that never update about whether monotonicity holds. We conclude that, despite the formal identification result, the ability to learn about monotonicity from data in practice is severely limited.
appendix boundary found by appendix_command · 49% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Copas, J. B (1973) Randomization models for the matched and unmatched 2 × 2 tables | 0.843 | 4 | 3 | 75% |
| 2 | Ding, Peng and Miratrix, Luke W (2019) Model-free causal inference of binary experimental data | 0.843 | 3 | 3 | 100% |
| 3 | Christy, Neil and Kowalski, Amanda Ellen (2025) Counting Defiers: A Design-Based Model of an Experiment Can Reveal Evidence Beyond the Average Effect | 0.811 | 4 | 2 | 100% |
| 4 | Heckman, James J. and Smith, Jeffrey and Clements, Nancy (1997) Making The Most Out Of Programme Evaluations and Social Experiments: Accounting For Heterogeneity in Programme Impacts | 0.644 | 2 | 2 | 100% |
| 5 | Caughey, Devin and Dafoe, Allan and Li, Xinran and Miratrix, Luke (2023) Randomisation inference beyond the sharp null: bounded null hypotheses and quantiles of individual treatment effects | 0.511 | 2 | 1 | 100% |
| 6 | Joshua Angrist and Guido Imbens (1994) Identification and Estimation of Local Average Treatment Effects | 0.405 | 1 | 1 | 100% |
| 7 | Abadie, Alberto and Athey, Susan and Imbens, Guido W. and Wooldridge… (2020) Sampling-Based versus Design-Based Uncertainty in Regression Analysis | 0.405 | 1 | 1 | 100% |
| 8 | Angrist, Joshua D and Imbens, Guido W (1995) Two-stage least squares estimation of average causal effects in models with variable treatment intensity | 0.405 | 1 | 1 | 100% |
| 9 | Gelman, Andrew and Mikhaeil, Jonas M (2025) Russian roulette: the need for stochastic potential outcomes when utilities depend on counterfactuals | 0.405 | 1 | 1 | 100% |
| 10 | Kitagawa, Toru (2015) A Test for Instrument Validity | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 18 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Finite Population Identification and Design-Based Sensitivity Analysis | 0.874 | 7 | 2 |
| 2 | Counting Defiers: A Design-Based Model of an Experiment Can Reveal Evidence Beyond the Average Effect | 0.481 | 6 | 2 |