Gentry Johnson, Brian Quistorff, Matt Goldman
arXiv 28 May 2019 · Econometrics
arXiv:1905.12020 · PDF · DOI · OpenAlex · Extracted main text
When pre-processing observational data via matching, we seek to approximate each unit with maximally similar peers that had an alternative treatment status--essentially replicating a randomized block design. However, as one considers a growing number of continuous features, a curse of dimensionality applies making asymptotically valid inference impossible (Abadie and Imbens, 2006). The alternative of ignoring plausibly relevant features is certainly no better, and the resulting trade-off substantially limits the application of matching methods to "wide" datasets. Instead, Li and Fu (2017) recasts the problem of matching in a metric learning framework that maps features to a low-dimensional space that facilitates "closer matches" while still capturing important aspects of unit-level heterogeneity. However, that method lacks key theoretical guarantees and can produce inconsistent estimates in cases of heterogeneous treatment effects. Motivated by straightforward extension of existing results in the matching literature, we present alternative techniques that learn latent matching features through either MLPs or through siamese neural networks trained on a carefully selected loss function. We benchmark the resulting alternative methods in simulations as well as against two experimental data sets--including the canonical NSW worker training program data set--and find superior performance of the neural-net-based methods.
appendix boundary found by appendix_command · 72% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Li, S. and Fu, Y (2017) Matching on balanced nonlinear representations for treatment effects estimation | 1.000 | 8 | 5 | 100% |
| 2 | Abadie, A. and Imbens, G. W (2006) Large sample properties of matching estimators for average treatment effects | 0.928 | 4 | 4 | 100% |
| 3 | Antonelli, J., Cefalu, M., Palmer, N., and Agniel, D (2016) Doubly robust matching estimators for high dimensional confounding adjustment | 0.405 | 1 | 1 | 100% |
| 4 | Cybenko, G (1989) Approximation by superpositions of a sigmoidal function | 0.405 | 1 | 1 | 100% |
| 5 | Dehejia, R. H. and Wahba, S (1999) Causal effects in nonexperimental studies: Reevaluating the evaluation of training programs | 0.405 | 1 | 1 | 100% |
| 6 | Deng, H. and Runger, G (2012) Feature selection via regularized trees | 0.405 | 1 | 1 | 100% |
| 7 | Gu, X. S. and Rosenbaum, P. R (1993) Comparison of multivariate matching methods: Structures, distances, and algorithms | 0.405 | 1 | 1 | 100% |
| 8 | Hansen, B. B (2008) The prognostic analogue of the propensity score | 0.405 | 1 | 1 | 100% |
| 9 | Hinton, G. E. and Salakhutdinov, R. R (2006) Reducing the dimensionality of data with neural networks | 0.405 | 1 | 1 | 100% |
| 10 | Johansson, F., Shalit, U., and Sontag, D (2016) Learning representations for counterfactual inference | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 24 scored citations.