Susan Athey, Guido Imbens, Jonas Metzger, Evan Munro
arXiv 5 Sep 2019 · Econometrics · publishedJournal of Econometrics (2021) · 26 citations (OpenAlex)
arXiv:1909.02210 · PDF · DOI · OpenAlex · Extracted main text
When researchers develop new econometric methods it is common practice to compare the performance of the new methods to those of existing methods in Monte Carlo studies. The credibility of such Monte Carlo studies is often limited because of the freedom the researcher has in choosing the design. In recent years a new class of generative models emerged in the machine learning literature, termed Generative Adversarial Networks (GANs) that can be used to systematically generate artificial data that closely mimics real economic datasets, while limiting the degrees of freedom for the researcher and optionally satisfying privacy guarantees with respect to their training data. In addition if an applied researcher is concerned with the performance of a particular statistical method on a specific data set (beyond its theoretical properties in large samples), she may wish to assess the performance, e.g., the coverage rate of confidence intervals or the bias of the estimator, using simulated data which resembles her setting. Tol illustrate these methods we apply Wasserstein GANs (WGANs) to compare a number of different estimators for average treatment effects under unconfoundedness in three distinct settings (corresponding to three real data sets) and present a methodology for assessing the robustness of the results. In this example, we find that (i) there is not one estimator that outperforms the others in all three settings, so researchers should tailor their analytic approach to a given setting, and (ii) systematic simulation studies can be helpful for selecting among competing methods in this situation.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Ward… (2014) Generative adversarial nets | 0.874 | 7 | 2 | 100% |
| 2 | Alberto Abadie and Guido W Imbens (2011) Bias-corrected matching estimators for average treatment effects self | 0.811 | 4 | 2 | 100% |
| 3 | Martin Arjovsky, Soumith Chintala, and Léon Bottou (2017) Wasserstein gan | 0.811 | 4 | 2 | 100% |
| 4 | Michael Knaus, Michael Lechner, and Anthony Strittmatter (2018) Machine learning estimation of heterogeneous causal effects: Empirical monte carlo evidence | 0.737 | 3 | 2 | 100% |
| 5 | Susan Athey, Guido W Imbens, and Stefan Wager (2018) Approximate residual balancing: debiased inference of average treatment effects in high dimensions self | 0.644 | 2 | 2 | 100% |
| 6 | Rajeev H Dehejia and Sadek Wahba (2002) Propensity score-matching methods for nonexperimental causal studies | 0.644 | 2 | 2 | 100% |
| 7 | Rajeev H Dehejia and Sadek Wahba (1999) Causal effects in nonexperimental studies: Reevaluating the evaluation of training programs | 0.644 | 2 | 2 | 100% |
| 8 | Keisuke Hirano, Guido W Imbens, and Geert Ridder (2003) Efficient estimation of average treatment effects using the estimated propensity score self | 0.644 | 2 | 2 | 100% |
| 9 | Robert J LaLonde (1986) Evaluating the econometric evaluations of training programs with experimental data | 0.644 | 2 | 2 | 100% |
| 10 | Alejandro Schuler, Ken Jung, Robert Tibshirani, Trevor Hastie, and N… (2017) Synth-validation: Selecting the best causal inference method for a given dataset | 0.585 | 3 | 1 | 100% |
Showing the top 10 of 60 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.