EconBase
← All papers

Using Wasserstein Generative Adversarial Networks for the Design of Monte Carlo Simulations

Susan Athey, Guido Imbens, Jonas Metzger, Evan Munro

arXiv 5 Sep 2019 · Econometrics · publishedJournal of Econometrics (2021) · 26 citations (OpenAlex)

arXiv:1909.02210 · PDF · DOI · OpenAlex · Extracted main text

Abstract

When researchers develop new econometric methods it is common practice to compare the performance of the new methods to those of existing methods in Monte Carlo studies. The credibility of such Monte Carlo studies is often limited because of the freedom the researcher has in choosing the design. In recent years a new class of generative models emerged in the machine learning literature, termed Generative Adversarial Networks (GANs) that can be used to systematically generate artificial data that closely mimics real economic datasets, while limiting the degrees of freedom for the researcher and optionally satisfying privacy guarantees with respect to their training data. In addition if an applied researcher is concerned with the performance of a particular statistical method on a specific data set (beyond its theoretical properties in large samples), she may wish to assess the performance, e.g., the coverage rate of confidence intervals or the bias of the estimator, using simulated data which resembles her setting. Tol illustrate these methods we apply Wasserstein GANs (WGANs) to compare a number of different estimators for average treatment effects under unconfoundedness in three distinct settings (corresponding to three real data sets) and present a methodology for assessing the robustness of the results. In this example, we find that (i) there is not one estimator that outperforms the others in all three settings, so researchers should tailor their analytic approach to a given setting, and (ii) systematic simulation studies can be helpful for selecting among competing methods in this situation.

Citation extraction

60
references
85
in-text mentions
60
distinct cited
9
self-citations
12,958
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Ward… (2014) Generative adversarial nets0.87472100%
2Alberto Abadie and Guido W Imbens (2011) Bias-corrected matching estimators for average treatment effects self0.81142100%
3Martin Arjovsky, Soumith Chintala, and Léon Bottou (2017) Wasserstein gan0.81142100%
4Michael Knaus, Michael Lechner, and Anthony Strittmatter (2018) Machine learning estimation of heterogeneous causal effects: Empirical monte carlo evidence0.73732100%
5Susan Athey, Guido W Imbens, and Stefan Wager (2018) Approximate residual balancing: debiased inference of average treatment effects in high dimensions self0.64422100%
6Rajeev H Dehejia and Sadek Wahba (2002) Propensity score-matching methods for nonexperimental causal studies0.64422100%
7Rajeev H Dehejia and Sadek Wahba (1999) Causal effects in nonexperimental studies: Reevaluating the evaluation of training programs0.64422100%
8Keisuke Hirano, Guido W Imbens, and Geert Ridder (2003) Efficient estimation of average treatment effects using the estimated propensity score self0.64422100%
9Robert J LaLonde (1986) Evaluating the econometric evaluations of training programs with experimental data0.64422100%
10Alejandro Schuler, Ken Jung, Robert Tibshirani, Trevor Hastie, and N… (2017) Synth-validation: Selecting the best causal inference method for a given dataset0.58531100%

Showing the top 10 of 60 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Policy Learning with $$-Expected Welfare0.87463
2PLRD: Partially Linear Regression Discontinuity Inference0.81142
3Using Forests in Multivariate Regression Discontinuity Designs0.778173
4Simultaneous Inference for Local Structural Parameters with Random Forests$^*$0.73732
5ASSESSING INFERENCE METHODS0.64422
6Time Series (re)sampling using Generative Adversarial Networks0.64422
7Semiparametric Estimation of Long-Term Treatment Effects$^*$0.51132
8Causal Inference for Spatial Treatments0.40511
9Gaussian Transforms Modeling and the Estimation of Distributional Regression Functions0.40511
10Invidious Comparisons: Ranking and Selection as Compound Decisions0.40511