Max H. Farrell, Tengyuan Liang, Sanjog Misra
arXiv 26 Sep 2018 · Econometrics · 44 citations (OpenAlex)
arXiv:1809.09953 · PDF · DOI · OpenAlex · Extracted main text
We study deep neural networks and their use in semiparametric inference. We establish novel rates of convergence for deep feedforward neural nets. Our new rates are sufficiently fast (in some cases minimax optimal) to allow us to establish valid second-step inference after first-step estimation with deep learning, a result also new to the literature. Our estimation rates and semiparametric inference results handle the current standard architecture: fully connected feedforward neural networks (multi-layer perceptrons), with the now-common rectified linear unit activation function and a depth explicitly diverging with the sample size. We discuss other architectures as well, including fixed-width, very deep networks. We establish nonasymptotic bounds for these deep nets for a general class of nonparametric regression-type loss functions, which includes as special cases least squares, logistic regression, and other generalized linear models. We then apply our theory to develop semiparametric inference, focusing on causal parameters for concreteness, such as treatment effects, expected welfare, and decomposition effects. Inference in many other semiparametric contexts can be readily obtained. We demonstrate the effectiveness of deep learning with a Monte Carlo analysis and an empirical application to direct mail marketing.
appendix boundary found by appendix_command · 72% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Hitsch, G. J. and S. Misra (2018) Heterogeneous Treatment Effects and Optimal Targeting Policy Evaluation | 1.000 | 7 | 3 | 100% |
| 2 | Farrell, M. H (2015) Robust Inference on Average Treatment Effects with Possibly More Covariates than Observations self | 1.000 | 6 | 3 | 100% |
| 3 | Yarotsky, D (2017) Error bounds for approximations with deep ReLU networks | 0.874 | 6 | 4 | 67% |
| 4 | Bartlett, P. L., N. Harvey, C. Liaw, and A. Mehrabian (2017) Nearly-tight VC-dimension bounds for piecewise linear neural networks, in | 0.874 | 6 | 3 | 67% |
| 5 | Chen, X. and H. White (1999) Improved rates and asymptotic normality for nonparametric neural network estimators | 0.874 | 5 | 2 | 100% |
| 6 | Yarotsky, D (2018) Optimal approximation of continuous functions by very deep ReLU networks | 0.843 | 5 | 4 | 60% |
| 7 | Belloni, A., V. Chernozhukov, I. Fernández-Val, and C. Hansen (2017) Program Evaluation and Causal Inference With High-Dimensional Data | 0.843 | 3 | 3 | 100% |
| 8 | Athey, S. and S. Wager (2018) Efficient Policy Learning | 0.811 | 4 | 2 | 100% |
| 9 | Anthony, M. and P. L. Bartlett (1999) Neural Network Learning: Theoretical Foundations | 0.737 | 3 | 3 | 67% |
| 10 | Belloni, A., V. Chernozhukov, and C. Hansen (2014) Inference on Treatment Effects after Selection Amongst High-Dimensional Controls | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 96 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.