Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
110,848 characters · 22 sections · 97 citation commands
Deep Neural Networks for Estimation and Inference
Keywords: Deep Learning, Neural Networks, Rectified Linear Unit, Nonasymptotic Bounds, Convergence Rates, Semiparametric Inference, Treatment Effects, Program Evaluation, Treatment Targeting.
\setcounter{page}{0} \thispagestyle{empty}
\doublespacing
Statistical machine learning methods are being rapidly integrated into the social and medical sciences. Economics is no exception, and there has been a recent surge of research that applies and explores machine learning methods in the context of econometric modeling, particularly in “big data” settings. Furthermore, theoretical properties of these methods are the subject of intense recent study. This has netted several breakthroughs both theoretically, such as robust, valid inference following machine learning, and in novel applications and conclusions. Our goal in the present work is to study a particular statistical machine learning technique which is widely popular in industrial applications, but less frequently used in academic work and largely ignored in recent theoretical developments on inference: deep neural networks. To our knowledge we provide the first inference results using deep learning methods.
Neural networks are estimation methods that model the relationship between inputs and outputs using layers of connected computational units (neurons), patterned after the biological neural networks of brains. These computational units sit between the inputs and output and allow data-driven learning of the appropriate model, in addition to learning the parameters of that model. Put into terms more familiar in nonparametric econometrics: neural networks can be thought of as a (complex) type of sieve estimation where the basis functions are flexibly learned from the data. Neural networks are perhaps not as familiar to economists as other methods, and indeed, were out of favor in the machine learning community for several years, returning to prominence only very recently in the form of deep learning. Deep neural nets contain many hidden layers of neurons between the input and output layers, and have been found to exhibit superior performance across a variety of contexts. Our work aims to bring wider attention to these methods and to take the first step toward filling the gaps in the theoretical understanding of inference using deep neural networks. Our results can be used in many economic contexts, including selection models, games, consumer surplus, and dynamic discrete choice.
Before the recent surge in attention, neural networks had taken a back seat to other methods (such as kernel methods or forests) largely because of their modest empirical performance and challenging optimization. However, the availability of scalable computing and stochastic optimization techniques lecun1998gradient,kingma2014adam and the change from smooth sigmoid-type activation functions to rectified linear units (ReLU), $x \mapsto \max(x,0)$ nair2010rectified, have seemingly overcome optimization hurdles and now this form of deep learning matches or sets the state of the art in many prediction contexts krizhevsky2012imagenet, he2016identity. Our theoretical results speak directly to this modern implementation of deep learning: we explicitly model the depth of the network as diverging with the sample size and focus on the ReLU activation function.
Further back in history, before falling out of favor, neural networks were widely studied and applied, particularly in the 1990s. In that time, shallow neural networks with smooth activation functions were shown to have many good theoretical properties. Intuitively, neural networks are a form of sieve estimation, wherein basis functions of the original variables are used to approximate unknown nonparametric objects. What sets neural nets apart is that the basis functions are themselves learned from the data by optimizing over many flexible combinations of simple functions. It has been known for some time that such networks yield universal approximations hornik1989multilayer. Comprehensive theoretical treatments are given by white1992artificial and Anthony-Bartlett1999_book. Of particular relevance in this strand of theoretical work is Chen-White1999_IEEE, where it was shown that single-layer, sigmoid-based networks could attain sufficiently fast rates for semiparametric inference (see Chen2007_handbook for more references).
We explicitly depart from the extant literature by focusing on the modern setting of deep neural networks with the rectified linear (ReLU) activation function. We provide nonasymptotic bounds for nonparametric estimation using deep neural networks, immediately implying convergence rates. The bounds and convergence rates appear to be new to the literature and are one of the main theoretical contributions of the paper. We provide results for a general class of smooth loss functions for nonparametric regression style problems, covering as special cases generalized linear models and other empirically useful contexts. In our application to causal inference we specialize our results to linear and logistic regression as concrete illustrations. Our proof strategy employs a localization analysis that uses scale-insensitive measures of complexity, allowing us to consider richer classes of neural networks. This is in contrast to analyses which restrict the networks to have bounded parameters for each unit (discussed more below) and to the application of scale sensitive measures such as metric entropy Chen-White1999_IEEE. These approaches would not deliver our sharp bounds and fast rates. Recent developments in approximation theory and complexity for deep ReLU networks are important building blocks for our results.
Our second main result establishes valid inference on finite-dimensional parameters following first-step estimation using deep learning. We focus on causal inference for concreteness and wide applicability, as well as to allow direct comparison to the literature. Program evaluation with observational data is one of the most common and important inference problems, and has often been used as a test case for theoretical study of inference following machine learning Belloni-Chernozhukov-Hansen2014_REStud,Farrell2015_arXiv,Belloni-etal2017_Ecma,Athey-Imbens-Wager2018_JRSSB. Causal inference as a whole is a vast literature; see Imbens-Rubin2015_book for a broad review and Abadie-Cattaneo2018_ARE for a recent review of program evaluation methods, and further references in both. Deep neural networks have been argued (experimentally) to outperform the previous state-of-the-art in causal inference westreich2010propensity,johansson2016learning,shalit2017estimating,hartford2017deep. To the best of our knowledge, ours are among the first theoretical results that explicitly deliver inference using deep neural networks.
We give specific results for average treatment effects, counterfactual expected utility/profits from treatment targeting strategies, and decomposition effects. Our results allow planners (e.g., firms or medical providers) to compare different strategies, either predetermined or estimated using auxiliary data, and recognizing that targeting can be costly, decide which strategy to implement. An interesting, and potentially useful, point we make in this context is that the selection on observables framework yields identification of counterfactual average outcomes without additional structural assumptions, so that, e.g., expected profit from a counterfactual treatment rule can be evaluated.
The usefulness of our deep learning results is of course not limited to causal inference. In particular, our results yield inference on essentially any estimand that admits a locally robust estimator Chernozhukov-etal2018_WP that depends only on target functions within our class of loss function (under appropriate regularity conditions). Our aim is not to innovate at the semiparametric step, for example by seeking weaker conditions on the first stage, but rather, we aim to utilize such results. Prior work has verified the high-level conditions for other first-stage estimators, such as traditional kernels or series/sieves, lasso methods, sigmoid-based shallow neural networks, and others (under suitable assumptions for each method). Our work contributes directly to this area of research by showing that deep nets are a valid and useful first-step estimator, in particular, attaining a rate of $o(n^{-1/4})$ under appropriate smoothness conditions. Finally, we do not rely on sample splitting or cross fitting. In particular, we use localization explicitly to directly verify conditions required for valid inference, which may be a novel application of this proof method that is useful in future semiparametric inference problems.
We numerically illustrate our results, and more generally the utility of deep learning, with a detailed simulation study and an empirical study of a direct mail marketing campaign. Our data come from a large US consumer products retailer and consists around to three hundred thousand consumers with one hundred fifty covariates. Hitsch-Misra2018_WP recently used this data to study various estimators, both traditional and modern, of heterogeneous treatment effects. We refer the reader to that paper for a more complete description of the data as well as results using other estimators (see also Hansen-Kozbur-Misra2017_WP). We study the effect of catalog mailings on consumer purchases, and moreover, compare different targeting strategies (i.e.\ to which consumers catalogs should be mailed). The cost of sending out a single catalog can be close to one dollar, and with millions being set out, carefully assessing the targeting strategy is crucial. Our results suggest that deep nets are at least as good as (and sometimes better than) the best methods in Hitsch-Misra2018_WP.
The remainder of the paper proceeds as follows. Next, we briefly review the related theoretical literature. Section (ref) introduces deep ReLU networks and states our main theoretical results: nonasymptotic bounds and convergence rates for general nonparametric regression-type loss functions. The semiparametric inference problem is set up in Section (ref) and asymptotic results are presented in Section (ref). The empirical application is presented in Section (ref). Results of a simulation study are reported in Section (ref). Section (ref) concludes. All proofs are given in the appendix.
Our paper contributes to several rapidly growing literatures, and we can not hope to do justice to each here. We give only those citations of particular relevance; more references can be found within these works. First, there has been much recent study of the statistical properties of the machine learning tools as an end in itself. Many studies have focused on the lasso and its variants Bickel-Ritov-Tsybakov2009_AoS,Belloni-Chernozhukov-Wang2011_Bmka,BCCH2012_Ecma,Farrell2015_arXiv and tree/forest based methods Wager-Athey2018_JASA, though earlier work studied shallow (typically with a single hidden layer) neural networks with smooth activation functions white1989learning, white1992artificial, Chen-White1999_IEEE. We fill the gap in this literature by studying deep neural networks with the non-smooth ReLU activation.
A second, intertwined strand of literature focuses on inference following the use of machine learning methods, often with a focus on average causal effects. Initial theoretical results were concerned with obtaining valid inference on a coefficient in a high-dimensional regression, following model selection or regularization, with particular focus on the lasso BCCH2012_Ecma,Javanmard-Montanari2014_JMLR,vandeGeer-etal2014_AoS. Intuitively, this is a semiparametric problem, where the coefficient of interest is estimable at the parametric rate and the remaining coefficients are collectively a nonparametric nuisance parameter estimated using machine learning methods. Building on this intuition, many have studied the semiparametric stage directly, such as obtaining novel, weaker conditions easing the application of machine learning methods Belloni-Chernozhukov-Hansen2014_REStud,Farrell2015_arXiv,Chernozhukov-etal2018_WP,Belloni-etal2018_handbook. Conceptually related to this strand are targeted maximum likelihood vanderLaan-Rose2011_book and the higher-order influence functions Robins-etal2008_IMS,Robins-etal2017_AoS. Our work builds on this work, employing conditions therein, and in particular, verifying them for deep ReLU nets.
Finally, our convergence rates build on, and contribute to, the recent theoretical machine learning literature on deep neural networks. Because of the renaissance in deep learning, a considerable amount of study has been done in recent years. Of particular relevance to us are Yarotsky2017_NN, Yarotsky2018_WP and Bartlett-etal2017_COLT; a recent textbook treatment, containing numerous other references, is given by Goodfellow-Bengio-Courville2016_book.
In this section we will give our main theoretical results: nonasymptotic bounds and associated convergence rates for deep neural network estimation. The utility of these results for second-step semiparametric causal inference (the downstream task), for which our rates are sufficiently rapid, is demonstrated in Section (ref). We view our results as an initial step in establishing both the estimation and inference theory for modern deep learning, i.e.\ neural networks built using the multi-layer perceptron architecture (described below) and the nonsmooth ReLU activation function. This combination is crucial: it has demonstrated state of the art performance empirically and can be feasibly optimized. This is in contrast with sigmoid-based networks, either shallow (for which theory exists, but may not match empirical performance) or deep (which are not feasible to optimize), and with shallow ReLU networks, which are not known to approximate broad classes functions.
As neural networks are perhaps less familiar to economists and other social scientists, we first briefly review the construction of deep ReLU nets. Our main focus will be on the fully connected feedfoward neural network, frequently referred to as a multi-layer perceptron, as this is the most commonly implemented network architecture and we want our results to inform empirical practice. However, our results are more general, accommodating other architectures provided they are able to yield a universal approximation (in the appropriate function class), and so we review neural nets more generally and give concrete examples.
Our goal is to estimate an unknown, assumed-smooth function $f_*(\bm{x})$, that relates the covariates $\bm{X} \in \mathbb{R}^d$ to an outcome $Y$ as the minimizer of the expectation of the per-observation loss function. Collecting these random variables into the vector $\bm{Z} = (Y, \bm{X}')' \in \mathbb{R}^{d+1}$, with $\bm{z} = (y, \bm{x}')'$ denoting a realization, we write \[ f_* = \operatornamewithlimits{arg\,min} \mathbb{E} \left[ \ell\left(f, \bm{Z}\right) \right]. \] We allow for any loss function that is Lipschitz in $f$ and obeys a curvature condition around $f_*$. Specifically, for constants $c_1$, $c_2$, and $C_\ell$ that are bounded and bounded away from zero, we assume that $\ell(f,\bm{z})$ obeys
Our results will be stated for a general loss obeying these two conditions.\footnote{We thank an anonymous referee for suggesting this approach.} We give a unified localization analysis of all such problems. This family of loss function covers many interesting problems. Two leading examples, used in our application to causal inference, are least squares and logistic regression, corresponding to the outcome and propensity score models respectively. For least squares, the target function and loss are
respectively, while for logistic regression these are
Lemma (ref) verifies, with explicit constants, that (ref) holds for these two. Losses obeying (ref) extend beyond these cases to other generalized linear models, such as count models, and can even cover multinomial logistic regression (multiclass classification), as shown in Lemma (ref).
For any loss, we estimate the target function using a deep ReLU network. We will give a brief outline of their construction here, paying closer attention to the details germane to our theory; complete introductions, and further references, are given by Anthony-Bartlett1999_book and Goodfellow-Bengio-Courville2016_book.
The crucial choice is the specific network architecture, or class. In general we will call this $\cF_{\rm DNN}$. From a theoretical point of view, different classes have different complexity and different approximating power. We give results for several concrete examples below. We will focus on feedforward neural networks. An example of a feedforward network is shown in Figure (ref). The network consists of $d$ input units, corresponding to the covariates $\bm{X} \in \mathbb{R}^d$, one output unit for the outcome $Y$. Between these are $U$ hidden units, or computational nodes or neurons. These are connected by a directed acyclic graph specifying the architecture. The key graphical feature of a feedforward network is that hidden units are grouped in a sequence of $L$ layers, the depth of the network, where a node is in layer $l = 1, 2, \ldots, L$, if it has a predecessor in layer $l-1$ and no predecessor in any layer $l' \geq l$. The width of the network at a given layer, denoted $H_l$, is the number of units in that layer. The network is completed with the choice of an activation function $\sigma: \mathbb{R} \mapsto \mathbb{R}$ applied to the output of each node as described below. In this paper, we focus on the popular ReLU activation function $\sigma(x) = \max(x, 0)$, though our results can be extended (at notational cost) to cover piecewise linear activation functions (see also Remark (ref)).
An important and widely used subclass is the one that is fully connected between consecutive layers but has no other connections and each layer has number of hidden units that are of the same order of magnitude. This architecture is often referred to as a Multi-Layer Perceptron (MLP) and we denote this class as $\cF_{\rm MLP}$. See Figure (ref), cf.\ Figure (ref). We will assume that all the width of all layers share a common asymptotic order $H$, implying that for this class $U \asymp L H$.
We will allow for generic feedforward networks in our results, but we present special results for the MLP case, as it is widely used in empirical practice. As we will see below, the architecture, through its complexity, and more importantly, approximation power, plays a crucial role in the final convergence rate. In particular, we find only a suboptimal rate for the MLP case, but our upper bound is still sufficient for semiparametric inference. As a note on exposition, while our main results are in fact nonasymptotic bounds that hold with high probability, for simplicity we will refer to them as “rates” in most discussion.
To build intuition on the computation, and compare to other nonparametric methods, let us focus on least squares for the moment, i.e.\ Equation (ref), with a continuous outcome using a multilayer perceptron with constant width $H$. Each hidden unit $u$ receives an input in the form of a linear combination $\tilde{\bm{x}}'\bm{w} + b$, and then returns $\sigma(\tilde{\bm{x}}'\bm{w} + b)$, where the vector $\tilde{\bm{x}}$ collects the output of all the units with a directed edge into $u$ (i.e., from prior layers), $\bm{w}$ is a vector of weights, and $b$ is a constant term. (The constant term is often referred to as the “bias” in the deep learning literature, but given the loaded meaning of this term in inference, we will largely avoid referring to $b$ as a bias.) The final layer's output is simply $\tilde{\bm{x}}'\bm{w} + b$ in the least squares case. The collection, over all nodes, of $\bm{w}$ and $b$, constitutes the parameters $\theta$ which are optimized in the final estimation. We denote $W$ as the total number of parameters of the network. For the MLP, $W = (d+1) H + (L-1) (H^2 + H) + H + 1$. In general, $W$, $U$, $L$, and $H$, may change with $n$, but we suppress this in the notation.
Optimization proceeds layer-by-layer using (variants of) stochastic gradient descent, with gradients of the parameters calculated by back-propagation (implementing the chain rule) induced by the network structure. To see this, let $\tilde{x}_{h,l}$ denote the scalar output of a node $u = (h,l)$, for $h = 1, \ldots H$, $l = 1, \ldots L$, and let $\tilde{\bm{x}}_{l} = (\tilde{x}_{1,l}, \ldots, \tilde{x}_{H,l})'$ for layer $l \leq L$. Each node thus computes $\tilde{x}_{h,l} = \sigma(\tilde{\bm{x}}_{l-1}'\bm{w}_{h,l-1} + b_{h,l-1})$ and the final output is $\hat{y} = \tilde{\bm{x}}_{L}'\bm{w}_{L} + b_{L}$. Once we recall that the network begins with the original observation $\bm{x}$, we can view $\tilde{\bm{x}}_{L} = \tilde{\bm{x}}_{L}(\bm{x})$, and thus the final output may be seen as a basis function approximation (albeit a complex and random one) written as $\hat{f}_{\rm MLP}(\bm{x}) = \tilde{\bm{x}}_{L}(\bm{x})'\bm{w}_{L} + b_{L}$, which is reminiscent of a traditional series (linear sieve) estimator. If all layers save the last were fixed, we could simply optimize using least squares directly: $(\bm{w}_{L}, b_{L}) = \operatornamewithlimits{arg\,min}_{\bm{w}, b} \| y_i - \tilde{\bm{x}}_{L}'\bm{w} - b\|_n^2$.
The crucial distinction is that the basis functions $\tilde{\bm{x}}_{L}(\cdot)$ are learned from the data. The “basis” is $\tilde{\bm{x}}_{L} = (\tilde{x}_{1,L}, \ldots, \tilde{x}_{H,L})'$, where each $\tilde{x}_{h,L} = \sigma(\tilde{\bm{x}}_{L-1}'\bm{w}_{h,L-1} + b_{h,L-1})$. Therefore, “before” we can solve the least squares problem above, we would have to estimate $(\bm{w}_{h,L-1}', b_{h,L-1}), h = 1, \ldots, H$, anticipating the final estimation. These in turn depend on the prior layer, and so forth back to the original inputs $\bm{X}$. Measuring the gradient of the loss with respect to each layer of parameters uses the chain rule recursively, and is implemented by back-propagation. This is simply a sketch of course; for further introduction, see Hastie-Tibshirani-Friedman2009_book and Goodfellow-Bengio-Courville2016_book.
To further clarify the use of deep nets, it is useful to make explicit analogies to more classical nonparametric techniques, leveraging the form $\hat{f}_{\rm MLP}(\bm{x}) = \tilde{\bm{x}}_{L}(\bm{x})'\bm{w}_{L} + b_{L}$. For a traditional series estimator, say smoothing splines, the two choices for the practitioner are the spline basis (the shape and the degree) and the number of terms (knots), commonly referred to as the smoothing and tuning parameters, respectively. In kernel regression, these would respectively be the shape of the kernel (and degree of local polynomial) and the bandwidth(s). For neural networks, the same phenomena are present: the architecture as a whole (the graph structure and activation function) are the smoothing parameters while the width and depth play the role of tuning parameters for a set architecture.
The architecture plays a crucial role in that it determines the approximation power of the network, and it is worth noting that because of the relative complexity of neural networks, such approximations, and comparisons across architectures, are not simple. It is comparatively obvious that quartic splines are more flexible than cubic splines (for the same number of knots) as is a higher degree local polynomial (for the same bandwidth). At a glance, it may not be clear what function class a given network architecture (width, depth, graph structure, and activation function) can approximate. As we will show below, the MLP architecture is not yet known to yield an optimal approximation (for a given width and depth) and therefore we are only able to prove a bound with slower than optimal rate. As a final note, computational considerations are important for deep nets in a way that is not true conventionally; see Remarks (ref), (ref), and (ref).
Just as for classical nonparametrics, for a fixed architecture, it is the tuning parameter choices that determine the rate of convergence (for a fixed smoothness of the underlying function). The recent wave of theoretical study of deep learning is still in its infancy. As such, there is no understanding yet of optimal architecture(s) or tuning parameters. Choices of both are quite difficult, and only preliminary research has been done daniely2017depth, telgarsky2016benefits, safran2016depth, Mhaskar-Poggio2016_WP,Raghu-etal2017_ICML. Further exploration of these ideas is beyond the current scope. It is interesting to note that in some cases, a good approximation can be obtained even with a fixed width $H$, provided the network is deep enough, a very particular way of enriching the “sieve space” $\cF_{\rm DNN}$; see Corollary (ref).
In sum, for a user-chosen architecture $\cF_{\rm DNN}$, encompassing the choices $\sigma(\cdot)$, $U$, $L$, $W$, and the graph structure, the final estimate is computed using observed samples $\bm{z}_i = (y_i, \bm{x}_i')'$, $i=1,2, \ldots, n$, of $\bm{Z}$, by solving
Recall that $\theta$ collects, over all nodes, the weights and constants $\bm{w}$ and $b$. When (ref) is restricted to the MLP class we denote the resulting estimator $\widehat{f}_{\rm MLP}$. The choice of $M$ may be arbitrarily large, and is part of the definition of the class $\cF_{\rm DNN}$. This is neither a tuning parameter nor regularization in the usual sense: it is not assumed to vary with $n$, and beyond being finite and bounding $\|f_*\|_\infty$ (see Assumption (ref)), no properties of $M$ are required. This is simply a formalization of the requirement that the optimizer is not allowed to diverge on the function level in the $l_\infty$ sense-- the weakest form of constraint. It is important to note that while typically regularization will alter the approximation power of the class, that is not the case with the choice of $M$ as we will assume that the true function $f_*(\bm{x})$ is bounded, as is standard in nonparametric analysis. With some extra notational burden, one can make the dependence of the bound on $M$ explicit, though we omit this for clarity as it is not related to statistical issues.
We can now state our main theoretical results: bounds and convergence rates for deep ReLU networks. All proofs appear in the Appendix. We study neural networks from a nonparametric point of view white1989learning, white1992artificial,schmidt2017nonparametric,Liang2018_GAN,bauer2017deep. Chen-Shen1998_Ecma and Chen-White1999_IEEE share our goal, fast convergence rates for use in semiparametric inference, but focus on shallow, sigmoid-based networks compared to our deep, ReLU-based networks, though they consider dependent data which we do not. Our theoretical approach is quite different. In particular, Chen-White1999_IEEE obtain sufficiently fast rates by following the approach of Barron1993_IEEE in using Maurey's method pisier1981remarques for approximation, but applying the refinement of makovoz1996random. Our analysis of deep nets instead employs localization methods koltchinskii2000rademacher, bartlett2005local, koltchinskii2006local, Koltchinskii2011_book, liang2015learning, along with the recent approximation work of Yarotsky2017_NN, Yarotsky2018_WP and complexity results of Bartlett-etal2017_COLT.
The regularity conditions we require are collected in the following.
This assumption is fairly standard in nonparametrics. The only restriction worth mentioning is that the outcome is bounded. In many cases this holds by default (such as logistic regression, where $\mathcal{Y} = \{0,1\}$) or count models (where $\mathcal{Y} = \{0,1,\ldots,M\}$, with $M$ limited by real-world constraints). For continuous outcomes, such as least squares regression, our restriction is not substantially more limiting than the usual assumption of a model such as $Y = f_*(\bm{X}) + \varepsilon$, where $\bm{X}$ is compact-supported, $f_*$ is bounded, and the stochastic error $\varepsilon$ possesses many moments. Indeed, in many applications such a structure is only coherent with bounded outcomes, such as the common practice of including lagged outcomes as predictors. Next, the assumption of continuously distributed covariates is quite standard. From a theoretical point of view, covariates taking on only a few values can be conditioned on and then averaged over, and these will, as usual, not enter into the dimensionality which curses the rates. Discrete covariates taking on many values may be more realistically thought of as continuous, and it may be more accurate to allow these to slow the convergence rates. Our focus on $L_2(\bm{X})$ convergence allows for these essentially automatically. Finally, from a practical point of view, deep networks handle discrete covariates seamlessly and have demonstrated excellent empirical performance, which is in contrast to other more classical nonparametric techniques that may require manual adaptation.
Proceeding now to our results, we begin with the most important network architecture, the multi-layer perceptron. This is the most widely used network architecture in practice and an important contribution of our work is to cover this directly, along with ReLU activation. MLPs are now known to approximate smooth functions well, leading to our next assumption: that the target function $f_*$ lies in a Sobolev ball with certain smoothness. Discussion of Sobolev spaces, and comparisons to H\"older and Besov spaces, can be found in Gine-Nickl2016_book.
Under Assumptions (ref) and (ref) we obtain the following result, which, to the best of our knowledge, is new to the literature. In some sense, this is our main result for deep learning, as it deals with the most common architecture. We apply this in Sections (ref) and (ref) for semiparametric inference.
Several aspects of this result warrant discussion. We build on the recent results of Bartlett-etal2017_COLT, who find nearly-tight bounds on the Vapnik-Chervonenkis (VC) and Pseudo-dimension of deep nets. One contribution of our proof is to use a scale sensitive localization theory with scale insensitive measures, such as VC- or Pseudo-dimension, for deep neural networks for general smooth loss functions. For the special case of least squares regression, Koltchinskii2011_book uses a similar approach, and a similar result to our Theorem (ref)(a) can be derived for this case using his Theorem 5.2 and Example 3 (p.\ 85f).
This approach has two tangible benefits. First, we do not restrict the class of network architectures to have bounded weights for each unit (scale insensitive), in accordance to standard practice zhang2016understanding and in contrast to the classic sieve analysis with scale sensitive measure such as metric entropy. Moreover, this allows for a richer set of approximating possibilities, in particular allowing more flexibility in seeking architectures with specific properties, as we explore in the next subsection. Second, from a technical point of view, we are able to attain a faster rate on the second term of the bound, order $n^{-1}$ in the sample size, instead of the $n^{-1/2}$ that would result from a direct application of uniform deviation bounds. This upper bound informs the trade offs between width and depth, and the approximation power, and may point toward optimal architectures for statistical inference.
This result gives a nonasymptotic bound that holds with high probability. As mentioned above, we will generally refer to our results simply as “rates” when this causes no confusion. This result relies on choosing $H$ appropriately given the smoothness $\beta$ of Assumption (ref). Of course, the true smoothness is unknown and thus in practice the “$\beta$” appearing in $H$, and consequently in the convergence rates, need not match that of Assumption (ref). In general, the rate will depend on the smaller of the two. Most commonly it is assumed that the user-chosen $\beta$ is fixed and that the truth is smoother; witness the ubiquity of cubic splines and local linear regression. Rather than spell out these consequences directly, we will tacitly assume the true smoothness is not less than the $\beta$ appearing in $H$ (here and below). Adaptive approaches, as in classical nonparametrics, may also be possible with deep nets, but are beyond the scope of this study.
Even with these choices of $H$ and $L$, the bound of Theorem (ref) is not optimal (for fixed $\beta$, in the sense of stone1982optimal). We rely on the explicit approximating constructions of Yarotsky2017_NN, and it is possible that in the future improved approximation properties of MLPs will be found, allowing for a sharpening of the results of Theorem (ref) immediately, i.e.\ without change to our theoretical argument. At present, it is not clear if this rate can be improved, but it is sufficiently fast for valid inference.
Theorem (ref) covers only one specific architecture, albeit the most important one at present. However, given that this field is rapidly evolving, it is important to consider other possible architectures which may be beneficial in some cases. To this end, we will state a more generic result and then two specific examples: one to obtain a faster rate of convergence and one for fixed-width networks. All of these results are, at present, more of theoretical interest than practical value, as they are either agnostic about the network (thus infeasible) or rely on more limiting assumptions.
In order to be agnostic about the specific architecture of the network we need to be flexible in the approximation power of the class. To this end, we will replace Assumption (ref) with the following generic assumption, rather more of a definition, regarding the approximation power of the network.
It may be possible to require only an approximation in the $L_2(\bm{X})$ norm, but this assumption matches the current approximation theory literature and is more comparable with other work in nonparametrics, and thus we maintain the uniform definition.
Under this condition we obtain the following generic result.
This is a more general than Theorem (ref), covering the general deep ReLU network problem defined in (ref), general feedforward architectures, and the general class of losses defined by (ref). The same comments as were made following Theorem (ref) apply here as well: the same localization argument is used with the same benefits. We explicitly use this in the next two corollaries, where we exploit the allowed flexibility in controlling $\epsilon_{\rm DNN}$ by stating results for particular architectures. The bound here is not directly applicable without specifying the network structure, which will determine both the variance portion (through $W$, $L$, and $U$) and the approximation error. With these set, the bound becomes operational upon choosing $\gamma$, which can be optimized as desired, and this will immediately then yield a convergence rate.
Turning to special cases, we first show that the optimal rate of stone1982optimal can be attained, up to log factors. However, this relies on a rather artificial network structure, designated to approximate functions in a Sobolev space well, but without concern for practical implementation. Thus, while the following rate improves upon Theorem (ref), we view this result as mainly of theoretical interest: establishing that (certain) deep ReLU networks are able to attain the optimal rate.
Next, we turn to very deep networks that are very narrow, which have attracted substantial recent interest. Theorem (ref) and Corollary (ref) dealt with networks where the depth and the width grow with sample size. This matches the most common empirical practice, and is what we use in Sections (ref) and (ref). However, it is possible to allow for networks of fixed width, provided the depth is sufficiently large. The next result is perhaps the largest departure from the classical study of neural networks: earlier work considered networks with diverging width but fixed depth (often a single layer), while the reverse is true here. The activation function is of course qualitatively different as well, being piecewise linear instead of smooth. Using recent results mhaskar2016deep,Hanin2017_WP,Yarotsky2018_WP we can establish the following rate for very deep, fixed-width MLPs.
This result is again mainly of theoretical interest. The class is only able to approximate well functions with $\beta = 1$ (cf. the choice of $L$) which limits the potential applications of the result because, in practice, $d$ will be large enough to render this rate, unlike those above, too slow for use in later inference procedures. In particular, if $d \geq 3$, the sufficient conditions of Theorem (ref) fail.
Finally, as mentioned following Theorem (ref), our theory here will immediately yield a faster rate upon discovery of improved approximation power of this class of networks. In other words, for example, if a proof became available that fixed-width, very deep networks can approximate $\beta$-smooth functions (as in Assumption (ref)), then Corollary (ref) will trivially be improvable to match the rate of Theorem (ref). Similarly, if the MLP architecture can be shown to share the approximation power with that of Corollary (ref), then Theorem (ref) will itself deliver the optimal rate. Our proofs will not require adjustment.
We will use the results above, in particular Theorem (ref), coupled with results in the semiparametric literature, to deliver valid asymptotic inference for causal effects. The novelty of our results is not in this semiparametric stage per se, but rather in delivering valid inference after relying on deep learning for the first step estimation. In this section we define the parameters of interest, while asymptotic inference is discussed next.
We will focus, for concreteness, on causal parameters that are of interest across different disciplines: average treatment effects, expected utility (or profits) under different targeting policies, average effects on (non-)treated subpopulations, and decomposition effects. Our focus on causal inference with observational data is due to the popularity of these estimands both in applications and in theoretical work, thus allowing our results to be put to immediate use and easily compared to prior literature. The average treatment effect in particular is often used as a benchmark parameter for studying inference following machine learning (see references in the Introduction). However, armed with our results for deep neural networks we can cover a great deal more (some discussion is in Section (ref)).
The estimation of average causal effects is a well-studied problem, and we will give only a brief overview here. Recent reviews and further references are given by Belloni-etal2017_Ecma,Athey-Imbens-Pham-Wager2017_AERPP,Abadie-Cattaneo2018_ARE. We consider the standard setup for program evaluation with observational data: we observe a sample of $n$ units, each exposed to a binary treatment, and for each unit we observe a vector of pre-treatment covariates, $\bm{X} \in \mathbb{R}^d$, treatment status $T \in \{0,1\}$, and a scalar post-treatment outcome $Y$. The observed outcome obeys $Y = T Y(1) + (1-T)Y(0)$, where $Y(t)$ is the (potential) outcome under treatment status $t \in \{0,1\}$. The “fundamental problem” is that only $Y(0)$ or $Y(1)$ is observed for each unit, never both.
The crucial identification assumptions, which pertain to all the parameters we consider, are selection on observables, also known as ignorability, unconfoundedness, missingness at random, or conditional independence, and overlap, or common support. Let $p(\bm{x}) = \mathbb{P}[T = 1 \vert \bm{X} = \bm{x}]$ denote the propensity score and $\mu_t(\bm{x}) = E[Y(t) \vert \bm{X}=\bm{x}], \ t \in \{0,1\}$ denote the two outcome regression functions. We then assume the following throughout, beyond which, we will mostly need only regularity conditions for inference.
It will be useful to divide our discussion between parameters that are fully marginal averages, such as the average treatment effect, and those which are for specific subpopulations. Here, “subpopulations” refer to the treated or nontreated groups, with corresponding parameters such as the treatment effect for the treated. Any parameter, in either case, can be studied for a suitable subpopulation defined by the covariates $\bm{X}$, such as a specific demographic group. Though causal effects as a whole share some structure, there are slight conceptual and notational differences. In particular, the form of the efficient influence function and doubly robust estimator is different for the two sets, but common within.
Here we are interested in averages over the entire population. The prototypical parameter of interest is the average treatment effect:
In the context of our empirical example, the treatment is being mailed a catalog and the outcome is dollars spent (results for the binary purchase decision are available on request). The average treatment effect, also referred to as “lift” in digital contexts, corresponds to the expected gain in revenue from an average individual receiving the catalog compared to the same person not receiving the catalog.
A closely related parameter of interest is the average realized outcome, which in general may be interpreted as the expected utility or welfare from a counterfactual treatment policy. In the context of our empirical application this is expected profits; in a medical context it would be the total health outcome. The question of interest here is whether a change in the treatment policy would be beneficial in terms of increasing outcomes, and this is judged using observational data. Intuitively, the average treatment effect is the expected gain from treating the “next” person, relative to if they had not been exposed. That is, it is the expected change in the outcome. Expected utility/profit, on the other hand, is concerned with the total outcome, not the difference in outcomes. In the context of our empirical application, we are interested in total sales rather than the change in sales. Our discussion is grounded in this language for easy comparison.
The parameter depends on a counterfactual/hypothetical treatment targeting strategy, which is often itself the object of evaluation. This is simply a rule that assigns a given set of characteristics (e.g.\ a consumer profile), determined by the covariates $X$, to treatment status: that is, a known function (which may include randomization but is not estimated from the sample) $s(\bm{x}): \operatorname*{supp}\{\bm{X}\} \mapsto \{0,1\}$. Note well that this is not necessarily the observed treatment: $s(\bm{x}_i) \neq t_i$. The policy maker may wish to evaluate the gain from targeting only a certain subset of customers, a price discrimination strategy, or comparisons of different such policies. Our assumptions, while standard, deliver identification of such counterfactuals at no cost.
The parameter of interest is expected utility, or profit, from a fixed policy, given by
where we make explicit the dependence on the policy $s(\cdot)$. Compare to Equation (ref) and recall that the observed outcome obeys $Y = T Y(1) + (1-T)Y(0)$. Whereas $\tau$ is the gain in assigning the next person to treatment and is given by the difference in potential outcomes, $\pi(s)$ is the expected outcome that would be observed for the next person if the treatment rule were $s(\bm{x})$.
A natural question is whether a candidate targeting strategy, say $s'(\bm{x})$, is superior to baseline or status quo policy, $s_0(\bm{x})$. This amounts to testing the hypothesis $H_0: \pi(s') \geq \pi(s_0)$. To evaluate this, we can study the difference in expected profits, which amounts to
Assumption (ref) provides identification for $\pi(s)$ and $\pi(s', s_0)$, arguing analogously as for $\tau$. Moreover, notice that $\pi(s', s_0) = \mathbb{E} [(s'(\bm{X}) - s_0(\bm{X})) (Y(1) - Y(0) )] = \mathbb{E} [(s'(\bm{X}) - s_0(\bm{X})) \tau(\bm{X})]$, where $\tau(\bm{x}) = \mathbb{E}[Y(1) - Y(0) \mid \bm{X} = \bm{x}]$ is the conditional average treatment effect. The latter form makes clear that only those differently treated, of course, impact the evaluation of $s'$ compared to $s_0$. The strategy $s'$ will be superior if, on average, it targets those with a higher individual treatment effect. Estimating the optimal treatment policy from the data is discussed briefly in Section (ref).
The common structure of these parameters is that they all involve full-population averages of the potential outcomes, possibly scaled by a known function. For these parameters, the influence function is known from Hahn1998_Ecma, and estimators based on the influence function are doubly robust, as they remain consistent if either the regression functions or the propensity score are correctly specified Robins-Rotnitzky-Zhao1994_JASA,Robins-Rotnitzky-Zhao1995_JASA. With a slight abuse of terminology (since we are omitting the centering), the influence function for a single average potential outcome, $t \in \{0,1\}$, is given by, for $\bm{z} = (y, t, \bm{x}')'$,
Our estimation of $\tau$, $\pi(s)$, and $\pi(s', s_0)$ will utilize sample averages of this function, with unknown objects replaced by estimators. Our use of influence functions here follows the recent literature in econometrics showing that the double robustness implies valid inference under weaker conditions on the first step nonparametric estimates Farrell2015_arXiv,Chernozhukov-etal2018_EJ.
The second type of causal effects of interest are based on potential outcomes averaged over only a specific treatment group. A single such average, for $t, t' \in \{0,1\}$, is denoted by
Many interesting parameters are linear combinations of these for different $t$ and $t'$. We focus on two for concreteness. (We could also consider averages restricted by targeting-type functions, as in expected utility/profit, but for brevity we omit this.) The most well-studied of these parameters is the treatment effect on the treated, given by
To appreciate the breadth of this framework, and the applicability of our causal inference results, we also consider a decomposition parameter, a semiparametric analogue of Oaxaca-Blinder Kitagawa1955_JASA,Oaxaca1973_IER,Blinder1973_JHR. In this context, the “treatment” variable $T$ is typically not a treatment assignment per se, but rather an exogenous covariate such as a demographic indicator, perhaps most commonly a male/female indicator. See Fortin-Lemieux-Firpo2011_handbook for a complete discussion and further references. The parameter of interest in this case is the decomposition of $\Delta = \mathbb{E}[Y(1) \mid T = 1] - \mathbb{E}[Y(0) \mid T = 0]$, into the difference in the covariate distributions and the difference in expected outcomes. These can be written as functions of different $\rho_{t,t'}$. For example, $\Delta_X = \mathbb{E}[Y(1) \vert T \!=\! 1] - \mathbb{E}[Y(1) \vert T \!=\! 0] = \mathbb{E}[\mu_1(\bm{X}) \vert T \!=\! 1] - \mathbb{E}[\mu_1(\bm{X}) \vert T \!=\! 0] = \rho_{1,1} - \rho_{1,0}$. We are in general interested in
Just as in the case of full-population averages, the influence function is known and leads to a doubly robust estimator. For a single $\rho_{t,t'}$, the (uncentered) influence function is (cf.\ (ref)):
Estimation and inference requires, as above, estimation of the propensity scores and regression functions, depending on the exact choices of $t$ and $t'$, and here we also require the marginal probability of treatments.
Moving beyond a fixed parameter, our results on deep neural networks can be used to address optimal targeting. In the notation of Section (ref), this amounts to finding a policy, say $s_\star(\bm{x})$, that maximizes a given measure of utility stemming from treatment, generally the expected gain relative to a baseline policy. In Section (ref) we considered the utility (or profit) difference between two given strategies, a candidate $s'(\bm{x})$ and a baseline $s_0(\bm{x})$. Instead of inference on $\pi(s', s_0)$, we can use the data to find the $s_\star(\bm{x})$ which maximizes the gain relative to the baseline. This problem has been widely studied in econometrics and statistics; for detailed discussion and numerous references see Manski2004_Ecma, Hirano-Porter2009_Ecma, Kitagawa-Tetenov2018_Ecma, and Athey-Wager2018_WP. In particular, the latter noticed using the locally robust framework allows policy optimization under nearly the same conditions as inference and proved fast convergence rates of the estimated policy in terms of regret.
More formally, we want to find the optimal choice $s_\star(\bm{x})$ in some policy/action space $\mathcal{S}$. The policy space, and thus its complexity, is user determined. Simple examples include simple decision trees or univariate-based strategies; more can be found in the references above. Recall that $\pi(s', s_0) = \mathbb{E}[Y(s')] - \mathbb{E}[Y(s_0)] = \mathbb{E} [(s'(\bm{X}) - s_0(\bm{X})) \tau(\bm{X})]$, where $\tau(\bm{x}) = \mathbb{E}[Y(1) - Y(0) \mid \bm{X} = \bm{x}]$ is the conditional average treatment effect. Given a space $\mathcal{S}$, we wish to find the policy $s_\star(\bm{x}) \in \mathcal{S}$ which solves $\max_{s' \in \mathcal{S}} \pi(s', s_0)$. The main result of Athey-Wager2018_WP is that replacing $\pi$ with the doubly-robust $\hat{\pi}$ of Equation (ref), and minimizing the empirical analogue of regret, one obtains an estimator $\hat{s}(\bm{x})$ of the optimal policy that obeys the regret bound $\pi(s_\star, s_0) - \pi(\hat{s}, s_0) = O_P(\sqrt{{\rm VC}(\mathcal{S})/n})$ (a formal statement would be notationally burdensome). The complexity of the user-chosen policy space enters the bound through its VC dimension. Simple, interpretable policy classes often have bounded or slowly-growing dimension, implying rapid convergence.
There are of course many other contexts where first-step deep learning is useful. Only trivial extensions to the above would be required for other causal effects, such as multi-valued treatments Cattaneo2010_JoE and others with doubly-robust estimators Sloczynski-Wooldridge2018_ET. Further, under selection on observables, treatment effects, missing data, measurement error, and data combination are equivalent, and thus all our results apply immediately to those contexts. For reviews of these and Assumption (ref) more broadly, see Chen-Hong-Tarozzi2004_WP,Tsiatis2006_book,Heckman-Vytlacil2007a_Handbook,Imbens-Wooldridge2009_JEL.
Moving beyond causal effects, any estimand with a locally/doubly robust estimator depending only on target functions falling into our class of losses can be covered using the results of Section (ref). For example, estimands requiring distribution estimation require further study; see Liang2018_GAN for recent results via Generative Adversarial Networks (GANs). More precisely, along with regularity conditions, our theory can be used to verify the conditions of Chernozhukov-etal2018_WP, who treat more general semiparametric estimands using local robustness, sometimes relying on sample splitting or cross fitting. Further in this vein, our results on deep neural networks can be used to address optimal targeting, i.e., finding the policy, say $s_\star(\bm{x})$, that maximizes a given measure of utility, by applying the results of Athey-Wager2018_WP, who noticed that using the locally robust framework allows policy optimization.
More broadly, the learning of features using deep neural networks is becoming increasingly popular and our results speak to this context directly. To illustrate, consider the simple example of a linear model where some predictors are features learned from independent data. Here, the object of interest is the fixed-dimension coefficient vector $\bm{\lambda}$, which we assume can be partitioned as $\bm{\lambda} = (\bm{\lambda}_1', \bm{\lambda}_2')'$ according to the model $Y = \bm{f}(\bm{X})'\bm{\lambda}_1 + W'\bm{\lambda}_2 + \varepsilon$. The features $\bm{f}(\bm{X})$, often a “score” of some type, are generally learned from auxiliary (and independent) data. For a recent example, see liu2017large. In such cases, inference on $\bm{\lambda}$ can proceed directly, as long as care is taken to interpret the results. See Section (ref).
We now turn to asymptotic inference for the causal parameters discussed above. We first define the estimators, which are based on sample averages of the (uncentered) influence functions (ref) and (ref). We then give a generic result for single averages which can then be combined for inference on a given parameter of interest. Below we discuss inference under randomized treatment and using sample splitting.
Throughout, we assume we have a sample $\{\bm{z}_i = (y_i, t_i, \bm{x}_i')'\}_{i=1}^n$ from $\bm{Z} = (Y, T, \bm{X}')'$. We then form
where $\hat{\mathbb{P}}[T = t \mid \bm{X} = \bm{x}_i] = \hat{p}(\bm{x}_i)$ for $t=1$ and $1-\hat{p}(\bm{x}_i)$ for $t=0$, and similarly
where $\hat{\mathbb{P}}[T = t']$ is simply the sample frequency $\mathbb{E}_n[\mathbbm{1}\{t_i = t'\}]$.
For the first stage estimates appearing in (ref) and (ref) we use our results on deep nets, and Theorem (ref) in particular. Specifically, the estimated propensity score, $\hat{p}(\bm{x})$, is the estimate that results from solving (ref), with the MLP architecture, for the logistic loss (ref) with $T$ as the outcome. Similarly, for each status $t \in \{0,1\}$, we can let $\hat{\mu}_t(\bm{x})$ be the deep-MLP estimate of $f_*(\bm{x}) = \mathbb{E}[ Y | T = t, \bm{X} = \bm{x}]$, solving (ref) for least squares loss, (ref), with outcome $Y$, using only observations with $t_i = t$. However, it is worth noting that the theoretically-equivalent joint estimation of Equation (ref) performs much better, as the two groups may share features. To state the results, let $\beta_p$ and $\beta_\mu$ be the smoothness parameters of Assumption (ref) for the propensity score and outcome models, respectively.
We then obtain inference using the following results, essentially taken from Farrell2015_arXiv. Similar results are given by Belloni-etal2017_Ecma and Chernozhukov-etal2018_EJ. All of these provide high-level conditions for valid inference, and none verify these for deep nets as we do here.
This result, our main inference contribution, shows exactly how deep learning delivers valid asymptotic inference for our parameters of interest. Theorem (ref) (a generic result using Theorem (ref) could be stated) proves that the nonparametric estimates converge sufficiently fast, as formalized by conditions (a), (b), and (c), enabling feasible efficient semiparametric inference. In general, these are implied by, but may be weaker than, the requirement of that the first step estimates converge faster than $n^{-1/4}$, which our results yield for deep ReLU nets. The first is a mild consistency requirement. The second requires a rate, but on the product of the two estimates, which can be satisfied under weaker conditions. Finally, the third condition is the strongest. Intuitively, this condition arises from a “leave-in” type remainder, and as such, it can be weakened using sample splitting Chernozhukov-etal2018_EJ,Newey-Robins2018_WP. We opt to maintain (c) exactly because deep nets are not amenable to either simple leave-one-out forms (as are, e.g., classical kernel regression) or to sample splitting, being a data hungry method the gain in theoretically weaker rate requirements may not be worth the price paid in constants in finite samples. Instead, we employ our localization analysis, as was used to obtain the results of Section (ref), to verify (c) directly (see Lemma (ref)); this appears to be a novel application of localization, and this approach may be useful in future applications of second-step inference using machine learning methods.
From this result we immediately obtain inference for all the causal parameters discussed above. For the full-population averages, for example, we would form
The estimator $\hat{\tau}$ is exactly the doubly/locally robust estimator of the average treatment effect that is standard in the literature. The estimators for profits can be thought of as the doubly robust version of the constructs described in Hitsch-Misra2018_WP. Furthermore, to add a per-unit cost of treatment/targeting $c$ and a margin $m$, simply replace $\psi_1$ with $m \psi_1 - c$ and $\psi_0$ with $m \psi_0$. Similarly, $\hat{\tau}_{1,0}$, $\hat{\Delta}_X$, and $\hat{\Delta}_\mu$ would be linear combinations of different $\hat{\rho}_{t,t'} = \mathbb{E}_n [ \hat{\psi}_{t,t'}(\bm{z}_i) ]$.
It is immediate from Theorem (ref) that all such estimators are asymptotically Normal. The asymptotic variance can be estimated by simply replacing the sample first moments of (ref) with second moments. That is, looking at $\hat{\pi}(s)$ to fix ideas, \[ \sqrt{n} \hat{\Sigma}^{-1/2}\left(\hat{\pi}(s) - \pi(s) \right) \stackrel{d}{\to} \mathcal{N}(0,1) , \quad \text{with}\quad \hat{\Sigma} = \mathbb{E}_n \left[ \left( s(\bm{x}_i)\hat{\psi}_1(\bm{z}_i) + (1-s(\bm{x}_i))\hat{\psi}_0(\bm{z}_i) \right)^2 \right] - \hat{\pi}(s)^2. \] The others are similar. Further, Theorem (ref) can be generalized straightforwardly to yield uniformly valid inference, following the approach of Romano2004_SJS, exactly as in Belloni-Chernozhukov-Hansen2014_REStud or Farrell2015_arXiv.
Finally, we note that our focus with Theorem (ref) is showcasing the practical utility of deep learning. Our use of local/double robustness here is toward the aim of attaining feasible inference without requiring more detailed assumptions on the machine learning step. This comes at the expense of, for example, stronger-than-minimal smoothness assumptions. That is, the requirement that $\beta_p \wedge \beta_\mu > d$ is not minimal, and moreover, neither is the weaker condition $\beta_p \wedge \beta_\mu > d/2$ that would be required after applying Corollary (ref) instead of Theorem (ref). Obtaining a Gaussian limit, and possibly semiparametric efficiency, under minimal conditions has been studied by many, dating at least to Bickel-Ritov1988_Sankhya; see Robins-etal2009_EJS for recent results and references on optimal estimation and minimal conditions. For causal inference, Chen-Hong-Tarozzi2008_AoS and Athey-Imbens-Wager2018_JRSSB obtain semiparametric efficiency under strictly weaker conditions than ours on $p(\bm{x})$ (the former under minimal smoothness on $\mu_t(\bm{x})$ and the latter under a sparsity in a high-dimensional linear model). Further, as above, cross-fitting Newey-Robins2018_WP paired with local robustness may yield weaker smoothness conditions by providing underfitting” robustness (i.e.\ weakening bias-related tuning parameter assumptions). On the other hand, weaker variance-related assumptions, or “overfitting” robust inference procedures, Cattaneo-Jansson2018_Ecma,Cattaneo-Jansson-Ma2018_RESTUD, may also be possible following deep learning, but are less automatic at present. Finally, other methods designed for causal inference under relaxed assumptions may be useful here, such as the recently developed extensions to doubly robust estimation tan2018model and inverse weighting Ma-Wang2018_IPW: pursuing these in the context of deep learning is left to future work.
Our analysis thus far has focused on observational data, but it is worth spelling out results for randomized experiments. This is particularly important in the Internet age, where experimentation is common, vast amounts of data are available, and effects are often small in magnitude Taddy-etal2015_WP. Indeed, our empirical illustration, detailed in the next section, stems from an experiment with 300,000 units and hundreds of covariates. When treatment is randomized, inference can be done directly using the mean outcomes in the treatment and control groups, such as the difference for the average treatment effect or the corresponding weighted sum for profit. However, pre-treatment covariates can be used to increase efficiency Hahn2004_REStat.
We will focus on the simple situation of a purely randomized binary treatment, but our results can be extended naturally to other randomization schemes. We formalize this with the following.
Under this assumption, the obvious simplification is that the propensity score need not be estimated using the covariates, but can be replaced with the (still nonparametric) sample frequency: $\hat{p}(\bm{x}_i) \equiv \hat{p} = \mathbb{E}_n [t_i]$. This is plugged into Equation (ref) and estimation and inference proceeds as above. Only rate conditions on the regression functions $\hat{\mu}_t(x)$ are needed. Further, conditions (a) and (b) of Theorem (ref) collapse, as $\hat{p}$ is root-$n$ consistent, leaving only condition (c) to be verified. Again, cross-fitting can be used in theory to remove this condition and thus weaken the requirement that $\beta_\mu > d$, but we maintain this for simplicity. We collect this into the following result, which is a trivial corollary of Theorem (ref).
Sample splitting may be used to obtain valid inference in cases, unlike those above, where the parameter of interest itself is learned from the data. For the causal estimands above, the regression functions and propensity score must be estimated, but these are nuisance functions. This is not true in the inference after policy or feature learning (Sections (ref) and (ref)). For policy learning, our results can be used to verify the high-level conditions of Athey-Wager2018_WP, though they require the additional condition of uniform consistency of the first stage estimators, and for machine learning estimators this is not clearly innocuous. However, this gives only point estimation.
Sample splitting is used in the obvious way: the first subsample, or more generally, independent auxiliary data, is used to learn the features or optimal policy, and then Theorem (ref) is applied in the second subsample, conditional on the results of the first. For policy learning this delivers valid inference on $\pi(\hat{s})$ or $\pi(\hat{s}, s_0)$, while for the simple example of feature learning in a linear model we obtain inference on the parameters defined by the “model” $Y = \bm{f}(\bm{X})'\bm{\lambda}_1 + W'\bm{\lambda}_2 + \varepsilon$, where $\bm{f}(\bm{X})$ is estimated from auxiliary data. Care must be taken in interpreting the results. The results of the first-subsample estimation are effectively conditioned upon in the inference stage, redefining the target parameter to be in terms of the learned object. In many contexts this may be sufficient Chernozhukov-etal2018_generic, but further assumptions will generally be needed to assume that the first subsample has recovered the true population object. To fix ideas, consider policy learning: inference on $\pi(\hat{s}, s_0)$, conditional on the map $\hat{s}(\bm{x})$ learned in the first subsample, is immediate and requires no additional assumptions, but inference on $\pi(s_\star, s_0)$ is not obvious without further conditions.
To illustrate our results, Theorems (ref) and (ref) in particular, we study, from a marketing point of view, a randomized experiment from a large US retailer of consumer products. The outcome of interest is consumer spending and the treatment is a catalog mailing. The firm sells directly to the customer (as opposed to via retailers) using a variety of channels such as the web and mail. The data consists of nearly three hundred thousand (292,657) consumers chosen at random from the retailer's database. Of these, 2/3 were randomly chosen to receive a catalog, and in addition to treatment status, we observe roughly one hundred fifty covariates, including demographics, past purchase behaviors, interactions with the firm, and other relevant information. For more on the data and a complete discussion of the decision making issues, we refer the reader to Hitsch-Misra2018_WP (we use the 2015 sample). That paper studied various estimators, both traditional and modern, of average and heterogeneous causal effects. Importantly, they did not consider neural networks. Our results show that deep nets are at least as good as (and sometimes better than) the best methods in Hitsch-Misra2018_WP.
In terms of motivation, a key element of a firm's toolkit is the design and implementation of targeted marketing instruments. These instruments, aiming to induce demand, often contain advertising and informational content about the firms offerings. The targeting aspect thus boils down to the selection of which particular customers should be sent the material. This is a particularly important decision since the costs of creation and dissemination of the material can accumulate rapidly, particularly over a large customer base. For a typical retailer engaging in direct marketing the costs of sending out a catalog can be close to a dollar per targeted customer. With millions of catalogs being sent out, the cost of a typical campaign is quite high.
Given these expenses, an important problem for firms is ascertaining the causal effects of such targeted mailing, and then using these effects to evaluate potential targeting strategies. At a high level, this approach is very similar to modern personalized medicine where treatments have to be targeted. In these contexts, both the treatment and the targeting can be costly, and thus careful assessment of $\pi(s)$ (interpreted as welfare) is crucial for decision making.
The outcome of interest for the firm is customer spending. This is the total amount of money that a given customer spends on purchases of the firm's products, within a specified time window. For the experiment in question the firm used a window of three months, and aggregated sales from all available purchase channels including phone, mail, and the web. In our data 6.2% of customers made a purchase. Overall mean spending is \$7.31; average spending conditional on buying is \$117.7, with a standard deviation of \$132.44. The idea then is to examine the incremental effect that the catalog had on this spending metric. Table (ref) presents summary statistics for the outcome and treatment. Figure (ref) displays the complete density of spending conditional on a purchase, which is quite skewed.
We estimated deep neural nets under a variety of architecture choices. In what follows we present eight examples and focus on one particular architecture to compute various statistics and tests to illustrate the use of the theory developed above. All computation was done using TensorFlow\textsuperscript{\tiny TM}.
For treatment effect and profit estimation we follow Equations (ref) and (ref). Because treatment is randomized, we apply Corollary (ref), and thus, only require estimates of the regression functions $\mu_t(\bm{x}) = E[Y(t) \vert \bm{X}=\bm{x}], \ t \in \{0,1\}$. An important implementation detail, from a computation point of view (recall Remark (ref)) is that we will estimate $\mu_0(\bm{x})$ and $\tau(\bm{x})$ (and thereby $\mu_1(\bm{x})$) jointly (results from separate estimation are available). To be precise, recalling Equations (ref) and (ref), we solve
where the minimization is over the relevant network architecture. Recall that, in the context of our empirical example $y_i$ is the customer's spending, $\bm{x}_i$ are her characteristics, and $t_i$ indicates receipt of a catalog. In this format, $\mu_0(\bm{x}_i)$ reflects base spending and $\tau(\bm{x}) = \mu_1(\bm{x}_i) - \mu_0(\bm{x}_i)$ is the conditional average treatment effect of the catalog mailing. In our application, this joint estimation outperforms separately estimating each $\mu_t(\bm{x})$ on the respective samples (though these two approaches are equivalent theoretically).
The details of the eight deep net architectures are presented in Table (ref). See Section (ref) for an introduction to the terminology and network construction. Most yielded similar results, both in terms of fit and final estimates. A key measure of fit reported in the final column of the table is the portion of $\hat{\tau}(\bm{x}_i)$ that were negative. As argued by Hitsch-Misra2018_WP, it is implausible under standard marketing or economic theory that receipt of a catalog causes lower purchasing. On this metric of fit, deep nets perform as well as, and sometimes better than, the best methods found by Hitsch-Misra2018_WP: Causal KNN with Treatment Effect Projections (detailed therein) or Causal Forests Wager-Athey2018_JASA. Figure (ref) shows the distribution of $\hat{\tau}(\bm{x}_i)$ across customers for each of the eight architectures. While there are differences in the shapes of the densities, the mean and variance estimates are nonetheless quite similar.
We present now results for treatment effects, utility/profits, and targeting policy evaluations. Table (ref) shows the estimates of the average treatment effect from the eight network architectures along with their respective $95\%$ confidence intervals. These results are constructed following Section (ref), using Equations (ref) and (ref) in particular, and valid by Corollary (ref). Because this is an experiment, we can compare to the standard unadjusted difference in means, which yields an average treatment effect of $2.561$.
Turning to expected profits, we estimate $\pi(s) = \mathbb{E} \big[s(\bm{X})(mY(1) - c) + \left(1 - s(\bm{X})\right) m Y(0) \big]$, adding a profit margin $m$ and a mailing cost $c$ to (ref) (our NDA with the firm forbids revealing $m$ and $c$). We consider three different counterfactual policies $s(\bm{x})$: (i) never treat, $s(\bm{x}) \equiv 0$; (ii) a blanket treatment, $s(\bm{x}) \equiv 1$; (iii) a loyalty policy, $s(\bm{x}_i) = 1$ only for those who had purchased in the prior calendar year. Results are shown in Table (ref). It is clear that profits from the three policies are ordered as $\pi(\text{never}) < \pi(\text{blanket}) < \pi(\text{loyalty})$.
For both the average effects of Table (ref) and the counterfactuals of Table (ref) there is broad agreement among the eight architectures both numerically and substantially. This may be due to the fact that the data is experimental, so that the propensity score is constant. In true observational data this may not be the case. We explore this issue in our Monte Carlo analysis below.
We conducted a set of placebo tests to examine whether the deep neural networks we use can truly recover causal effects. In particular, we take only the untreated customers in the data and randomly assign half to treated status.\footnote{We thank Guido Imbens suggesting this analysis.} We then ran the eight architectures of Table (ref), as in the true data. The conditional average treatment effects across the architectures are plotted in Figure (ref). We see that the “true” zero average effect is recovered precisely and with the expected distribution. The average treatment effect across all models is estimated to be around -0.024, compared to 2.56 in the original data. Exercises with different proportions of (placebo) treated customers revealed similar results.
To explore further, we focus on architecture \#3 and study subpopulation treatment targeting strategies following the ideas of Section (ref). (The other architectures yield similar results, so we omit them.) Architecture \#3 has depth $L=2$ with widths $H_1 = 30$ and $H_2 = 20$. The learning rate was set at 0.0001 and the specification had a total of 4,952 parameters. For this architecture, recalling Remark (ref), we added dropout for the second layer with a fixed probability of $1/2$. Using this architecture, we compare the blanket strategy (so $s_0(\bm{x})=1$) to targeting customers with spend of at least $\bar{y}$ dollars in the prior calendar year (prior spending is one of the covariates), in \$50 increments to \$1200. The policy class is therefore $\mathcal{S} = \{ s(\bm{x}) = \mathbbm{1}({\tt prior \ spend} > \bar{y}), \bar{y} = 0, 50, 100, \ldots, 1150, 1200\}$. Figure (ref) presents the results. The black dots show the difference $\big\{\hat{\pi}(\text{spend}>\bar{y}) - \hat{\pi}(\text{blanket})\big\}$ and the shaded region gives a pointwise 95% confidence band (to ease presentation, sample splitting is not used). We see that there is a significant difference between various choices of $\bar{y}$. Initially, targeting customers with higher spend yields higher profits, as would be expected, but this effect diminishes beyond a certain $\bar{y}$, roughly \$500, as fewer and fewer are targeted. The optimal policy estimate is $\hat{s}(\bm{x}) = \mathbbm{1}({\tt prior \ spend} > 400)$. In general, simpler policy classes may yield better decisions, but it is certainly possible to expand our search to different $\mathcal{S}$ by considering further covariates and/or transformations.
We conducted a set of Monte Carlo experiments to evaluate our theoretical results. We study inference on the average treatment effect, $\tau$ of (ref), under different data generating processes (DGPs). In each DGP we take $n=10,000$ i.i.d.\ samples and use 1,000 replications. For either $d=20$ or 100, $\bm{X}$ includes a constant term and $d$ independent uniform random variables, $\mathcal{U}(0,1)$. Treatment assignment is Bernoulli with probability $p(\bm{x})$, where $p(\bm{x})$ is the propensity score. We consider both (i) randomized treatments with $p(\bm{x}) = 0.5$ and (ii) observational data with $p(\bm{x})= ( 1+\exp(-\bm{\alpha}_p'\bm{x}) )^{-1}$, where $\alpha_{p,1} = 0.09$ and the remainder are drawn once as $\mathcal{U}(-0.55,0.55)$, and then fixed for the replications. For $d=100$, we maintain $\|\bm{\alpha}_p\|_0 = 20$ for the simplicity. These generate propensities with an approximate range of approximately $(0.30,0.75)$ and mean roughly $0.5$.
Given covariates and treatment assignment, the outcomes are generated according to \[ y_i = \mu_0(\bm{x}_i) + \tau(\bm{x}_i) t_i + \varepsilon_i , \qquad \mu_0(\bm{x}) = \bm{\alpha}_\mu'\bm{x} + \bm{\beta}_{\mu}'\varphi(\bm{x}) , \qquad \tau\left(x_{i}\right) = \bm{\alpha}_\tau'\bm{x} + \bm{\beta}_{\tau}'\varphi(\bm{x}), \] where $\varepsilon_i \sim \mathcal{N}(0,1)$ and $\varphi(\bm{x})$ are second-degree polynomials including pairwise interactions. For $\mu_0(\bm{x})$ and $\tau(\bm{x})$ we consider two cases, linear and nonlinear models. In both cases the intercepts are $\alpha_{\mu,1} = 0.09$ and $\alpha_{\tau,1} = -0.05$ and slopes are drawn (once) as $\alpha_{\mu,k} \sim \mathcal{N}(0.3,0.7)$ and $\alpha_{\tau,k} \sim \mathcal{U}(0.1,0.22)$, $k = 2, \ldots, d+1$. The linear models set $\bm{\beta}_{\mu} = \bm{\beta}_{\tau} = \bm{0}$ while the nonlinear models take $\beta_{\mu,k} \sim \mathcal{N}(0.01,0.3)$ and $\beta_{\tau,k} \sim \mathcal{U}(-0.05,0.06)$. Altogether, this yields eight designs: $d=20$ or 100, $p(\bm{x})$ constant or not, and outcome models linear or nonlinear.
For each design, we consider a variety of network architectures, all ReLU-based MLPs. These architectures are variants of the ones used in the empirical application (which were customized for the application). All networks vary in their depth and width, as spelled out in Table (ref).
Tables (ref) and (ref) show the results for all eight DGPs. Table (ref) shows randomized treatment while Table (ref) shows results mimicking observational data. Overall, the results reported show excellent performance of deep learning based semiparametric inference. The bias is minimal and the coverage is quite accurate, while the interval length is under control. Notice that the most architectures yield similar results with no architecture dominating the others. Further, the coverage and interval length are fairly similar with the more complex architecture not exhibiting any systematic patterns of length inflation.
None of the architectures we presented earlier used regularization. In typical empirical applications, including our own, researchers adopt architectures that employ dropout, a common method of regularization; see Remark (ref). Our own preliminary exploration of dropout and other forms of regularization found expected departures form nonregularized models. In most, but not all, cases the coverage remained accurate, but with increased bias and interval length compared to Table (ref) and (ref). The results preach caution when applying regularization in applications.
The utility of deep learning in social science applications is still a subject of interest and debate. While there is an acknowledgment of its predictive power, there has been limited adoption of deep learning in social sciences such as economics. Some part of the reluctance to adopting these methods stems from the lack of theory facilitating use and interpretation. We have shown, both theoretically as well as empirically, that these methods can offer excellent performance.
In this paper, we have given a formal proof that inference can be valid after using deep learning methods for first-step estimation. To the best of our knowledge, ours is the first inference result using deep nets. Our results thus contribute directly to the recent explosion in both theoretical and applied research using machine learning methods in economics, and to the recent adoption of deep learning in empirical settings. We obtained novel bounds for deep neural networks, speaking directly to the modern (and empirically successful) practice of using fully-connected feedfoward networks. Our results allow for different network architectures, including fixed width, very deep networks. Our results cover general nonparametric regression-type loss functions, covering most nonparametric practice. We used our bounds to deliver fast convergence rates allowing for second-stage inference on a finite-dimensional parameter of interest.
There are practical implications of the theory presented in this paper. We focused on semiparametric causal effects as a concrete illustration, but deep learning is a potentially valuable tool in many diverse economic settings. Our results allow researchers to embed deep learning into standard econometric models such as linear regressions, generalized linear models, and other forms of limited dependent variables models (e.g. censored regression). Our theory can also be used as a starting point for constructing deep learning implementations of two-step estimators in the context of selection models, dynamic discrete choice, and the estimation of games.
To be clear, we see our paper as a first step in the exploration of deep learning as a tool for economic applications. There are a number of opportunities, questions, and challenges that remain. For example, factor models in finance might benefit from the use of auto-encoders and recurrent neural nets may have applications in time series. For some estimands, it may be crucial to estimate the density as well, and this problem can be challenging in high dimensions. Deep nets, in the form of GANs are a promising tool for distribution estimation. There are also interesting questions remaining as to an optimal network architecture, and if this can be itself learned from the data, as well as computational and optimization guidance. Research into these further applications and structures is underway.
\singlespacing \begingroup \endgroup
\onehalfspacing