Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
94,775 characters · 23 sections · 77 citation commands
Structural Sieves
\newtheorem{dfn}{Definition}[section] \newtheorem{rem}{Remark}[section] \newtheorem{cor}{Corollary}[section] \newtheorem{thm}{Theorem}[section] \newtheorem{lem}{Lemma}[section] \newtheorem{notn}{Notation}[section] \newtheorem{con}{Condition}[section] \newtheorem{prp}{Proposition}[section] \newtheorem{pty}{Property}[section] \newtheorem{ass}{Assumption}[section] \newtheorem{ex}{Example}[section] \newtheorem*{cst1}{Constraint S} \newtheorem*{cst2}{Constraint U} \newtheorem{qn}{Question}[section]
\onehalfspacing
Artificial Neural Networks (ANN) have been extremely successful at solving certain statistical tasks involving the interpretation of sensory data and replication of human cognition. This paper pursues the question whether a potential advantage relative to other flexible approximation devices may carry over to certain types of economic data. We consider plausible generative models of production and discrete choice that have a latent structure satisfying shape or other qualitative constraints which do not directly translate into manageable restrictions on a “reduced form" in terms of the manifest (observable) variables of the model.
We propose a flexible, “deep" modeling approach for estimation of economic models. We envision a scenario in which reality is complex but modular in a way that is best captured by a, possibly nonlinear, latent variable model. That is, mappings transforming observable inputs into outputs can be disaggregated into simpler components with a simpler structure. We focus in particular on the role of shape, sparsity and separability restrictions that are imposed globally, or at intermediate stages of that transformation. These properties are in general not inherited by a composition of multiple such stages (“layers") consisting of multiple parallel units (“neurons"), but can be imposed as sign or exclusion restrictions in estimation approaches that replicate this modular structure.
In contrast to general-purpose sieves to estimate reduced-form relationships between the observed variables nonparametrically, our approach consists in using a parametric model for estimation which can then be made arbitrarily flexible. We see several benefits to such an approach - for one, if economic behavior that is best described in terms of a latent variable structure, a reduced form will typically display interaction effects in forcing variables if the relationship is not fully linear. A model with a similar latent variable structure may more readily reproduce such interaction effects even at a fairly low degree of approximation, especially if these are disciplined by a fairly low-dimensional latent factor structure. Furthermore, common qualitative model restrictions - such as monotonicity, convexity, separability, or sparsity - are usually imposed on economic primitives but have no analog in the resulting reduced form. We show that shape restrictions of this kind can be easily imposed as sign or exclusion restrictions in the deep generating model and also greatly restrict its expressivity, resulting in superior theoretical performance. Finally, we can also report ancillary aspects or predictions of the estimated model to help interpret and provide context to the main empirical results.
A particular feature of our approach is that we impose qualitative shape restrictions on the function of interest, concavity or quasi-concavity, and monotonicity, which are motivated by economic theory. These restrictions substantially reduce the complexity of the class of functions that can be represented by the network and therefore improve our ability to approximate the function flexibly. Furthermore, convexity may also reduce computational challenges from multiple local minima of the loss function when training the neural network.
The dimension-dependent optimal rates of nonparametric estimation derived in Sto80 impose an absolute constraint on how well an otherwise unconstrained statistical relationship between multiple variables can be estimated. With that in mind, our results speak to three key scenarios, depending on sample size and the researcher's confidence in the modular structure of the underlying DGP and the importance of shape restrictions.
Our approach therefore seeks to combine a parametric “structural" modeling philosophy with a nonparametric approach, where we do not assume a “correct" specification of the data generating process but an approximation that can be made arbitrarily flexible depending on the available amount of data. Through a nonparametric lens, it is known that the choice of estimators for nonparametric model components does not affect first-order asymptotic properties of regular estimators (see e.g. New94), however we are interested in settings when the data may not be sufficiently rich for these formal conclusions to be taken at face value. Rather, we propose a “regularization path" along a sequence of parsimonious parametric models that are adapted to the economic structure of the problem but also have the usual properties of a nonparametric sieve. Results can also be interpreted “parametrically" by reporting “projections" of components of the full model onto a nested, lower-dimensional version of the model in addition to a main effect of functional of primary interest.
We will develop our main ideas for nested CES models which are particularly suited for problems with convexity constraints. It would be possible to consider different sieves for modular problems without convexity, however at a greater challenge for controlling statistical complexity. A theory for estimation of nested separable nonparametric models using B-splines has been derived by HMa07.
Our setup has obvious parallels with popular methods in deep learning, most importantly multilayer feedforward neural networks (MFNN) and deep Boltzmann machines (DBM) consisting of multiple hidden layers, each of which processes inputs from the preceding layers into a vector of outputs. One key difference of our approach is that we consider settings in which these nested transformations are not mechanical but involve decisions by economic agents who may also anticipate their effect on subsequent layers of the process. We also allow for unobservables to enter at each stage of this process rather than treating the nested model as deterministic.
Various authors, most importantly HMa07, ESh16, MPo16, KKr17, BKo19, and SHi20 have shown that deep network architectures have a superior performance in uncovering compositional models of nested layers of smooth functions that satisfy certain separability or sparsity conditions. Compositionality of this type is common in latent variable models which have historically played an important role in describing economic decisions, see e.g. McF74, Tra09, or Hec79. Also, key concepts in economic models - economic activity, human or physical capital, etc. - are often not directly observable as a scalar quantity but are typically inferred as indices or proxies constructed from measurements or components of that latent concept. Deep learning has so far been most fruitfully applied in image and video processing which can also be viewed as latent variable problems, where surfaces defining an object in space can typically not all be seen at the same time on an image or may be occluded by other objects.
This paper explores the applicability of techniques from deep learning to economic problems, when observable outcomes of economic activity are best described as a result of multi-stage decision or production processes and when it is impractical to model the intermediate stages of that process explicitly e.g. due to the lack of direct measurements or the need to specify additional model components parametrically or nonparametrically. For an overview over recent developments in deep learning, see GBC16 and ZLLS20. Recent work by Yar17 and FLM19 derives inferential results for semiparametric inference using deep neural networks with a ReLU (rectified linear units) activation function.
The paper is organized as follows: Section 2 gives the general framework for the model and estimation, section 3 analyzes nested models of production as approximations for a general technology, and section 4 gives rates for approximating a discrete choice model with general dependence among taste shocks using multi-layer cross-nested Logit models. Section 5 discusses network architectures to impose additional qualitative constraints on the approximating network. Section 6 gives our main asymptotic result regarding the rate of nonparametric estimation with and without additional shape restrictions, section 7 concludes.
We consider the problem of flexible estimation of a reduced-form relationship between variables $\mathbf{X}:=(X_{1},\dots,X_{K})'$ (“explanatory variables", “covariates", or “inputs") and $\mathbf{Y}:=(Y_{1},\dots,Y_{M})'$ (“outcomes", “outputs"). This could be for the purpose of prediction, causal inference, or in the context of a structural model. We also let $\mathbf{Z}:=(\mathbf{Y}',\mathbf{X}')'$.
We assume that the researcher observes a sample of $n$ i.i.d. units $i$ from a distribution with joint p.d.f. \[f_0(\mathbf{y},\mathbf{x}) = f_{Y|X}(\mathbf{y}|\mathbf{x})f_X(\mathbf{x})\] assuming that the conditional p.d.f. is also well-defined for $\mathbf{x}$. We allow for the possibility of missing data, where certain components of $\mathbf{Y}$ and/or $\mathbf{X}$ may not be observed for some units in the sample. We assume that the object of interest is a function depending on that distribution, \[\mu_0(\mathbf{z}):=\mu(\mathbf{z};f_0)\] Here we will primarily consider cases in which $\mu_0(\mathbf{z})$ denotes a conditional or unconditional expectation or probability for an event in $\mathbf{Y}$ given $\mathbf{X}$.
We assume that $\mu_0$ is defined as the minimizer of expected loss
where the random vector $\mathbf{Z}$ is distributed according to $f_0(\mathbf{z})$. Following FLM19 we assume that the loss function $\ell(\mu,\mathbf{z})$ is Lipschitz \[|\ell(\mu,\mathbf{z})-\ell(\mu',\mathbf{z})|\leq C_l|\mu(\mathbf{z})-\mu'(\mathbf{z})|\] for some constant $C_l<\infty$, and satisfies \[c_1\mathbb{E}[(\mu(\mathbf{Z})-\mu_0(\mathbf{Z}))^2]\leq\mathbb{E}[\ell(\mu,\mathbf{Z})-\mathbb{E}[\ell(\mu_0,\mathbf{Z})]\leq c_2\mathbb{E}[(\mu(\mathbf{Z})-\mu_0(\mathbf{Z}))^2]\] for finite constants $c_2>c_1>0$.
For the purposes of this paper, the first leading case is that of a conditional expectation function, $\mu_0(\mathbf{x}):=\mathbb{E}[Y|\mathbf{X}=\mathbf{x}]$ for scalar $Y$, which satisfies ((ref)) with $\ell(\mu,\mathbf{z}):=(y-\mu(\mathbf{x}))^2$. A second case of interest is that of a conditional choice probability where $\mathbf{Y}$ takes values in a finite set $\{y_1,\dots,y_M\}$ and $\mu(y,\mathbf{x}):=\mathbb{P}(Y=y|\mathbf{X}=\mathbf{x})$ is the conditional probability of $Y=y$ given $\mathbf{X}$. One loss function satisfying ((ref)) for this problem is $\ell(\mu,\mathbf{z}):=-\sum_{m=1}^M1\hspace{-2.5pt}\textnormal{l}\{Y=y_m\}\log\mu(y_m,\mathbf{x})$.
Nonparametric estimation of conditional mean or distribution functions is known to be subject to a curse in dimensionality in the number of covariates and/or outcomes, see Sto80, which is inherent in the estimation problem and not the particular technique that is used for estimation. The key challenge here is that in the absence of additional separability restrictions, a flexible model for this relationship would have to account for any possible interaction effects between functions of two or more components of $\mathbf{X}$.
The approach put forward in this paper does not sidestep or remedy this challenge, however we propose a nonlinear sieve that (a) incorporates common shape constraints on the underlying economic model (most importantly convexity), and (b) specifically targets the interaction effects that would result from standard models of economic decisions and optimization. We argue that when observable data results from nested nonlinear models, then separability of the reduced form will be the exception and not the rule, and “deep" architectures may have an advantage at replicating these interaction effects with a smaller number of parameters.
To appreciate this point, consider a linear index model with a scalar outcome variable $Y_i$ and regression function \[\mu(\mathbf{x}):=\mathbb{E}[Y_i|\mathbf{X}_i=\mathbf{x}] \equiv G(\mathbf{x}'\boldsymbol\beta)\] with coefficient $\beta\in\mathbb{R}^K$ and link function $G:\mathbb{R}\rightarrow\mathbb{R}$ that is twice continuously differentiable. The cross-partial derivatives of this model, \[\frac{\partial^2}{\partial x_k\partial x_l}\mu(\mathbf{x})=\beta_k\beta_lG''(\mathbf{x}'\boldsymbol\beta)\] are generally non-zero, unless the function $G(\cdot)$ is affine. As this simple example illustrates, a single nonlinear transformation (“activation") may be sufficient to mask any separability properties that may have been satisfied in preceding stages of this model. However those interaction effects are also tightly constrained by the simple parametric structure of this model, so an estimation approach exploiting that index structure may in fact estimate the conditional mean function at a much faster rate.
A similar point was established formally for a class of nonlinear compositional or hierarchical interaction models by BKo19 and SHi20. The potential benefits of deeper, rather than shallow, network architectures to approximate compositional functions have recently been analyzed by ESh16 and MPo16. Specifically, if the model in $K$ regressors is a composition of a bounded number of functions of at most $d^*\leq K$ variables at each step, estimation by a sufficiently deep network can achieve estimation at a nonparametric rate depending on $d^*$ rather than $K$. This suggests that nested models may have an advantage at leveraging sparsity or separability restrictions at intermediate stages to achieve faster convergence rates and mitigate the curse of dimensionality.
As a second key feature, analyzing the data-generating process as a latent variable model naturally incorporates missing or mismeasured data into estimation. This obviously includes the cases where the available data are complete or are thought to conform to a “missing at random" assumption (see Rub76), but also situations where missing data is at the heart of the problem, including imperfect measurement (JGo75, CHS10), self-selected samples (Hec79,APo93), and causal inference (Ney23,Rub78). More generally, such an approach can be adapted to situations in which data quality is uneven or observed variables are only proxies the relevant economic concept.
We develop a technique for solving and estimating models involving a large number of nested - discrete or continuous - intermediate decisions to approximate complex statistical relationships between inputs and outputs. One key distinguishing feature of our approach is that “activations" at intermediate stages incorporate not only states in past layers that are carried forward mechanically, but also constitute the result of choices by an optimizing agent who anticipates outputs in subsequent layers, iterating future states backward.
This results in an multilayer perceptron (MLP) with possibly nonlinear activation functions in which intermediate states can be fed both forward and backwards, so that the resulting graph is not necessarily acyclic. For the purposes of estimation, these intermediate stages are latent and not directly observable to the researcher. We then show that such a model can be made arbitrarily flexible by increasing the number of units (“nodes") in each hidden stage (“layer").
For expositional clarity, we focus on two prototypes for such generative models with latent discrete or continuous decisions, which may be adapted or combined flexibly. The first model concerns a model for a production technology, where initial inputs are transformed into intermediate goods, which in turn serve as inputs in subsequent intermediate and final stages of production. Specifically, we assume that production takes place in $S+1$ stages (“layers"), where at the $s$th layer there are separate technologies (“neurons") to produce intermediate goods $k=1,\dots,K_s$ for $s=0,\dots,S$.
The technology for producing intermediate good $k$ is given by the production function \[w_k^{(s)} = \tilde{F}_k^{(s)}(w^{(s-1)}_1,\dots,w^{(s-1)}_{K_{s-1}})\] where $w^{(s-1)}_l$ denotes the quantity of the $l$th intermediate good from the $(s-1)$ stage employed in the production of the $k$th intermediate good at stage $s$. We show below that the optimal (cost-minimal) production plan can be characterized recursively as the cost-minimization problem at neuron $k$ in layer $s$, with input prices $\pi_1^{(s-1)},\dots,\pi_{K_{s-1}}^{(s-1)}$ given by marginal cost of production at stage $s-1$, and given desired output level $v_k^{(s)}$ determined by factor demand at stage $s$. The values of solution of $v_k^{(s)},\pi_k^{(s)}$ resulting from this constrained optimization problem are determined recursively via the activation mappings
that take states $\left(h_l^{(s-1)}\right)_{l=1}^{K_{s-1}}:=\left(\pi_l^{(s-1)},v_l^{(s-1)}\right)_{l=1}^{K_{s-1}}$ and $\left(h_l^{(s+1)}\right)_{l=1}^{K_{s+1}}:=\left(\pi_l^{(s+1)},v_l^{(s+1)}\right)_{l=1}^{K_{s+1}}$, respectively, as inputs and produce outputs $\left(h_k^{(s)}\right)_{k=1}^{K_s}$.
The second model is a model of nested discrete decisions where at the $k$th neuron in the $s$th layer an agent chooses among $K_{s+1}$ discrete nests in the subsequent layer. There are no flow payoffs, but each terminal node in the top layer is associated with a random utility where the joint distribution of taste shocks corresponds to that for a cross-nested Logit (CNL) with that nesting structure (see Vov97 and WKo01).
This model can be reparametrized as sequential choice among nests starting at the bottom layer, given the inclusive (continuation) values $v_k^{(s)}$ for the $k$th nest in the $s$th layer. We also denote the unconditional probability of reaching the $k$th nest in layer $s$ with $\pi_k^{(s)}$. Inclusive values are determined by iterating the activation mappings
backwards from the $S$th (top) layer to the root node of the graph, and the conditional choice probabilities for the nests can be obtained by iterating forward over
from the root node. For the CNL specification the mappings $\psi_k^{(s)}(\cdot),\phi_k^{(s)}(\cdot)$ are available in closed form.
The primitive functions $F_k^{(s)}(\cdot)$ in the generative model are generally not known and are typically not nonparametrically identified (see e.g. HMa07 for feedforward networks). We therefore do not aim to estimate intermediate activation functions but are primarily interested in estimating the reduced form $f_0(\mathbf{y},\mathbf{x})$ given shape restrictions on the primitive functions. To this end we propose an artificial neural network with the capacity to approximate the activation functions $\psi_k^{(s)},\phi_k^{(s)}$ flexibly while incorporating qualitative constraints on $F_k^{(s)}(\cdot)$.
We construct a network of $\tilde{S}+1$ layers, where layer $s$ has $\tilde{K}_s$ neurons for $s=0,\dots,\tilde{S}$. For neuron $k$ in the $s$th layer we choose the production function (or flow utility for a network of nested discrete decision) from a parametric family, \[\tilde{F}_k^{(s)}(w_1,\dots,w_{\tilde{K}_{s-1}}) = \tilde{f}\left(w_1,\dots,w_{\tilde{K}_{s-1}};\boldsymbol\theta_k^{(s)}\right)\] where the vector $\boldsymbol\theta_{k}^{(s)}$ consists of parameters $\boldsymbol\beta_k^{(s)}=(\beta_{k1}^{(s)},\dots,\beta_{kK_{s-1}}^{(s)})$ governing the effect of the preceding layer on activation of neuron $k$ in layer $s$, and potentially additional shape parameters $\boldsymbol\varrho_k^{(s)}$.
We then solve the constrained optimization problem analogous to the generative model given state variables $v_1^{(s+1)},\dots, v_{\tilde{K}_{s+1}}^{(s+1)}$ and $\pi_1^{(s-1)},\dots,\pi_{\tilde{K}_{s-1}}^{(s-1)}$ to obtain the activation functions
for some collection of parameters $\boldsymbol\theta_{k-}^{(s)},\boldsymbol\theta_{k+}^{(s)}$ from the adjacent layers, $s-1$ and $s+1$.
Such a structure with $S$ layers and activation functions $\psi_k^{(s)},\phi_k^{(s)}$ defines a class $\mathcal{H}_{\mathbf{K}}^S$ of neural networks that is indexed by the free parameters $\boldsymbol\theta$. We denote the activations of the hidden units that are consistent with the recursions ((ref)) and ((ref)) with $h_k^{(s)}(\boldsymbol\theta)$.
Given the sample $\mathbf{z}_1,\dots,\mathbf{z}_n$, we can train this network, where we identify $\mu(\mathbf{z}_i)\equiv\mu(\mathbf{z}_i;\boldsymbol\theta)$ with activations $h_{ki}^{(s)}(\boldsymbol\theta)$ of certain neurons, in most cases in the top and/or bottom layer of the network. Specifically, we estimate $\mu_0(\mathbf{z})$ by minimizing the average training error,
for some bound $B<\infty$. Here, minimization over the class $\mathcal{H}_{\mathbf{K}}^{S}$ is equivalent to minimization over the parameter $\boldsymbol\theta$. Note that training error in ((ref)) is the empirical analog of expected population loss in ((ref)), the criterion defining the target $\mu_0$. The trained network can then be used to compute $\hat{\mu}(\mathbf{z})$ for arbitrary values of $\mathbf{z}$.
Our main theoretical results concern statistical properties of $\hat{\mu}$ as an estimator for $\mu_0$, and recommendations for the choice of the number and size of hidden layers. The minimization problem in ((ref)) can be solved adapting commonly used algorithms for conventional feedforward networks, including stochastic gradient descent, or Adam. We argue in Appendix (ref) that the gradient can be approximated using recursive application of the chain rule even when the graph is not acyclic.
The hidden layers of this model are generally not identified in the sense that the minimum in ((ref)) may be attained (exactly or to an approximation) at multiple values of $\boldsymbol\theta$. We do not address this issue explicitly in this paper, rather the main object of interest for the purposes of this paper is a reduced-form relationship characterizing the joint distribution of the observable variables $\mathbf{z}_i$. A structural interpretation of the hidden layers may require additional restrictions and normalizations, see also HMa07 for a discussion for the case of nonparametric estimation of nested regression functions.
Consider a model for production of outputs $Y_i^*$ from inputs $X_i$ where we do not observe $Y_i^*$ directly, but rather a collection of measurements $Y_i:=(Y_{1i},\dots,Y_{Mi})$. For example there is no direct agreed upon measure for human capital, rather the term captures groups of distinct skills and attitudes, see e.g. CHS10. In a typical setting, we observe parental investments, indicators of educational (e.g. test scores) and economic achievement (e.g. labor force status or income) and want to determine the effects of various educational investments or interventions.
JGo75 proposed the parametric Multiple Indicators and Multiple Causes (MIMIC) model with a single hidden variable to estimate such a production process. CHS10 analyzed a semi-parametric model for human capital production with two latent factors, cognitive and non-cognitive skills which could in principle be further disaggregated into more distinct latent skills to explain the link between childhood investments and testing or economic performance.
Our approach involves a disaggregate description of a technology where we the number of latent factors can be chosen flexibly, and we may observe only some inputs directly as quantities, and possibly prices or cost shifters for some others. This approach involves a trade-off between the predictive performance of the model and interpretability, which should be resolved depending on the ultimate objective and the functionals/parameters $\mu$ used to summarize this technology.
Suppose that we are interested in the causal effect of a binary intervention (“treatment") $D\in\{0,1\}$ on a scalar outcome variable $Y$. Using the Neyman-Rubin potential outcomes framework (see e.g. IRu15), each unit $i$ is associated with two possible outcomes, $Y(1)$ with the intervention $D=1$, and $Y(0)$ in the absence of an intervention, $D=0$. For any unit in the sample, only the potential outcome corresponding to the realized intervention, $Y=Y(D)=(1-D)Y(0) + DY(1)$, is observed. In this setup, the unit-level treatment effect $\Delta:=Y(1)-Y(0)$ can not be determined from the observable data, so that the problem of causal inference can be viewed as a missing-data problem.
Suppose that we have a sample of $n$ units for which we observe all relevant covariates $X$, the realized treatment $D$, and the outcome $Y=Y(D)$ Furthermore, assume that we observe additional variables $\mathbf{Z}$ that are conditionally independent of $Y(0),Y(1)$ given $X$, but not of $D$ (under the ignorability assumption for observational designs, $D$ itself would meet that requirement). We can then represent this problem in our framework with a bivariate outcome vector $(Y(1),Y(0))'$, where the network architecture can be constrained such that $Z$ serves as an indirect input determining $D$, but is excluded from the latent factor model determining potential outcomes $(Y(0),Y(1))$.
After training the MLP relating $X_i,\mathbf{Z}_i$ to $Y_i,D_i$, we can simulate from the estimated distribution of $Y_i,D_i$ given $X_i$ and different values of $\mathbf{Z}_i=z$. In particular we can evaluate the conditional expectation of $Y_i(1)-Y_i(0)$ among the compliers according to that estimated distribution. We may in principle also compute other functionals of the estimated distribution that are not nonparametrically identified, however these will in general not be consistent estimates of their population counterparts but vary with the parametrization and the choice of starting values during training of the MLP.
Following the seminal work by Bec64 and Bec65, decisions on fertility, investment in education, and other forms of economic activity have been cast as a problem of optimal allocation of resources towards efficient home production. Nested models of production have been considered by McF78b, PWa87, or PWa75 in the context of household production, where PRu95 established a general approximation property for nested Constant Elasticity of Substitution (CES) function with regards to the cost functions.
Misspecified production functions can easily lead to inaccurate policy conclusions or obscure economically relevant features of real-world processes. As an important example, CHe07 and CHS10 show in their seminal work on skill formation that an extension of a single-skill model with a single period of production to two skill types (cognitive and non-cognitive) and multiple successive periods of investments and accumulation uncovers richer dynamic complementarities in the skill-formation process, and matches important empirical facts about skill formation that can't be properly explained by the simpler model. Our approach would allow to disaggregate latent production stages into a larger number of distinct skills (“intermediate goods") and stages of production. However our focus is on approximating a reduced form mapping between observable inputs and outputs, so we do not analyze identification of intermediate technologies, which was central to their contribution.
We consider a model of production with a large number of unobserved intermediate goods to approximate a general technology to approximate a mapping from partially observed factor quantities and/or prices to multiple outputs. There are $K$ basic inputs $\mathbf{x}=\left(x_1,\dots,x_K\right)'$ and $M$ final outputs $\mathbf{y}=(y_1,\dots,y_M)'$. We describe the technology as the production possibility set $\mathcal{Y}_0\subset\mathbb{R}^{M+K}$, where a pair of $(\mathbf{y}',\mathbf{x}')'$ is interpreted as a transformation of inputs $-\mathbf{x}$ to produce outputs $\mathbf{y}$. We also let \[\mathcal{Y}_0(\mathbf{x}):=\left\{\mathbf{y}:(\mathbf{y}',\mathbf{x}')'\in\mathcal{Y}_0\right\}\] denote the set of feasible outputs given inputs $\mathbf{x}$. We maintain the following assumptions on the technology:
While these assumptions are common in classical production theory, they are also restrictive. Taken together, parts (1) and (3) preclude the existence of fixed costs of production. Convexity of the production possibility set also restricts returns to scale to be nonincreasing. A commonly imposed, weaker version of (3) only assumes that input sets are convex from below, which is still sufficient for duality theory to apply (see e.g. McF78b). For our approximation results, weakening parts (1) and (3) would require recentering of intermediate inputs, thus introducing additional free parameters and complicating the approximating argument. We therefore maintain the stronger set of assumptions throughout.
We can characterize the convex hull of $\mathcal{Y}_0(\mathbf{x})$ in terms of its support function, \[\mu_0(\mathbf{u},\mathbf{x}):=\sup\left\{\langle\mathbf{y},\mathbf{u}\rangle:(\mathbf{y},\mathbf{x})\in\mathcal{Y}_0\right\},\hspace{0.5cm} \mathbf{u}\in\mathbb{R}_+^M\] For the case of a single output, $\mu_0(y,\mathbf{x})$ corresponds to the usual production function, and for the general case $M\geq1$, the support function $\mu_0(\mathbf{e}_m,\mathbf{x})$ evaluated at the $m$th unit vector $\mathbf{e}_m$ yields the highest achievable level for the $m$th output given inputs $\mathbf{x}$. We generally assume efficient production, $\langle\mathbf{y},\mathbf{y}\rangle =\mu_0(\mathbf{y},\mathbf{x})$, however we do not explicitly model selection among feasible output combinations on the frontier.
To approximate the production function associated with $\mathcal{Y}_0$, we consider a production network of the following form: production takes place in $S$ stages, where at stage $s$, the $K_{s-1}$ intermediate inputs $\mathbf{w}^{(s-1)}$ produced during the preceding stage are transformed into $K_{s}$ intermediate outputs $\mathbf{w}^{(s)}=(w_1^{(s)},\dots,w_{K_s}^{(s)})'$.
At each stage, the transformation of intermediate inputs into intermediate outputs is described by Constant Elasticity of Substitution (CES) production functions. Specifically, we define
for parameters $\boldsymbol\theta_k^{(s)}:=(\beta_{k1}^{(s)},\dots,\beta_{kK_{s-1}}^{(s)},\varrho_k^{(s)},\tau_k^{(s)})'$, and assume that output at the $k$th neuron in the $s$th layer is given by
Here, $w_{lk}^{(s-1)}$ is the quantity of intermediate good $w_l^{(s-1)}$ committed to the production of the intermediate output $w_k^{(s)}$, and we set $w_0^{(s-1)}\equiv1$ for each $s=1,\dots,S$, which can be thought of as a non-discretionary input.
Intermediate goods may be rival or non-rival as inputs for future stages of production, and we assume that they cannot be acquired from outside sources, but can only be produced using this technology. Note that the stage technologies in ((ref)) include the identity $w_k^{(s)} = w_l^{(s-1)}$ for any $k,l$, so that intermediate outputs could be passed through multiple stages of production unaltered.
We then identify the intermediate inputs at the first stage with the basic inputs, so that any feasible production plan must satisfy the resource constraints \[w_l\geq\left\{
\right.\] Similarly, the intermediate outputs at the last stage are identified with the final outputs, \[y_m\leq w_m^{(S)}\hspace{0.3cm}\textnormal{for any }m=1,\dots,M=:K_S\] Furthermore, any feasible production plan must satisfy the resource constraints for intermediate goods,
For simplicity, the remainder of this paper will only consider the case in which all intermediate goods are rival.
We denote the resulting technology with $\mathcal{Y}^{S}_{\mathbf{K}}\subset\mathbb{R}_+^{M+K}$, that is the set of input/output vectors $(\mathbf{y}',\mathbf{x}')'$ that can be achieved via a feasible production plan with intermediate production functions ((ref)), satisfying constraints ((ref)), where the vector $\mathbf{K}=(K_1,\dots,K_S)'$. We denote the corresponding support function with \[\mu_{\mathbf{K}}^S(\mathbf{u},\mathbf{x}):=\sup\left\{\langle\mathbf{y},\mathbf{u}\rangle:(\mathbf{y},\mathbf{x})\in\mathcal{Y}_{\mathbf{K}}^S\right\},\hspace{0.5cm} \mathbf{u}\in\mathbb{R}_+^M\]
The feasible set $\mathcal{Y}^{S}_{\mathbf{K}}\subset\mathbb{R}_+^{M+K}$ can be described recursively as follows: at stage $s$, we let $\mathcal{W}^{(s)}$ denote the feasible set of quantities for the intermediate goods from stages $s'=0,\dots,s$, that is the set of values $(\mathbf{w}^{(s)},\mathbf{w}^{(s-1)},\dots,\mathbf{w}^{(0)})$ such that there exist $\tilde{w}_{kl}^{(s-1)}$ for $k=1,\dots,K_s$ and $l=1,\dots,K_{s-1}$ such that $F_k^{(s)}(w_{k1}^{(s-1)},\dots,w_{kK_{s-1}}^{(s-1)})\geq w_k^{(s)}$ and $(\mathbf{w}^{(s-1)}+\tilde{\mathbf{w}}^{(s-1)},\mathbf{w}^{(s-2)},\dots,\mathbf{w}^{(0)})\in\mathcal{W}^{(s-1)}$, where $\tilde{\mathbf{w}}^{(s-1)}=\sum_{k=1}^{K_s}\left(\tilde{w}_{k1}^{(s-1)},\dots,\tilde{w}_{kK_{s-1}}^{(s-1)}\right)'$. Iterating from $\mathcal{W}^{(0)}:=\mathbb{R}^K$ we obtain the set $\mathcal{W}^{(S)}$ of feasible combinations of inputs and intermediate goods for all $S$ stages, so that we can define $\mathcal{Y}^{S}_{K_1,\dots,K_S}$ as the intersection of $\mathcal{W}^{(S)}$ with $\mathbf{R}^K\times\{0\}\times\dots\times\{0\}\times\mathbb{R}^M$, projected on its $K+M$ nontrivial coordinates.
By construction, $\mathcal{W}^{(s)}$ is convex for each $s$. Since $\mathcal{Y}^{S}_{K_1,\dots,K_S}$ is the intersection of that set with a linear subspace, it is also convex. Furthermore, if $\varrho_k^{(s)}>-\infty$ for all $s,k$, the efficient frontier of $\mathcal{W}^{(s)}$ is continuously differentiable to any order, implying that the boundary of the feasible set $\mathcal{Y}^{S}_{K_1,\dots,K_S}$ is also arbitrarily smooth. If $\varrho_k^{(s)}=-\infty$ for some stages and intermediate goods, we will instead consider limits as $\varrho_k^{(s)}\rightarrow-\infty$ for arguments relying on differentiability of the boundary.
In sum our approximating model consists of nested CES technologies, where the elasticity of substitution $\frac{1}{1-\varrho}$ and the degree of homogeneity $\tau$ are allowed to differ across intermediate technologies. Our main claim in this section is that this technology of nested CES aggregators is sufficient to approximate any technology $\mathcal{Y}_0$ satisfying our main assumptions.
See the appendix for a proof. The argument is in fact constructive and establishes that only two stages of production are needed, where basic inputs are perfect complements ($\varrho_k^{(1)}=-\infty$ for each $k$) in the production of the intermediate outputs $\mathbf{w}^{(1)}$, which in turn are perfect substitutes ($\varrho_k^{(2)}=1$ for all $k$) in the production of final outputs $\mathbf{y}$. This nested production function spans a polytope of production possibilities in $\mathbb{R}_+^{K+M}$. The vertices spanning that polytope correspond to production plans involving only a single intermediate output, with all remaining components of $\mathbf{w}^{(1)}$ equal to zero. Since the technology for intermediate output $w_k^{(1)}$ has constant returns to scale up to output $\beta_{k0}^{(1)}$, the second-stage technology then convexifies that vertex set. We can then approximate $\mathcal{Y}_0$ arbitrarily well by adding a sufficient number of appropriately chosen vertices to that polytope. It is also instructive to compare the approximation rate to the approximation bound for RELu approximation of smooth functions in Yar17, where the assumption of convexity results in a rate comparable to that for the case of weak derivatives up to order 2.
We next derive the activation functions $\phi(\cdot),\psi(\cdot)$ implied by maximizing behavior in the production model. Given the technology $\mathcal{Y}^{S}_{K_1,\dots,K_S}$ and input prices $\mathbf{p}=(p_1,\dots,p_K)$ we assume that outputs $\mathbf{\hat{y}}$ are produced using a cost-optimal, feasible production plan. That is, \[(\mathbf{\hat{y}}',\mathbf{\hat{x}}')' :=\arg\min_{\mathbf{x}}\left\{\mathbf{p'x}:(\mathbf{\hat{y}}',\mathbf{x}')'\in\mathcal{Y}^{S}_{K_1,\dots,K_S}\right\}\] We characterize the optimal production plan at each of the $S$ stages both in terms of quantities of intermediate inputs employed in each intermediate technology, as well as the implied price of each intermediate good.
At stage $s$, the optimal production plan for $w_k^{(s)}\equiv v_k^{(s)}$ units of the $k$th intermediate good given implied input prices $\pi_0^{(s-1)},\dots,\pi_{K_{s-1}}^{(s-1)}$ for the intermediate outputs from the $(s-1)$st stage is the solution to the cost-minimization problem \[\min_{v_{1k}^{(s-1)},\dots,v_{K_{s-1}k}^{(s-1)}}\sum_{l=0}^{K_{s-1}}\pi_l^{(s-1)}v_{lk}^{(s-1)}\hspace{0.5cm}\textnormal{subject to } F_k^{(s)}(v_{0k}^{(s-1)},\dots,v_{K_{s-1}k}^{(s-1)}) = v_k^{(s)}\] where the non-discretionary input is held fixed at $v_{0k}^{(s-1)}=1$.
The value of this constrained optimization problem is given by the cost function \[C_k^{(s)}(\boldsymbol{\pi}^{(s-1)};v_k^{(s)})\equiv C\left(\boldsymbol{\pi}^{(s-1)};v_k^{(s)};\boldsymbol\theta_k^{(s)}\right)\] where
Our framework treats quantities and implicit prices of the intermediate goods as latent (hidden) variables which are determined recursively by solving the cost-minimization problem for stages $s=1,\dots,S$. Specifically, the price of intermediate good $k$ corresponds to the marginal cost of production of an additional unit of the good,
where we define
Given prices, optimal factor demand for intermediate input $k$ for intermediate technology $l$ at stage $s+1$ is \[v_{kl}^{(s)}=\left(\frac{\pi_k^{(s)}}{\beta_{lk}^{(s+1)}}\right)^{\frac1{\varrho_l^{(s+1)}-1}} \left[\sum_{j=1}^{K_{s}}\left(\frac{(\pi_j^{(s)})^{\varrho_l^{(s+1)}}}{\beta_{lj}^{(s+1)}}\right)^{\frac1{\varrho_l^{(s+1)}-1}} \right]^{-\frac1{\varrho_l^{(s+1)}}} \left((v_l^{(s+1)})^{\frac{\varrho_l^{(s+1)}}{\tau_l^{(s+1)}}}-(\beta_{l0}^{(s+1)})^{\varrho_l^{(s+1)}}\right)^{\frac1{\varrho_l^{(s+1)}}} \] so that total factor demand for intermediate good $k$ is
where we define
For input levels $\mathbf{x}$ and outputs $\mathbf{y}$, we constrain $v_k^{(0)}\equiv x_k$ for $k=1,\dots,K$ and $v_m^{(S)}\equiv y_m$ for $m=1,\dots,M$. If factor prices for inputs are observed, we also set $\pi_k^{(0)}\equiv p_k$, otherwise we treat prices as latent as well. In particular, we choose good 1 as the num\'eraire, $\pi_1^{(0)}=1$, so that for an interior solution to the cost minimization problem at node $1$ in layer $s=1$, the first-order conditions for optimal input choices require \[\pi_k^{(0)}= \left(\frac{\beta_{1k}^{(1)}}{\beta_{11}^{(1)}}\right)^{\varrho_1^{(1)}}\left(\frac{v_{k1}^{(1)}}{v_{11}^{(1)}}\right)^{\varrho_1^{(1)}-1}\\ \nonumber=:\phi_{k}^{(0)}(\mathbf{v}^{(1)};\boldsymbol\theta_k^{(1)})\] for $k=2,\dots,K$. Intermediate good 1 in layer 1 was chosen arbitrarily, so if only a subset of inputs are employed in nonzero quantities to produce that good, implied factor prices for the $K$ initial inputs can be inferred from the first-order conditions at other intermediate production nodes in layer $1$. Additional normalizations may be required if nodes in that layer only use non-overlapping subsets of the initial inputs.
To summarize, the production system is characterized by hidden states $h_k^{(s)}:=(v_k^{(s)},\pi_k^{(s)})$, where at the first stage $\pi_k^{(0)} = p_k$ and $v_k^{(0)}=x_k$ for $k=1,\dots,K_0\equiv K$, and at the last stage, $v_k^{(S)}=y_k$ for $k=1,\dots,K_{S}\equiv M$. Latent prices $\pi_k^{(s)}$ are determined recursively by iterating ((ref)) forward from $s-1$ to $s$ starting at $s=1$, and quantities $v_k^{(s)}$ by iterating ((ref)) backward from $s+1$ to $s$, starting at $s=S$.
We consider training the network based on a sample of $n$ observed combinations $\mathbf{z}_i=(\mathbf{y}_i,\mathbf{x}_i')'$ of inputs and outputs. We allow for the possibility of measurement error in output levels as well as hidden inputs which we treat as stochastic. We therefore treat the technology $\mathcal{Y}^*$ as a random closed set and seek to approximate its (selection) expectation $\mathcal{Y}_0:=\mathbb{E}[\mathcal{Y}^*]$.\footnote{Following Mol05 the selection expectation of the random set $\mathcal{Y}^*:\Omega\rightarrow \mathcal{C}$ for some probability space $\Omega$ and the closed sets $\mathcal{C}$ in $\mathbb{R}^{K+M}$ is the closure of the set of all expectations of integrable selections, $\upsilon(\omega)\in\mathcal{Y}(\omega)$ for $\omega\in\Omega$.} It is known that when the distribution of $\mathcal{Y}^*$ is non-atomic, $\mathcal{Y}_0$ is a convex subset of $\mathbb{R}^{K+M}$ (see Theorem 1.15 in chapter 2 of Mol05).
We furthermore assume that production is efficient in that any combinations of inputs (including unobserved inputs) and correctly measured outputs correspond to points on the efficient frontier of $\mathcal{Y}^*(\omega)$. We do not explicitly model the choice of a particular output combination on the efficient frontier in the multiple output case. Rather we assume that production and demand are separable, and at least in principle a model for demand could be estimated separately. We then match observed combinations of quantities to points on the frontier predicted by the ANN. Training loss is given by $\ell(\mathbf{z}_i,\boldsymbol\theta):=\sum_{m=1}^M(y_{mi} - v_{mi}^{(S)}(\boldsymbol\theta))^2$ where $v_{mi}^{(S)}(\boldsymbol\theta)$ is the activation of the $m$th neuron in the top layer given inputs $x_{1i},\dots,x_{Ki}$.
More generally, the $k$th observable covariate $X_{ki}$ could interpreted either as “price"/cost shifter or the quantity of an input for production, and matched accordingly to either the price $\pi_{ki}^{(0)}$ or the quantity $v_{ki}^{(0)}$ in the initial layer. An implementation of this neural network could also incorporate additional unobserved inputs at intermediate neurons, with one or multiple independent random draws from some continuous distribution with non-negative support. This would also define smooth likelihoods/energy functions for the intermediate stages for an implementation as a Deep Boltzmann machine. We are not aware of a “conjugate" distribution for such an unobserved input that would yield closed-form solutions for factor demand or price equations at intermediate stages, but such an approach would in general have to rely on simulation by generating several replications for each unit, where data could also be permuted at random to reflect other invariance restrictions. While implementation is less straightforward, such an approach would at least conceptually be analogous to convolutional neural networks commonly used for image and video processing.
Models for nested discrete decisions - binary or multinomial - have long been used to model individual choice behavior, see BAk73, McF74, and McF78. HSW89 established that multi-layer feedforward networks with a “squashing" activation function - a nondecreasing function mapping its scalar argument to the unit interval - can approximate any Borel measurable function on compacta. Our focus is on a flexible framework for multinomial discrete decisions, generalizing BAk73's nested Logit model to approximate all members of McF78's Generalized Extreme Value (GEV) class.
There are many examples when a single nesting structure may be too restrictive to model realistic cross-substitution patterns among a population of agents, especially when alternatives are complex or bundles of more elementary choices. For example when choosing a travel destination, Barcelona may appear to be a close substitute for Lisbon or Istanbul if the purpose of the trip is to combine a city trip with a beach vacation, however it may be a closer substitute to Amsterdam or Prague as a destination for party travel. Vov97 also motivates overlapping nests in a model of transportation choice when commuters may combine various modes of transportation for different legs (“trunk", “egress") of the trip. Categorical outcomes may also result from choices among a richer set of heterogeneous alternatives - for example labor force status is determined by the decision whether to accept a particular job offer.
Standard random utility models (RUM) for discrete choice can be made more flexible to accommodate these richer substitution patterns, either by allowing for dependence among taste shocks, or modeling decisions sequentially via a richer latent hierarchy of nests of alternatives. Our modeling approach follows the second route but we also show that within McF78's Generalized Extreme Value (GEV) framework, the two are equivalent in terms of the distributions that a fully flexible model can generate.
We consider a model of multinomial choice $Y$ among $M$ options $\{y_1,\dots,y_M\}$ given agent and alternative specific attributes $\mathbf{x}$. The object of interest are the conditional choice probabilities \[\mu_0(y_m,\mathbf{x}):=\mathbb{P}(Y=y_m|\mathbf{X}=\mathbf{x}),\hspace{0.5cm}m=1,\dots,M\] which are assumed to result from a random utility model (RUM) with a flexible dependence structure among taste shocks. Specifically, we let \[U_{im}:=U_{im}^* + \varepsilon_{im},\hspace{0.3cm}m=1,\dots,M\] with systematic parts $U_{im}^*:=U^*(\mathbf{x}_{im})$ that are nonstochastic functions of agent/alternative specific characteristics, and idiosyncratic taste shocks $\varepsilon_{im}$ that are independent of $\mathbf{x}_{im}$ with a common marginal distribution $G(\varepsilon)$. The systematic part $U_{im}^*$ can itself be specified flexibly, e.g. as an additively separable function of elements of $\mathbf{x}_{im}$.
By Sklar's theorem, the joint distribution of taste shocks for the $M$ alternatives can be written as \[ G(\varepsilon_{1},\dots,\varepsilon_{M}) = C(G(\varepsilon_1),\dots,G(\varepsilon_{M}))\] where the copula $C:[0,1]^{M}\rightarrow[0,1]$ is a $M$-non-decreasing function that is onto, where $C(0,\dots,0)=0$ and $C(1,\dots,u_m,\dots,1)=u_m$ for each $k=1,\dots,M$. Taking logs of joint and marginal c.d.f.s, we can rewrite the copula as \[G(\varepsilon_{1},\dots,\varepsilon_{M}) = \exp\left\{-F_0\left(-\log G(\varepsilon_1),\dots,-\log G(\varepsilon_{M})\right)\right\}\] where we refer to the mapping $F_0:\mathbb{R}_+^{M}\rightarrow\mathbb{R}_+$, $F_0:(w_1,\dots,w_{M})\mapsto F_0(w_1,\dots,w_{M})$ as the generating function associated with the copula $C(\cdot)$.
For the case when the marginal distribution $G(\varepsilon)=\exp\{-e^{-\varepsilon}\}$ is extreme-value type I, Theorem 1 in McF78 shows that any conditional choice probabilities of the form
can be generated by such a RUM under additional regularity conditions on the generating function $F_0(\cdot)$. Specifically, he assumes the following
The alternating sign condition on partial derivatives is satisfied by any function $F_0(\cdot)$ that is $M$ times differentiable and $M$-nondecreasing (see Nel06, p.43 for a definition). Differentiability of degree greater than 1 is not needed for our results, whereas $M$-monotonicity is needed to ensure that the generating function defines a proper copula for the joint distribution of taste shocks under a random-utility interpretation of the conditional choice probabilities. In particular, these restrictions ensure that the mapping \[C(u_1,\dots,u_{M})=\exp\left\{-F_0(-\log u_1,\dots,-\log u_M)\right\}\] of marginal ranks $u_1,\dots,u_M$ to $[0,1]$ is $M$-increasing and therefore yields a well-defined joint distribution function. However, these restrictions do not guarantee that $C(u_1,\dots,u_{M})$ satisfies the boundary conditions $C(1,\dots,u_m,\dots,1)=u_m$, so the function need not be a proper copula. Rather, permissible generating functions under Assumption (ref) may produce joint distributions with marginals that are different from the extreme value type-I distribution. We discuss the possibility of allowing for flexible, smooth specifications of the marginal distributions $G_1(\varepsilon),\dots,G_M(\varepsilon)$ in Appendix (ref).
We now propose a strategy to approximate the generating function flexibly using a model of multiple stages of discrete decisions. Following Vov97 and WKo01 we assume a nesting structure with multiple layers where alternatives can be grouped according to various, not necessarily exclusive aspects, rather we allow for each alternative to belong to multiple nests at a time. Overlapping nests may occur in settings when a heterogeneous population of agents with different nesting structures - in an example by Vov97 on the choice of transportation modes, an agent may assign a “Park and Ride" commute either to a “private" or “public transit" nest, depending on which leg of the commute is more prominent.
We consider nested generating functions $F_k^{(s)}:\mathbb{R}_+^{K_{s-1}}\rightarrow\mathbb{R}_+$, where
where $\boldsymbol\theta_k^{(s)}:=(\beta_{k1}^{(s)},\dots,\beta_{kK_{s+1}}^{(s)},\varrho_k^{(s)})'$. The nesting coefficients $\varrho_k^{(s)}\geq1$ parameterize dependence of taste shocks within nests, where $\varrho_k^{(s)}=1$ corresponds to the case in which shocks are independent across nests, and $\varrho_k^{(s)}=\infty$ to the case of perfect rank dependence. The weights $\beta_{kl}^{(s)}$ can be interpreted as allocation parameters, measuring the relevance of nest $k$ for choosing alternative $l$. We then construct a generating function $F_{\mathbf{K}}^S(w_1,\dots,w_{M})$ recursively, where
We can interpret this structure as a cross-nested Logit model with $S$ layers, where in each layer $s$ the agent chooses among $K_s$ latent nests $k_s=1,\dots,K_s$, and we also assume that the $0$th (bottom) layer consists of a single root node, i.e. $K_0=1$. The resulting nesting tree is a weighted directed, acyclic graph where a terminal node may be reached through various paths.
The conditional choice probabilities $\mu(y_m,\mathbf{x})$ are then approximated by
in analogy to ((ref)). We allow for the possibility that $F_{\mathbf{K}}^S(\mathbf{w})$ is only directionally differentiable, in which case the partial derivative is taken to be the derivative from the right, $\frac{\partial}{\partial w_m}F_{\mathbf{K}}^S(w_1,\dots,w_m,\dots,w_M):= \lim_{t\downarrow0}\frac{F_{\mathbf{K}}^S(w_1,\dots,w_m+t,\dots,w_M)-F_{\mathbf{K}}^S(w_1,\dots,w_m,\dots,w_M)}{t}$.
We next show that with a sufficient number of latent nests, this model can in fact approximate the conditional choice probabilities resulting from the random utility model satisfying Assumption (ref) arbitrarily well.
See the appendix for a proof. For the restriction of the approximation to arguments $w_1,\dots,w_M$ in the interior of the positive orthant, notice that ((ref)) evaluates the generating function and its first partial derivatives only at arguments $w_m=\exp\{U_m^*(\mathbf{x})\}$ so that we can restrict our attention to compact subsets of $\textnormal{int}(\mathbb{R}_{+}^M)$ as long as $U_1^*(\mathbf{x}),\dots,U_{M}^*(\mathbf{x})$ are bounded functions of $\mathbf{x}$.
By ((ref)) the conditional choice probabilities can be expressed as continuous functions of $F_0(w_1,\dots,w_M)$ and its first partial derivatives. Our proof is constructive in that we propose a four-layer neural network which approximates the subgraph of $F_0(w_1,\dots,w_M)$ on $\mathcal{C}$ with vertices corresponding to the $K_2$ neurons in the second layer, where the free weights determine the location of those support points, and the remaining layers generate the surface of that polytope using parameters which depend on $F(\mathbf{w})$ only through those weights from the second layer. We then use known results on approximations of convex sets by polyhedra to bound the resulting error regarding the function $F_0(w_1,\dots,w_M)$ and its first partial derivatives.
It is also important to note that for this approximation, nest-specific weights $\beta_{kl}^{(s)}$ are non-negative, and nesting coefficients satisfy $\varrho_k^{(s)}\geq 1$, so that the approximating model is consistent with random utility maximization. One interpretation of the conventional nested Logit model represents the taste shifters for the alternatives in the final layer as a sum of independent alternative- and nest-specific random utility shocks (see BLe85 and Gal21). We can therefore think of the deep network approximating the “true" error distribution with a convolution of independent shocks some of which are shared by subsets of the $M$ alternatives.
To nest this model into the recursive framework in ((ref)) and ((ref)), we parameterize conditional choice probabilities from maximizing behavior in terms of inclusive values. Specifically we define the inclusive value for nest $k$ in the $s$th layer recursively via
where
Note that we can take limits \[\lim_{\varrho\rightarrow\infty} \left[\sum_{l=1}^{K}\left(\beta_{l}e^{v_l}\right)^{\varrho}\right]^{1/\varrho} =\max\left\{\beta_le^{v_1},\dots,\beta_Ke^{v_K}\right\}\] so that for large values of $\varrho_{k}^{(s)}$, we can interpret $\psi_k^{(s)}$ in ((ref)) as a softmax mapping from $e^{v_1^{(s+1)}},\dots, e^{v_{K_{s+1}}^{(s+1)}}$ to $e^{v_k^{(s)}}$.
We can next characterize the conditional choice probabilities from this random utility model recursively in terms of these inclusive values. Specifically, we define the probabilities $\pi_k^{(s)}$ of reaching the $k$th nest in the $s$th layer via the recursion
initialized at the root node. Lemma (ref) below shows that this mapping is given by
See the appendix for a proof. It is interesting to note that the recursion defining inclusive values and conditional choice probabilities is triangular in that the mapping iterating inclusive values ((ref)) does not take conditional choice probabilities $\pi_k^{(s)}$ as an argument, so that all hidden states can be calculated with a single backward pass followed by a forward pass of the recursions $\psi_k^{(s)}(\cdot)$ and $\phi_k^{(s)}(\cdot)$, respectively.
The main object of interest in the nested GEV model is the conditional choice probability \[\mu_0(y,\mathbf{x}):=\mathbb{P}(Y=y|\mathbf{X}=\mathbf{x})\] Given the trained network, the model prediction for $\mu_0(y,\mathbf{x})$ is then given by the activations \[\mu_{\mathbf{K}}^S(y,\mathbf{x};\boldsymbol\theta):=\left\{
\right.\] In order to train the network, we can therefore use the activations $\pi_1^{(S)}(\mathbf{x},\boldsymbol\theta),\dots,\pi_M^{(S)}(\mathbf{x},\boldsymbol\theta)$ in the top layer in a regression layer with loss $\ell(\mathbf{z},\boldsymbol\theta):=-\sum_{m=1}^M1\hspace{-2.5pt}\textnormal{l}\{y=y_m\}\log\pi_{m}^{(S)}(\mathbf{x},\boldsymbol\theta)$ and deeper layers specified according to ((ref)) and ((ref)).
Generally speaking, whether or not there is an advantage in approximating the relationship between inputs and outputs using a nested model which mimics this structure depends on whether there are any meaningful constraints on the nested mappings $\phi_k(\cdot)$ and $\psi_k(\cdot)$ in the generative model. Specifically, we consider qualitative restrictions on the intermediate transformation functions $\left\{F_k^{(s)}(w_1,\dots,w_{K_{s-1}})\right\}_{k,s}$ either in production or the construction of the GEV generating function, which are the main economic primitives of either problem.
Our general approach already makes use of global shape restrictions - monotonicity, convexity, homogeneity where applicable - on certain economic primitives. Generally speaking, qualitative restrictions on components of a structural model do not necessarily translate into analogous restrictions that could be imposed on the reduced form in a straightforward manner. However, for the flexible approximating models proposed in this paper we show that these shape constraints can be incorporated quite naturally into estimation as sign restrictions on model parameters and yield a substantial reduction of its statistical complexity.
In this section we consider two additional qualitative/nonparametric restrictions that formalize how economic primitives at individual stages of this model may be more “fundamental" than their composition.
In models of production, sparsity reflects that intermediate technologies may only use some of the available inputs for production. For nested models of discrete choice, sparsity restricts the cross-nesting structure by allowing any nest to have at most $d^*$ successors in the following nesting layer. Sparsity is assumed for the results in MPo16 and BKo19 with $d^*=2$, and general $d^*\leq K$, respectively.
Separability is satisfied by index models $F(\mathbf{v}'\boldsymbol\beta)$ for arbitrary link functions $F(\cdot)$ and the constant-elasticity of substitution (CES) family of production functions, including the cases of perfect complements and perfect substitutes. When the functions $g_1^{(s)}(\cdot),\dots,g_{K_{s-1}}^{(s)}(\cdot)$ are the identity, this corresponds to the nested nonparametric regression model analyzed in HMa07. In principle, this definition can be extended to include models satisfying separability among partitions of $w_1,\dots,w_{K_{s-1}}$ into subsets of multiple variables.
One key observation is that neither of these properties of link functions $F_k^{(s+1)},F_l^{(s)},\dots,F_{K_{s}}^{(s)}$ is generally inherited by compositions \[F_k^{(s,s+1)}(w_1,\dots,w_{K_{s-1}}):=F_k^{(s+1)}(F_1^{(s)}(w_1,\dots,w_{K_{s-1}}),\dots,F_{K_{s}}^{(s)}(w_1,\dots,w_{K_{s-1}}))\] Moreover, our main focus will be on cases in which the relevant intermediate state variables $v_k^{(s)},\pi_k^{(s)}$ are the result of optimization with an objective function $F_k^{(s)}(\cdot)$ satisfying sparsity or separability, whereas the activation functions $\psi_k^{(s)}(\cdot),\phi_k^{(s)}$ derived from these primitives need in general not exhibit either property. It is for these reasons that we aim to construct the approximating network to mirror the nested architecture of the data generating process when applicable.
We now propose network architectures which directly impose either sparsity or separability within components of an $S$-layer generative model with neurons of unspecified functional form. Within our framework, it is then possible to construct a common architecture which would nest all three cases - sparse, separable, and unrestricted - where either restriction can be imposed by setting weights for certain connections equal to zero. For simplicity, we state our results for estimation of production models, the case of discrete choice models is entirely analogous.
We first consider approximation of a model with stage technologies that are $d^*$-sparse according to Definition (ref) with dimension $d^*\leq K$, i.e. a production or generating function of the form \[F_{k}^{(s)}(w_1,\dots, w_{K_{s-1}})=f_k^{(s)}(w_{k_1},\dots w_{k_{d^*}}) \] Following the construction in BKo19, we propose the following approximating network:
We can now state the main result regarding the rate of approximation for a sparse stage production technology given that network architecture:
See the appendix for a proof. Comparing this result to the rates in Proposition (ref), all rates are in terms of the effective dimensionality $d^*$ under sparsity rather than $K$. An analogous result can be given for the discrete choice model under Assumption (ref) when the generating function is $d^*$-sparse.
We next consider separability restrictions in the production model. For models of discrete decisions, these restrictions can be imposed in an entirely analogous manner. In order to approximate an intermediate production technology that is separable in the sense of Definition (ref), i.e. satisfying \[F_k^{(s)}(w_1,\dots,w_{K_{s-1}}) = f_k^{(s)}(g_{k1}^{(s)}(w_1)+ \dots + g_{kK_{s-1}}^{(s)}(w_{K_{s-1}}))\] we propose the following Four-layer CES (4L-CES) specification of the production network uses three layers of nested CES production functions to approximate $F_k^{(s)}$.
Using this construction we can achieve the following approximation rates for separable stage technologies:
See the appendix for a proof.
This section gives convergence rates for estimating the function $\mu_0(\mathbf{y},\mathbf{x})$, where we approximate the underlying model primitives - the feasible set $\mathcal{Y}_0$ or the generating function $F_0(\mathbf{w})$, respectively - with a multilayer perceptron. We consider the class $\mathcal{H}_{\mathbf{K}}^S$ of multilayer neural networks with $S$ layers, $\mathbf{K}=(K_1,\dots,K_S)$ hidden nodes, and $W$ free parameters (weights $\beta_k^{(s)}$ and coefficients $\varrho_k^{(s)}$). We denote the resulting approximation to the target function with $\mu_{\mathbf{K}}^S(\mathbf{y},\mathbf{x})$. Note that in implementing this approach we do not explicitly compute and report the full set $\mathcal{Y}_0$ or generating function $F_0(\mathbf{w})$, rather the trained network only evaluates whether any given point in the sample belongs to that set. Global shape restrictions on that set then only restrict complexity of ways of assigning points in and outside of that set.
Estimation of $\mu_0$ is based on a sample of $n$ observations, where we assume the following:
The assumption of a rectangular support for $X$ is primarily for analytical convenience and can be generalized to other compact subsets of $\mathbb{R}^K$ with nonempty interior. The condition that we observe input/output combinations over the entire efficient frontier of $\mathcal{Y}_0$ requires additional assumption on the mechanism for selection among multiple efficient input/output combinations, e.g. due to sufficient variation in input and output prices for a profit-maximizing producer. A more rigorous formulation in terms of economic primitives is beyond the scope of this paper and will be left for future research. If that condition fails, the researcher may still learn about a segment of the efficient frontier determining $\mu_0(\mathbf{u},\mathbf{x})$ for certain directions $\mathbf{u}$ and a restricted set of input combinations $\mathbf{x}$.
Our asymptotic results then concern the estimate of $\mu_0(\mathbf{y},\mathbf{x})$ based on that sample. Following FLM19, suppose that we train the network according to a loss function $\ell(\mu,\mathbf{z})$ such that \[\mu_0=\operatorname*{arg\,min}_{\mu}\mathbb{E}[\ell(\mu,\mathbf{Z})].\] In addition we assume that $\ell(\cdot,\cdot)$ is Lipschitz and that there exist $c_1,c_2>0$ such that \[c_1\mathbb{E}\left[(\mu-\mu_0)^2\right]\leq\mathbb{E}\left[\ell(\mu,\mathbf{Z})\right]-\mathbb{E}\left[\ell(\mu_0,\mathbf{Z})\right] \leq c_2\mathbb{E}\left[(\mu-\mu_0)^2\right]\] After training the network $\mathcal{H}_{\mathbf{K}}^S$ we obtain the estimator
To characterize the asymptotic properties of $\hat{\mu}_n$ we start by stating a general VC bound of the approximating neural network which applies to general convex approximating architectures.
See the appendix for a proof. The proof of this result applies the generic bound in Theorem 2 by KMa97, noting that for any concave function all level and contour sets are connected, so that in the notation of their paper $B=1$. Furthermore, it is important to note that their results apply to Boolean combinations of “atomic formulas" which are arbitrary, infinitely differentiable functions of the data and parameters. In particular, the theorem does not require the $\mu_{\mathbf{K}}^S(\mathbf{y},\mathbf{x})$ to be defined in terms of an acyclic (feedforward) multilinear perceptron.
Other restrictions, like sparsity or separability enter through the architecture of the neural network, where the constrained network requires only a smaller number $W$ of adjustable weights. Since Lemma (ref) depends only on $W$ with no additional assumptions on network architecture, VC bounds for shape constrained networks can be obtained from that same result.
Most notably the bound in Lemma (ref) implies no penalty for network depth besides through the number of adjustable parameters, whereas all of the generic results (allowing for negative coefficients) do. The main reason for this is that without convexity or monotonicity, lower contour sets of the functions a neural network may generate consist of multiple connected components, where the number of connected components increases with network depth.
Given the previous results regarding the rate of approximation and VC dimension of these networks, we can adapt the proofs for the main results in FLM19 to obtain the asymptotic rate for estimation of $\mu_0(\cdot)$. Our approximation results in Propositions (ref) and (ref) and Lemma (ref) will simply take the place of Lemmas 6 and 7 in their argument, which is otherwise not specific to the case of ReLU feedforward networks.
We state the result for the smallest number of hidden layers required for the approximation in (ref) and (ref), noticing that neither result suggests an added benefit to depth for the case in the absence of additional shape restrictions. However, that conclusion changes once we assume that stage technologies are sparse or separable, and we derive convergence rates for that case separately below. Note also that the rate for $\hat{\mu}_n$ matches the minimax risk bound for nonparametric estimation of convex functions in Theorem 2.4 by HWe16, up to logarithmic terms. Interestingly, they also point out that the default nonparametric estimator, bounded least squares, fails to achieve that bound for $d>4$.
As discussed in the context of Theorem 3 in FLM19, the approximation bounds underlying this result can also be strengthened in order to establish asymptotic normality and variance estimation for certain functionals of $\mu_0(\cdot)$. Since apart from the separate derivation of the VC dimension of $\mathcal{H}_{\mathbf{K}}^S$ and its approximating properties the argument is completely analogous to their case we do not prove the analogous conclusions separately for the present framework.
In order to appreciate the gains from imposing additional shape restrictions, we next state the convergence rates under sparsity and separability restrictions for the stage functions in an $S$-layer model with neurons of unknown functional form.
This result follows immediately from the proof of Theorem (ref), noting that $Q$ can be chosen according to the approximation rate in Proposition (ref) and that the resulting approximating network has less than $3Sd^*Q$ free parameters.
As for the $d^*$-sparse case, this result follows again from the proof of Theorem (ref), where $Q$ can be chosen according to the approximation rate in Proposition (ref). The resulting approximating network has less than $(K+5)SQ$ free parameters, so that we obtain the conclusion by applying that rate to the generic VC bound in Lemma (ref).
This paper explores the use of artificial neural networks to approximate the reduced form of models of production and discrete decisions. For one we illustrate the expressive capacity of nested models of optimizing behavior, where in the absence of additional restrictions on the number of hidden units, such a model can generate any reduced form within a nonparametric class satisfying only broad shape constraints. Conversely that approximating property can be used for estimation, where nested models can be used to approximate an otherwise unconstrained reduced form. Here our results imply that the reduced form can be estimated at an asymptotic rate corresponding to the optimal nonparametric rate for estimating regression functions with bounded partial derivatives up to order two. Furthermore, monotonicity and convexity can be imposed as simple sign restrictions on model parameter.
One important benefit of using an approximating function class that mirrors a structural model for the data generating process is that additional sparsity or separability restrictions can be imposed directly on the network architecture, resulting in a faster rate of convergence for the estimator. This raises the question whether the network architecture can be made adaptive to the unknown structure of the underlying model, e.g. using $L^{1}$-penalization of weights, which will be left for future research.