Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
114,526 characters · 11 sections · 101 citation commands
Ensemble Learning with Statistical and Structural Models
\pagenumbering{gobble}
\pagenumbering{arabic}
In economics as well as many other scientific disciplines, statistical and structural modeling represent two distinct approaches to data analysis heckman_causal_2000. The structural approach draws a direct link between data and theory. It estimates structural models, or scientific models shalizi_advanced_2013, that specify the causal mechanisms generating the observed data. A complete structural model in economics describes economic and social phenomena as the outcomes of individual behavior in specific economic and social environments heckman_econometric_2007,reiss_structural_2007. Once estimated, these models can be used for making predictions, evaluating causal effects, and conducting normative welfare analyses low_use_2017.
In contrast to the structural approach, the statistical approach to data analysis relies on the use of statistical models for prediction and causal inference. While recent advances in machine learning have focused on predictive tasks athey_beyond_2017, a large literature in causal inference across multiple disciplines\footnote{e.g. the social sciences, the biomedical sciences, statistics, and computer science.} has proposed statistical methods for estimating causal effects from experimental and observational data imbens_causal_2015\footnote{In the statistical approach to causal inference, causal knowledge is used not to specify a complete structural model, but to inform research designs that can identify the causal effects of interest by exploiting exogenous variations in the data.}. In economics, this statistical approach to causal inference is informally referred to as the reduced-form approach chetty_sufficient_2009\footnote{As chetty_sufficient_2009 pointed out, the term “reduced-form” is largely a misnomer, whose meaning in the econometrics literature today has departed from its historical root. Historically, a reduced-form model is an alternative representation of a structural model. Given a structural model $\mathcal{M}(x,y,\epsilon)=0$, where $x$ is exogenous, $y$ is endogenous, and $\epsilon$ is unobserved, if we write $y$ as a function of $x$ and $\epsilon$, $y=f(x,\epsilon)$, then $f$ is the reduced-form of $\mathcal{M}$ reiss_structural_2007. Today, however, applied economists typically refer to nonstructural, statistical treatment effect models as “reduced-form” models. Perhaps reflecting the informal nature of the terminology today, rust_limits_2014 gave the following definitions of the two approaches: “At the risk of oversimplifying, empirical work that takes theory \textquotedblleft seriously\textquotedblright is referred to as structural econometrics whereas empirical work that avoids a tight integration of theory and empirical work is referred to as reduced form econometrics.”}. Methods such as controlling for observed confounding and instrumental variables regression are widely used in applied economic analyses athey_state_2017.
Which approach should be preferred -- the statistical or the structural -- has been the subject of a long-standing debate within the economics profession angrist_credibility_2010,deaton_instruments_2010,keane_structural_2010,keane_structural_2010-1,nevo_taking_2010,wolpin_limits_2013. For predictive tasks, statistical and machine learning models often fit the observed data well and have advantages in in-domain prediction, where the training and the test data have the same distribution\footnote{Using the terminology of transfer learning, a domain is a joint distribution governing the input and output variables muandet_domain_2013. A key limitation with most statistical and machine learning models is that they require the distributions governing the training data (the source domain) and the test data (the target domain) to be the same in order to guarantee performance ben-david_theory_2010. }. On the other hand, a main advantage of structural estimation lies in its ability to make out-of-domain predictions\footnote{In this paper, we distinguish between the notions of out-of-domain and \emph{out-of-sample}. Out-of-sample data are test data drawn from the same distribution as the training data.}. As long as the same causal mechanism governs data generation, a correctly specified structural model provides a way to extrapolate from the training data to the test data even if the distributions have changed\footnote{Traditionally, economists emphasize the ability of structural models to make \emph{counterfactual} predictions. We note that counterfactual predictions can be viewed as a special type of out-of-domain predictions.}. Similarly, in causal inference, reduced-form methods that exploit credible sources of identifying information deliver estimates of causal effects with high \emph{internal validity}\footnote{angrist_credibility_2010 offered an account of what they call “the credibility revolution” -- the increasing popularity of quasi-experimental methods that seek natural experiments as sources of identifying information. Our definition of reduced-form methods include both quasi-experimental and more traditional, non-quasi-experimental statistical methods that use expert knowledge to locate exogenous sources of variation.}, while structural estimates may have more claims to \emph{external validity}.
The relative strengths of the two approaches point to a complementarity that provides the motivation for this paper. Of course, the reason that any approach may outperform the other in certain aspects of data analysis is fundamentally due to model misspecification\footnote{By model misspecification, we refer to both incorrect functional form and distributional assumptions and, in the case of causal inference, incorrect causal assumptions.} -- if any model captures the true distributions governing the source and the target domains, then no improvement is possible. Indeed, one can argue that researchers on both sides of the methodological debate are motivated by a shared concern over model misspecification. Proponents for the statistical approach are concerned about misspecifications due to the often strong and unrealistic assumptions -- both causal and parametric -- made in structural models, while those advocating for the structural approach are concerned about misspecifications due to not incorporating theoretical insight -- functional forms such as constant elasticity of substitution (CES) aggregation and the gravity equation of trade, for example, often encode important prior economic knowledge that sophisticated statistical and machine learning methods would not be able to capture based on training data alone\footnote{rust_limits_2014: \textquotedblleft Notice the huge difference in world views. The primary concern of Leamer, Manski, Pischke, and Angrist is that we rely too much on assumptions that could be wrong, and which could result in incorrect empirical conclusions and policy decisions. Wolpin argues that assumptions and models could be right, or at least they may provide reasonable first approximations to reality.\textquotedblright}.
In this paper, we propose a set of methods for combining the statistical and structural approaches for improved prediction and causal inference. Our first proposed estimator, which we call the doubly robust statistical-structural (DRSS) estimator, provides a consistent in-domain estimate as long as either the structural or the (reduced-form) statistical model is correctly specified. Our second proposed estimator, which we call the ensemble statistical-structural (ESS) estimator, is a weighted ensemble that has the ability to outperform both the structural and the (reduced-form) statistical model, both in-domain and out-of-domain, when both are misspecified.
Our methods build on several intuitions. First, statistically speaking, a structural model is a generative model jebara_machine_2012. Given a structural model that specifies the data-generating mechanism of $\left(x_{1},\ldots,x_{p}\right)\in\mathcal{O}$, we can generate predictions of discriminative relationships $\mathbb{E}\left[\left.x_{j}\right|x_{i}\right]$ or $\mathbb{E}\left[x_{j}^{x_{i}=a}\right]$ for any $\left(x_{i},x_{j}\right)\subset\left(x_{1},\ldots,x_{p}\right)$, where $x_{j}^{x_{i}=a}$ denotes the potential outcome of $x_{j}$ under the intervention of $x_{i}=a$\footnote{In this paper, we mainly adopt the notations of the Rubin causal model rubin_estimating_1974 in discussing causal inference. Equivalently, using the notation of pearl_causality_2009, $\mathbb{E}\left[x_{j}^{x_{i}=a}\right]$ can be expressed as $\mathbb{E}\left[\left.x_{j}\right|\text{do}\left(x_{i}=a\right)\right]$. }. These structurally derived relationships can then be considered as competitors to (reduced-form) statistical models that explicitly model these relationships. This allows us to leverage the large statistical literature on dealing with competing models. One popular method used in causal inference is the doubly robust estimator that combines an outcome regression model with a treatment assignment model in the estimation of causal effects bang_doubly_2005. The doubly robust estimator is consistent if either of the two models is correctly specified, thus providing an insurance against model misspecification. lewbel_general_2019 generalized the classic doubly robust method to allow the combination of any parametric models. Their method provides a basis for our DRSS estimator.
Second, the complementary properties of statistical and structural models suggest that a model combination approach may yield superior results kellogg_combining_2020. In the Bayesian paradigm, model averaging has long been proposed as an alternative to model selection hoeting_bayesian_1999. Given a set of candidate models, bayesian model averaging produces a weighted average, with each model weighted by its posterior probability. Doing so accounts for the model uncertainty that is ignored by the standard practice of selecting a single model. More recently, in the machine learning literature, ensemble methods such as stacking, bagging, and boosting are proposed that seek to combine models to improve prediction so that the ensemble performs better than any of its individual members dietterich_ensemble_2000. These methods work by not only incorporating model uncertainty but expanding the space of representable functions minka_bayesian_2000\footnote{When the models being combined are complex and high-dimensional for which global optima are hard to obtain, the ensemble approach also produces gains by averaging local optima produced by local search dietterich_ensemble_2000.}. As breiman_stacked_1996 pointed out, ensemble methods benefit the most from the use of diverse and dissimilar models, which is exactly the case when we combine statistical and structural models.
In this paper, we provide two ensemble estimators. The first, which we call ESS-LN, is a linear ensemble based on the method of stacking wolpert_stacked_1992, or jackknife averaging hansen_jackknife_2012, which produces an optimal linear combination of a set of models by minimizing a cross-validated loss criterion such as expected mean squared error. We show how to use the method both for prediction and causal inference. Our second ensemble estimator, ESS-NP, goes beyond linear combinations and builds a nonparametric ensemble of statistical and structural models. For conditional mean estimation, it employs the random forest algorithm introduced by breiman_random_2001, which allows for the modeling of nonlinear relationships and complex interactions by building a large number of regression trees that adaptively partition the input space and combining them through bootstrap aggregation. The method can be viewed as an adaptive locally weighted estimator athey_generalized_2019, allowing us to assign different weights to different regions of the input space depending on which model -- the statistical or the structural -- performs better in that region. The resulting ensemble has the ability to combine the strengths of statistical and structural models while defending against their weaknesses.
\paragraph*{Example}
To illustrate our methods, consider the setting of a simple demand estimation problem. We observe the prices and quantities sold of a good $x$, as plotted in Figure (ref). Suppose the data are generated by the consumption decisions of $n$ consumers who purchased $x$ at different prices. Each consumer had fixed income $I$ and decided how much to purchase by solving the problem:
, where $\left(p_{i},q_{i},p_{i}^{o},q_{i}^{o}\right)$ denote respectively the price and quantity of good $x$ and of an outside good $o$. The consumer utility function is given by the following CES function:
, where $\rho=\frac{1}{2}$, suggesting an elasticity of substitution of $2$\footnote{$\left\{ \alpha_{i}\right\} $ are generated as follows: \[ \alpha_{i}=\frac{\exp\left(\xi_{i}\right)}{1+\exp\left(\xi_{i}\right)},\quad\xi_{i}\sim\mathcal{N}\left(0,0.5\right) \] }.
We can fit the following statistical model to the data:
The result is plotted in Figure (ref). Under the causal assumption that prices are exogenous to the consumers, (ref) represents a reduced-form estimate of the individual demand curve. The model appears to fit the data quite well. However, once we extrapolate beyond the observed ranges of prices, its predictions become very bad (Figure (ref)). On the other hand, structurally estimating the parameters of model (ref) would yield a demand curve that has both internal and external validity (Figure (ref)). This is not surprising as (ref) describes the true data-generating mechanism. In practice, given two competing models, the (reduced-form) statistical model (ref) and the structural model (ref), we may not know which one is correctly specified. The DRSS resolves this issue by combining the two models and providing a consistent estimate as long as one of them is correctly specified. Figure (ref) plots the DRSS fit. In this case, the DRSS estimator is able to “pick up” the right model and hews closely to the true structural fit.
In reality, of course, most often all our models are misspecified. In Figure (ref), we plot the results of estimating model (ref) but assuming $\rho=-\frac{1}{2}$\footnote{That is, instead of estimating both $\left(\alpha_{i},\rho\right)$ from the data, we estimate $\alpha_{i}$ only while treating $\rho=-0.5$ as an assumption of the model. The assumption, of course, is incorrect in this case.}. The resulting structural fit now deviates pronouncedly from the true model, highlighting the fact that the validity of the structural approach hinges crucially on the model being correct. The DRSS estimator that combines this misspecified structural model with the (reduced-form) statistical model (ref) now puts most of its weight on the latter and is no longer consistent (Figure (ref)). Note, however, compared to (ref), the misspecified structural model has worse fit in-domain, but still performs significantly better out-of-domain. This provides the motivation for our ensemble approach. Intuitively, although we misspecify the utility function, the theory of consumer utility maximization subject to budget constraints still provides important prior information on the likely shape of the demand curve -- such as its downward-slopingness -- that can be used to regulate the behavior of statistical models. In Figure (ref), we show the results of our ESS-NP estimator based on a random forest ensemble of the misspecified structural model and the (reduced-form) statistical model. The ESS-NP fit is closer to the true model and performs well both in-domain and out-of-domain. Thus in this example, the ensemble approach\footnote{The ESS-LN method produces similar results as the ESS-NP in this example.} is able to deliver optimal performance when both the structural and the (reduced-form) statistical models are incorrect.
In section (ref), we demonstrate the effectiveness of our methods using a set of simulation experiments under a variety of more realistic settings in applied economic analyses, including first-price auctions and dynamic models of entry and exit. We also revisit this demand estimation problem and show how to apply our methods to estimating the demand curve with the help of instrumental variables when prices are endogenous. For each experiment, we report the performance of the DRSS and ESS estimators when either or both of a structural model and a (reduced-form) statistical model is misspecified.
\paragraph*{Related Literature}
This paper is related to several strands of literature. The doubly robust estimator was proposed by robins_estimation_1994,robins_semiparametric_1995,scharfstein_adjusting_1999 as a means of estimating the average treatment effect by combining an outcome regression model with a treatment assignment model so that the estimator remains consistent as long as one of the models is correctly specified. In general, an estimator is said to have the doubly robustness property if it is consistent for the target parameter when any one of two nuisance parameters is consistently estimated benkeser_doubly_2017. Subsequent developments in doubly robust estimation include bang_doubly_2005,tan_bounded_2010,okui_doubly_2012,farrell_robust_2015,vermeulen_bias-reduced_2015,benkeser_doubly_2017,arkhangelsky_double-robust_2019. chernozhukov_double_2016,chernozhukov_doubledebiasedneyman_2017 showed that the doubly robust estimator can be viewed as being based on Neyman-orthogonal moment conditions that are first-order robust to errors in nuisance parameter estimation. More recently, lewbel_general_2019 proposed the general doubly robust (GDR) method that provides a general technique for constructing a doubly robust combination out of any parametric models, which forms the basis of our DRSS estimator.
Our paper is also related to the literature on model averaging and ensemble methods. Model averaging provides a natural response to model uncertainty in the Bayesian framework and has long been considered an alternative to model selection. See hoeting_bayesian_1999 for a comprehensive review of bayesian model averaging methods. In machine learning, wolpert_stacked_1992 proposed the method of stacking, or stacked generalization\footnote{Also see breiman_stacked_1996. When weights are restricted under a simplex constraint, stacking can be considered a frequentist model averaging technique. van_der_laan_super_2007 and hansen_jackknife_2012 provided theory on its asymptotic optimality. These authors also gave different names to the method: super learning van_der_laan_super_2007 and jackknife model averaging hansen_jackknife_2012.}. breiman_bagging_1996 proposed bagging, or bootstrap aggregation. freund_experiments_1996 introduced boosting. These ensemble methods are constructed with the explicit goal of maximizing predictive accuracy and achieve their effectiveness by incorporating model uncertainty, averaging local optima, and enriching the model space dietterich_ensemble_2000. More recently, there has also been a growing body of research in the statistics and econometrics literature on asymptotically optimal frequentist model averaging. See claeskens_focused_2003,hjort_frequentist_2003,hansen_least_2007,hansen_jackknife_2012,kitagawa_model_2016,zhang_optimal_2016,ando_weight-relaxed_2017. moral-benito_model_2015,steel_model_2019 provided overviews of the use of model averaging in economics.
Both the DRSS and the ESS estimators can be used to improve out-of-domain statistical predictions relative to a pure statistical approach. Our paper thus makes a contribution to the literature on transfer learning, which studies the problem of applying a model trained on a source domain to a target domain where the data-generating distribution may have changed\footnote{The problem of transfer learning is closely related to the problem of sampling bias or the sample selection problem -- a general problem that arises when we try to make inference, whether statistical or causal, about a population using data collected from another population.}. See pan_survey_2010 for a survey on transfer learning and ben-david_theory_2010 for theory on learning from different domains. A majority of research on transfer learning so far has focused on domain adaptation, where the marginal distributions of the input variables vary across domains and are observed, but the conditional outcome distribution is assumed to be the same. Methods that have been proposed aim to reduce the difference in input distributions either by sample-reweighting zadrozny_learning_2004,huang_correcting_2007,jiang_instance_2007,sugiyama_direct_2008 or by finding a domain-invariant transformation pan_domain_2010,gopalan_domain_2011\footnote{This includes the more recent deep domain adaptation literature that employs deep neural networks for domain adaptation. See glorot_domain_2011,chopra_dlid_2013,ganin_unsupervised_2014,tzeng_deep_2014,long_learning_2015. wang_deep_2018 provides an overview of this literature in the context of computer vision.}. Our methods, however, can be viewed as tackling the more difficult problem of domain generalization, where the target domain is unknown at the time of training and where both the marginal and the conditional distributions are allowed to vary. Intuitively, we achieve this by incorporating theory into statistical modeling\footnote{Transfer learning has also been referred to knowledge transfer pan_survey_2010. We note, however, that true knowledge transfer must involve causal knowledge as encapsulated in theory.}. The effectiveness of our approach hinges on the stability of the underlying causal mechanism and on the availability of a structural model that is informative, if not correctly specified\footnote{rojas-carulla_invariant_2018,kuang_stable_2020 also proposed methods for domain generalization by assuming stability in causal relationships. Both studies rely on the assumption that a subset of the input variables $v\subseteq x$ have a causal relation with the outcome $y$ and the conditional probability $p\left(y|v\right)$ is invariant across domains. However, it is \emph{not} true that having a causal relationship implies $p\left(y|v\right)$ is domain-invariant. Let $w=x\backslash v$. The assumption only holds under very limited and untestable conditions, namely that $y\perp w|v$ and that the causal effect of $v$ on $y$ is homogeneous.}.
A main contribution of this paper is to the literature on combining structural and reduced-form estimation. Many authors in economics have called for combining these two approaches to harness their respective strengths\footnote{chetty_sufficient_2009: “The structural and statistic methods can be combined to address the short-comings of each strategy ... By combining the two methods in this manner, researchers can pick a point in the interior of the continuum between reduced-form and structural estimation, without being pinned to one endpoint or the other.”}$^{,}$\footnote{Mirroring the debate in economics on structural vs. reduced-form estimation, there has long been a debate in the machine learning literature on generative vs. discriminative models as well as efforts to combine them. See ng_discriminative_2002,bishop_generative_2007. }. Early efforts include chetty_sufficient_2009,heckman_building_2010. Their solution is to use structural models to derive sufficient statistics for the intended analysis and then use reduced-form methods to estimate them. In comparison, we offer a set of general algorithms rather than relying on ad hoc derivations\footnote{However, our method cannot be used to conduct welfare analysis, which is the focus of chetty_sufficient_2009.}. More recently, fessler_how_2019,mao_structural_2020 proposed shrinkage methods that combine statistical and structural models by shrinking the former toward the latter. Their methods can be viewed as complementary to ours. Indeed, there is a connection between shrinkage and model averaging hansen_least_2007. By combining models of different complexities, a model averaging procedure effectively shrinks the more complex models toward the less complex ones.
Compared to fessler_how_2019,mao_structural_2020, our approach arguably also has several advantages. First, their methods are asymmetric with respect to the complexities of statistical and structural models. Specifically, they require the specification of complex statistical models to be regularized with structural models. In contrast, our approach is symmetric, allowing researchers to combine structural models with simple linear reduced-form models frequently used in applied research. Second, when the structural models are complex and high-dimensional, our ensemble methods can provide effective regularization. This can be most easily seen in the case of the stacking estimator ESS-LN. When the structural model is more complex than the statistical model, the ESS-LN effectively regularizes the former with the latter by averaging the two. This is relevant since many structural models used in empirical applications today are highly complicated and prone to overfitting as researchers strive for ever more “realistic” models\footnote{Importantly, the best model to describe a given data set may not be the model that truthfully describes the data-generating mechanism. This is because the true model may well be too complex for the amount of the data we have, in which case the model will be poorly fit on the limited sample and generate unreliable predictions. We therefore echo hansen_method_2015: “it remains an important challenge for econometricians to devise methods for infusing empirical credibility into \textquoteleft highly stylized\textquoteright models of dynamical economic systems. Dismissing this problem through advocating only the analysis of more complicated \textquoteleft empirically realistic\textquoteright models will likely leave econometrics and statistics on the periphery of important applied research.”}.
The rest of this paper is organized as follows. Section (ref) lays out the details of our algorithm. In section (ref) we apply our method to three sets of simulation experiments in the settings of first-price auctions, dynamic models of entry and exit, and demand estimation with instrumental variables and report their results. Section (ref) concludes.
The DRSS builds on the GDR method of lewbel_general_2019. In this section, we discuss the estimator first in the context of statistical prediction and then in causal inference. In both contexts, we first assume that we have access to a representative data set, i.e. the target domain on which we wish to make inference is the same as the source domain from which the data are drawn. We then consider the case that our data is non-representative and discuss its implications on the external validity or out-of-domain performance of our algorithms.
\paragraph*{Statistical Prediction}
Given variables $\left(x,y\right)\in\mathcal{X\times\mathbb{R}}$, assume first that our goal is to learn the conditional expectation function $\mu\left(x\right)=\mathbb{E}\left[y|x\right]$. We have at our disposal two parametric models for $\mu\left(x\right)$: $h\left(x;\theta_{h}\right)$ and $g\left(x;\theta_{g}\right)$, where $\theta_{h}\in\mathbb{R}^{p_{h}},\theta_{g}\in\mathbb{R}^{p_{g}}$. One of these models is correctly specified, but we do not know which one. Let $f\in\left\{ h,g\right\} $ index the correct model. Suppose the true parameter $\theta_{f}^{0}$ is identified by a set of $\ell_{f}\times1,\ \ell_{f}>p_{f}$ moment conditions $\mathbb{E}\left[\psi_{f}\left(x,y;\theta_{f}^{0}\right)\right]=0$. Given a sample of $n$ i.i.d. observations, we can then construct the following (adjusted) moment distance functions:
, where $\overline{\psi}_{m}\left(\theta_{m}\right)\doteq\frac{1}{n}\sum_{i=1}^{n}\psi_{m}\left(x_{i},y_{i};\theta_{m}\right)$, $\Omega_{m}$ is a $\ell_{m}\times\ell_{m}$ positive definite weight matrix\footnote{lewbel_general_2019 recommend the use of $\Omega=\widehat{\mathbb{E}}\left[\psi\left(\theta_{0}\right)\psi\left(\theta_{0}\right)'\right]^{-1}$, the (estimated) efficient GMM weight of hansen_large_1982. However, it may not be the optimal weight for the GDR or for our DRSS. We leave the characterization of the optimal weight matrix to future work. }, and $\kappa_{m}=\ell_{m}-p_{m}$ is the degrees of freedom of the $\chi^{2}$ statistic that the unadjusted $\text{\ensuremath{Q_{m}}}$ equals if $m$ is the true model.
Let $\widehat{\theta}_{m}=\underset{\theta_{m}}{\arg\min}\ Q_{m}\left(\theta_{m}\right),\ m\in\left\{ h,g\right\} $. A doubly robust estimator for $\mu\left(x\right)$ can be constructed as follows:
, where
Under regularity conditions, as long as one of the two models, $h$ or $g$, is correctly specified, it can be shown that $\widehat{\mu}\left(x\right)\rightarrow^{p}\mu\left(x\right)$. The proof is based on Theorem 1 of lewbel_general_2019 (see Appendix \hyperlink{APP}{A.1}). The intuition is simple: if one of the models, say $h$, is correctly specified but $g$ is not, then $Q_{h}\left(\widehat{\theta}_{h}\right)\rightarrow^{p}0$ while $Q_{g}\left(\widehat{\theta}_{g}\right)$ will have a nonzero limit. Thus in the limit, $w_{h}$ will be $1$ and $\widehat{\mu}\left(x\right)$ becomes $h\left(x;\widehat{\theta}_{h}\right)$ -- the consistently estimated correct model for $\mu\left(x\right)$.
Adapting the doubly robust estimator (ref) to combining statistical and structural models is straightforward: let $\mathcal{M}\left(x,y;\theta_{\mathcal{M}}\right)$ be a structural model that specifies the data-generating mechanism of $\left(x,y\right)$. From this generative structural model, we can derive its prediction of the discriminative function $\mu\left(x\right)$. Let $g\left(x;\theta_{\mathcal{M}}\right)=\mathbb{E}^{\mathcal{M}}\left[y|x\right]$ be the implied conditional mean of $y$ according to $\mathcal{M}$. We can then combine $g\left(x;\theta_{\mathcal{M}}\right)$ with any statistical model $h\left(x;\theta_{h}\right)$ according to (ref). The resulting estimator is the DRSS estimator for $\mu\left(x\right)$.
In practice, there are two ways to construct $\psi_{g}\left(x,y;\theta_{\mathcal{M}}\right)$ for the structurally derived discriminative model $g\left(x;\theta_{\mathcal{M}}\right)$. If $\mathcal{M}$ is the true model and $\theta_{\mathcal{M}}^{0}$ is the true parameter value, $\psi_{g}$ needs to satisfy $\mathbb{E}\left[\psi_{g}\left(x,y;\theta_{\mathcal{M}}^{0}\right)\right]=0$. Therefore, we can either directly specify a set of moment conditions that identify $\mathcal{M}$ or let $\psi_{g}\left(x,y;\theta_{\mathcal{M}}\right)=\phi\left(x\right)\left(y-g\left(x;\theta_{\mathcal{M}}\right)\right)$ for any function $\phi\left(.\right)$. We can then construct $Q_{g}\left(\theta_{\mathcal{M}}\right)$ based on $\text{\ensuremath{\psi_{g}\left(x,y;\theta_{\mathcal{M}}\right)}}$ and compute $\left(w_{h},w_{g}\right)$ based on $\text{\ensuremath{\left(Q_{h}\left(\widehat{\theta}_{h}\right),Q_{g}\left(\widehat{\theta}_{\mathcal{M}}\right)\right)}}$, where $\left(\widehat{\theta}_{h},\widehat{\theta}_{\mathcal{M}}\right)$ are obtained from separate first stage estimation of the statistical model $h$ and the structural model $\mathcal{M}$.
\paragraph{Sample Splitting}
The DRSS method as outlined above is a two-stage procedure, where $\left(\widehat{\theta}_{h},\widehat{\theta}_{\mathcal{M}}\right)$ are obtained in a first stage and the estimator is constructed according to (ref) in a second stage. If both stages are conducted on the same sample of data, however, finite sample bias from the first stage will be carried over to the second stage, especially when complex statistical or structural models, prone to overfitting, are estimated in the first stage. To avoid bias from overfitting and ensure good statistical behavior, we can use separate data sets for the two stages of the procedure. This can be accomplished by, for example, splitting the observed data randomly into two parts. This is known as sample-splitting angrist_split-sample_1995\footnote{The idea of sample-splitting is of course closely related to the idea of using separate training and validation data sets for fitting model- and hyper-parameters in machine learning. Indeed, the weights $\left(w_{h},w_{g}\right)$ can be viewed as the hyperparameters of the DRSS model.}. This way, from the perspective of the second stage, $\left(\widehat{\theta}_{h},\widehat{\theta}_{\mathcal{M}}\right)$ are exogenously given, so that when we evaluate the moment distance functions $Q_{h}$ and $Q_{g}$ -- critical for computing the DRSS weights -- we do not suffer an optimistic bias due to $\left(\widehat{\theta}_{h},\widehat{\theta}_{\mathcal{M}}\right)$ being obtained from the same data.
There is an efficiency cost involved in sample-splitting, as half of the data are wasted in each stage. The results can also be highly variable due to the whims of a single random split. To improve efficiency, we can perform sample-splitting multiple times and average their results. This is the idea behind cross-validation and cross-fitting chernozhukov_double_2016,chernozhukov_doubledebiasedneyman_2017 and can be described as follows for our DRSS estimator: randomly partition the data into $K$ equal-sized parts. For $k=1,\cdots,K$, let $\mathcal{D}_{k}$ denote the data of the $k$th partition and let $\mathcal{D}_{-k}$ denote the data not in $\mathcal{D}_{k}$. We use $\mathcal{D}_{-k}$ for the first stage estimation of $\theta_{h}$ and $\theta_{\mathcal{M}}$. This gives us $\left(\widehat{\theta}_{h}^{\left(-k\right)},\widehat{\theta}_{\mathcal{M}}^{\left(-k\right)}\right)$. We then use $\mathcal{D}_{k}$ to evaluate $Q_{h}$ and $Q_{g}$ at $\left(\widehat{\theta}_{h}^{\left(-k\right)},\widehat{\theta}_{\mathcal{M}}^{\left(-k\right)}\right)$. This gives us $\text{\ensuremath{\left(Q_{h}^{\left(k\right)}\left(\widehat{\theta}_{h}^{\left(-k\right)}\right),Q_{g}^{\left(k\right)}\left(\widehat{\theta}_{\mathcal{M}}^{\left(-k\right)}\right)\right)}}$. Finally, for cross-validation, $w$ is determined as
, where $\overline{Q}_{m}\doteq\frac{1}{K}\sum_{k=1}^{K}Q_{m}^{\left(k\right)}\left(\widehat{\theta}_{m}^{\left(-k\right)}\right),\ m\in\left\{ h,g/\mathcal{M}\right\} $ are cross-validated moment distances. For cross-fitting, let $w_{h}^{\left(k\right)}$ be constructed from $\text{\ensuremath{\left(Q_{h}^{\left(k\right)}\left(\widehat{\theta}_{h}^{\left(-k\right)}\right),Q_{g}^{\left(k\right)}\left(\widehat{\theta}_{\mathcal{M}}^{\left(-k\right)}\right)\right)}}$ according to (ref). Then the cross-fitted weight is\footnote{Both methods are consistent. See li_asymptotic_1987,chernozhukov_double_2016. Although to our knowledge, their asymptotic efficiency and finite sample performance have not been compared in existing studies.}
\paragraph*{Causal Inference}
We now discuss the problem of causal effect estimation under unconfoundedness. Let the observed variables be $\left(y,d,v\right)\in\mathbb{R}\times\mathbb{R}\times\mathcal{V}$, where $y$ is the outcome variable, $d$ is the treatment variable, and $v$ is a set of control variables. We are interested in the causal effect of $d$ on $y$. Specifically, let our target be the average treatment effect (ATE) denoted by $\tau$. We allow $\tau$ to be fully nonlinear and heterogeneous, i.e. $\tau=\tau\left(d,v\right)$. Then
, where $y^{d}$ is the potential outcome of $y$ under treatment $d$.
Under the unconfoundedness assumption of rosenbaum_central_1983\footnote{Suppose the treatment variable $d$ takes on a discrete set of values, $d\in\left\{ 1,\ldots,D\right\} $, then the unconfoundedness -- or conditional exchangeability -- assumption can be stated as \[\left.d\perp \!\!\! \perp\left(y^{d=1},\ldots,y^{d=D}\right)\right|v\]This assumption is satisfied if $d$ is not associated with any other causes of $y$ conditional on $v$, in which case we say $d$ is exogenous to $y$ conditional on $v$. A more precise statement on the sufficient conditions for satisfying this assumption, made in the language of causal graphical models based on directed acyclic graphs (DAGs), is that $v$ satisfies the back-door criterion pearl_causality_2009.}, $\mathbb{E}\left[\left.y^{d}\right|v\right]=\mathbb{E}\left[\left.y\right|d,v\right]$. Let $x=\left(d,v\right)$. The task of estimating $\tau\left(d,v\right)$ is thus equivalent to the task of estimating $\mathbb{E}\left[y|x\right]$. Suppose now that we have a reduced-form model $h\left(x;\theta_{h}\right)$ for $\mathbb{E}\left[y|x\right]$ and a structural model $\mathcal{M}\left(x,y;\theta_{\mathcal{M}}\right)$, both supporting the unconfoundedness condition\footnote{i.e. (1) the design of $h$ is based on the unconfoundedness condition; (2) in the causal structure assumed by $\mathcal{M}$, $v$ satisfies the back-door criterion.}, then we can use the DRSS to produce an estimate of $\mathbb{E}\left[y|x\right]$ by combining these two models, from which we can derive $\widehat{\tau}\left(d,v\right)$\footnote{Technically, $\tau\left(d,v\right)$ is the conditional ATE. With a slight abuse of notation, the population ATE $\tau\left(d\right)=\mathbb{E}_{v}\left[\tau\left(d,v\right)\right]$.}.
When the unconfoundedness condition does not hold so that $d$ is endogenous conditional on $v$, one of the most widely used strategies in reduced-form inference is to rely on the use of instrumental variables, which are auxiliary sources of randomness that can be used to identify causal effects. Let $h\left(x;\theta_{h}\right),\ x=\left(d,v\right)$ be a reduced-form model for $\mathbb{E}\left[\left.y^{d}\right|v\right]$. We can write $y=h\left(x;\theta_{h}\right)+\epsilon$, where $\epsilon$ is defined as $y-h\left(x;\theta_{h}\right)$ and may be correlated with $d$\footnote{By definition, when $\mathbb{E}\left[d\epsilon\right]\ne0$, the received treatment $d$ is related to unobserved factors that affect potential outcomes $y^{d}$, thus violating the unconfoundedness condition.}. If we have access to a variable $z$ that is correlated with $d$ (conditional on $v$) and satisfies $\mathbb{E}\left[z\epsilon\right]=0$, then $z$ can serve as an instrument for $d$\footnote{On a causal graph, this translates into the requirement that $z$ is correlated with $d$ and that every open path connecting $z$ and $y$ has an arrow pointing into $d$.}. In general, given $\theta_{h}\in\mathbb{R}^{p_{h}}$, let $\psi_{h}\left(x,y,z;\theta_{h}\right)=\phi\left(z\right)\left(y-h\left(x;\theta_{h}\right)\right)$ be a set of $\ell_{h}>p_{h}$ functions, where $\phi\left(z\right)$ is any function of $z$. If $h$ is the true model and $\theta_{h}^{0}$ is the true parameter, then $\theta_{h}^{0}$ can be identified via the following moment conditions:
Now let $\mathcal{M}\left(x,y,z;\theta_{\mathcal{M}}\right)$ be a structural model for the data-generating mechanism of the observed variables\footnote{$\mathcal{M}$ does not have to contain $z$. See e.g. section (ref) for an example. If $\mathcal{M}$ does contain $z$, $z$ needs to satisfy the IV requirement in the causal structure of $\mathcal{M}$, i.e. $z$ is correlated with $d$ and that every open path connecting $z$ and $y$ has an arrow pointing into $d$. If $\mathcal{M}$ is a model for $\left(x,y\right)$ only, in the case that it is the true model, the DRSS estimator for $\mathbb{E}\left[\left.y^{d}\right|v\right]$ will be based both on the causal assumptions in $\mathcal{M}$ and on the additional assumption that $z$ is a variable satisfying the IV requirement.}. Let $g\left(x;\theta_{\mathcal{M}}\right)=\mathbb{E}^{\mathcal{M}}\left[\left.y^{d}\right|v\right]$ be the model derived conditional expectation of the potential outcome under treatment $d$. Let $\psi_{g}\left(x,y,z;\theta_{\mathcal{M}}\right)$ be either a set of moment functions for $\mathcal{M}$ or let $\psi_{g}\left(x,y,z;\theta_{\mathcal{M}}\right)=\phi\left(z\right)\left(y-g\left(x;\theta_{\mathcal{M}}\right)\right)$. We can then construct $Q_{h}\left(\theta_{h}\right)$ and $Q_{g}\left(\theta_{\mathcal{M}}\right)$ based on $\psi_{h}\left(x,y,z;\theta_{h}\right)$ and $\text{\ensuremath{\psi_{g}\left(x,y,z;\theta_{\mathcal{M}}\right)}}$, and combine $h\left(x;\theta_{h}\right)$ and $g\left(x;\theta_{\mathcal{M}}\right)$ according to (ref) to produce a DRSS estimate of $\mathbb{E}\left[\left.y^{d}\right|v\right]$\footnote{The difference is that in (ref), by combining $h$ and $g$, we get $\widehat{\mathbb{E}}\left[\left.y\right|d,v\right]$. Here we get $\widehat{\mathbb{E}}\left[\left.y^{d}\right|v\right]$.}, from which we can obtain $\widehat{\tau}\left(d,v\right)$.
\paragraph*{Discussion}
The goal of doubly robust estimation is to ensure consistency when one of two candidate models is correctly specified but we do not know which one. When both models are misspecified, however, doubly robust estimators can perform poorly kang_demystifying_2007. This is not surprising as these estimators are not constructed to optimize performance based on a loss criterion such as expected mean squared error. In fact, the DRSS estimator can be viewed as a weighted average of its candidate models (see (ref)) and bears a close resemblance to bayesian model averaging, which is known to be flawed in $\mathcal{M}$-open settings in which none of the candidate models is true clyde_bayesian_2013,yao_using_2018\footnote{More precisely, bayesian model averaging is appropriate for $\mathcal{M}$-closed settings rather than $\mathcal{M}$-complete or $\mathcal{M}$-open settings. Following the definitions of bernardo_bayesian_2009, given a list of candidate models, the $\mathcal{M}$-closed setting is the one in which the true model is in the list. In the $\mathcal{M}$-complete setting, the true model can be specified but for tractability of computations or other reasons is not included in the model list. The $\mathcal{M}$-open setting refers to the situation in which we know the true model is not in the list and have no idea what it looks like.}.
In our presentation so far, we have also assumed that we have access to a representative sample drawn from the population of interest, i.e. the source domain is the same as the target domain. In practice, however, this is often not the case. In particular, we are often interested in making inference on populations that are much larger than the population from which we draw our sample, i.e. we care about the external validity or out-of-domain performance of our estimators. The DRSS however assures only in-domain consistency if one of its candidate models is correctly specified. In general, no similar guarantees on out-of-domain consistency can be obtained without further assumptions\footnote{This can be readily seen by considering two models that produce the same fit in-domain but behave completely differently out-of-domain. Without further assumptions, there is no way to tell them apart using observed data.}.
If our goal is not to achieve consistency on a target population, but rather to improve predictive accuracy as much as possible, then note that simply averaging a statistical model that fits well in-domain with an approximately correct structural model could improve the in-domain fit of the latter and the out-of-domain fit of the former. This observation applies to the DRSS as well, as it is also a weighted average method. The weights of the DRSS, however, are not constructed to optimize a performance criterion. This brings us to the ensemble estimators that we introduce in the next section, which are explicitly constructed to do so. As we will see, even though the criteria are evaluated on observed data, the ensemble estimators often produce superior in-domain and out-of-domain results relative to both of its candidate models and the DRSS approach, especially when both individual models are misspecified.
Given variables $\left(x,y\right)\in\mathcal{X\times\mathbb{R}}$, again assume that our goal is to learn the conditional expectation function $\mu\left(x\right)=\mathbb{E}\left[y|x\right]$ and we have at our disposal two parametric models $h\left(x;\theta_{h}\right)$ and $g\left(x;\theta_{g}\right)$. Let $\widehat{h}\left(x\right)\doteq h\left(x;\widehat{\theta}_{h}\right)$ and $\widehat{g}\left(x\right)\doteq g\left(x;\widehat{\theta}_{g}\right)$ be their fitted values on the observed sample. The linear ensemble, ESS-LN, combines the two linearly to form an estimate of $\mu\left(x\right)$:
To choose the optimal weights $w=\left(w_{0},w_{1},w_{2}\right)$, we can simply run a least squares regression of $y$ on $\widehat{h}\left(x\right)$ and $\widehat{g}\left(x\right)$. At the population level, combining models this way never make things worse hastie_elements_2009. On finite sample, however, we need to take into consideration differences in model complexity and avoid carrying over any biases in the first stage estimation of $\left(\widehat{\theta}_{h},\widehat{\theta}_{g}\right)$ into the choice of $w$. To this end, one can use the method of stacking wolpert_stacked_1992 and obtain $w$ via leave-one-out cross validation:
, where $\widehat{h}^{-i}\left(x_{i}\right)$ and $\widehat{g}^{-i}\left(x_{i}\right)$ are respectively the predictions at $x_{i}$ using $h$ and $g$ that are estimated on the training data with the $i$th observation removed. The cross-validated error gives a better approximation of the expected error, allowing an optimal combination. In practice, one can also account for model complexity via the use of sample-splitting or cross-fitting, or use $K-$fold instead of leave-one-out cross validation.
To adapt the stacking method to combining statistical and structural models, as in the construction of the DRSS estimator, we let $g\left(x;\theta_{\mathcal{M}}\right)=\mathbb{E}^{\mathcal{M}}\left[y|x\right]$ be the implied conditional mean of $y$ according to the structural model $\mathcal{M}\left(x,y;\theta_{\mathcal{M}}\right)$. We then combine $g\left(x;\theta_{\mathcal{M}}\right)$ with statistical model $h\left(x;\theta_{h}\right)$ according to (ref). With regard to the choice of $w$, in wolpert_stacked_1992, no restrictions are placed and $\widehat{w}$ is given by least squares regression of $y_{i}$ on $\widehat{h}^{-i}\left(x_{i}\right)$ and $\widehat{g}^{-i}\left(x_{i}\right)$\footnote{The stacking method as proposed by wolpert_stacked_1992 is therefore a general model combination or ensemble method rather than a model averaging method.}. hansen_jackknife_2012 proved the asymptotic optimality of stacking for linear models under a model averaging constraint that $w_{0}=0,w_{1},w_{2}\ge0,w_{1}+w_{2}=1$. ando_weight-relaxed_2017 proved asymptotic optimality for generalized linear models with weight restrictions relaxed to $w_{0}=0,w_{1},w_{2}\in\left[0,1\right]$. In this paper, we follow the original stacking method and do not place restrictions on $w$\footnote{In particular, both hansen_jackknife_2012 and ando_weight-relaxed_2017 assumed individual (generalized) linear models with intercept terms, so that their prediction errors have mean $0$. In our case, we do not require misspecified structural models to generate predictions of $y$ that have mean $0$ error. We thus need an additional intercept term $w_{0}$.}.
We now discuss the use of ESS-LN for causal effect estimation. As discussed in section (ref), given treatment variable $d$, outcome variable $y$, and control variables $v$, the task of estimating the conditional ATE under unconfoundedness is equivalent to the task of estimating the conditional expectation $\mathbb{E}\left[y|d,v\right]$\footnote{Technically, the conditional ATE $\tau\left(d,v\right)=\left.\partial\mathbb{E}\left[y|d,v\right]\right/\partial d$ under unconfoundedness.}$^{,}$\footnote{When the unconfoundedness condition does not hold, a number of reduced-form strategies are often employed to identify causal effects. In addition to the use of instrumental variables, which we detail below, these methods include difference-in-differences (DID) and regression discontinuity (RD). Statistically, both DID and RD can be cast as a conditional mean estimation problem given specific designs and thus can be combined with their structurally-derived counterpart using the ensemble method we have described.}. Procedurally, the causal inference problem is thus the same as the statistical prediction problem in this case\footnote{We note that in current practice, the goal of causal inference is typically to produce an unbiased estimate of the treatment effect, while in predictive modeling, the goal is to often to minimize an expected $\mathcal{L}_{2}$ loss. However, whether causal effect estimation should aim for unbiasedness or precision remains an unsettled question.}$^{,}$\footnote{Importantly, in the case of ensemble estimators, even if the ensemble model estimates causal effects based on the unconfoundedness assumption, the structural model in the ensemble does not have to support the assumption. Whatever the causal assumptions are made by the structural model, we use its derived functional form for $\mathbb{E}\left[y|d,v\right]$ as an input into the ensemble. Thus, the final ensemble estimate is still based on the unconfoundedness assumption. If this assumption holds true but is unsupported by a member model in the ensemble, then that model is simply misspecified. }.
In general, however, without assuming unconfoundedness, our goal is to produce an estimate of $\mathbb{E}\left[\left.y^{d}\right|v\right]$ based on a reduced-form model $\widehat{h}\left(x\right)=h\left(x;\widehat{\theta}_{h}\right)$ and a structurally-derived model $\widehat{g}\left(x\right)=g\left(x;\widehat{\theta}_{\mathcal{M}}\right)$:
, from which we can obtain $\widehat{\tau}\left(d,v\right)=\left.\partial\widehat{\mathbb{E}}\left[\left.y^{d}\right|v\right]\right/\partial d$.
When $d$ is endogenous -- when there is unmeasured confounding, if we observe a variable $z$ that can serve as a valid instrument for $d$, then we can specify the following $\ell\times1,\ \ell\ge3$ moment conditions:
, where $\phi\left(z\right)$ is any function of $z$ and $w^{0}=\left(w_{0}^{0},w_{0}^{1},w_{0}^{2}\right)$ are the true values of $w$\footnote{Assuming that (ref) is the true model.}$^{,}$\footnote{The structural model $\mathcal{M}$ from which $\widehat{g}\left(x\right)$ is derived does not have to contain $z$, and if it does, $z$ does not need to satisfy the IV requirement in the causal structure assumed by $\mathcal{M}$. See footnote (ref).}.
Let $\psi\left(x,y,z;w\right)\doteq\phi\left(z\right)\left(y-w_{0}+w_{1}\widehat{h}\left(x\right)+w_{2}\widehat{g}\left(x\right)\right)$. Let $\overline{\psi}\left(w\right)\doteq\frac{1}{n}\sum_{i=1}^{n}\psi\left(x_{i},y_{i},z_{i};w\right)$. Let $Q\left(w\right)\doteq\overline{\psi}\left(w\right)'\Omega\overline{\psi}\left(w\right)$, where $\Omega$ is a $\ell\times\ell$ positive definite weight matrix\footnote{e.g. the efficient GMM weight of hansen_large_1982.}. The optimal $w$ can then be obtained by minimizing the GMM objective function:
In practice, as in the case of conditional mean modeling, given finite sample, we want to account for model complexity and avoid carrying any bias in the first stage estimation of $\widehat{h}$ and $\widehat{g}$ into the determination of $w$. This can be accomplished by using the strategies of either sample-splitting, cross-validation, or cross-fitting.
The ESS-LN is a linear ensemble. Our ESS-NP estimator goes one step further and allows any nonlinear combinations of individual models. In conditional mean estimation, let
, where $f\left(.,.\right)$ is any function. Statistically, this amounts to regressing the outcome $y$ nonparametrically on the predictions obtained from individual models $h$ and $g$.
While a large class of nonparametric models can be used for $f$, in this paper we adopt the random forest model of breiman_random_2001. The random forest is based on decision tree models. A decision tree is constructed by repeatedly splitting or partitioning the predictor space into different regions in order to maximize fit. In each region, a constant model is fit so that the predicted value is simply the mean of the observed outcomes in that region. Thus, in its simplest form, with a predetermined number of splits (such as in the case of a stump), a decision tree is a piecewise-constant model. When splits are adaptively chosen to minimize prediction error, the decision tree becomes a nonparametric model whose complexity grows with data and is related to kernels and nearest-neighbor methods in that its predictions are based on the values of neighborhood observations, except that it chooses the neighborhoods (regions) in a data-driven way athey_generalized_2019.
In contrast to conventional trees, in the ESS-NP, the predictor space is formed by $\widehat{h}\left(x\right)$ and $\widehat{g}\left(x\right)$ -- the predictions obtained from statistical model $h$ and structurally-derived model $g$. A tree constructed out of $\widehat{h}\left(x\right)$ and $\widehat{g}\left(x\right)$ carves up the space formed by $\widehat{h}\left(x\right)$ and $\widehat{g}\left(x\right)$, which in turn, implies a partition of the underlying input space $x$. The ESS-NP can therefore be viewed as allowing us to adaptively assign different weights to different regions of the input space depending on which model -- the statistical or the structural -- performs better.
While decision trees are powerful tools for capturing nonlinear relations and complex interactions, they tend to suffer from high variance and instability. Random forests improve upon decision trees by building and combining a large number of trees through bootstrap aggregation, thereby reducing variance and increasing predictive accuracy\footnote{The random forest is an ensemble of individual trees. In our ESS-NP estimator, each tree is in turn an ensemble of $h$ and $g$. The ESS-NP is therefore an “ensemble of ensembles”.}. Additional randomness can be introduced to further de-correlate individuals trees via random split selection that restricts the variables available for consideration in each split\footnote{See loh_fifty_2014,biau_random_2016 for overviews of decision trees and forest-based methods. Consistency results on random forests are obtained in biau_analysis_2012,scornet_consistency_2015,scornet_asymptotics_2016.}. In the ESS-NP estimator (ref), $f$ is therefore based on the random forest model.
The conditional mean ESS-NP estimator can be used for prediction and causal effect estimation under unconfoundedness\footnote{The estimator can also be used to combine structural models with reduced-form models based on statistical designs such as DID and RD when there is unmeasured confounding.}. When there is unmeasured confounding, as in the case of ESS-LN, it is conceptually possible to adapt the ESS-NP to perform instrumental variables estimation based on the following conditional moment restrictions:
, where $f\left(.,.\right)$ is again any function. The type of nonparametric IV regression defined by (ref), however, is known to suffer from poor statistical performance due to the ill-posed inverse problem newey_nonparametric_2013. Applying the random forest method to this task is also not straight-forward\footnote{Methods for estimating heterogeneous causal effects with semiparametric IV regression based on random forests have recently been proposed in athey_generalized_2019.}. Therefore, in this paper, we do not propose an ESS-NP method for IV estimation.
In this section, we demonstrate the effectiveness of our methods and compare their finite-sample performances using three sets of simulated experiments. Taken together, these exercises cover prediction and causal inference problems, static and dynamic settings, and individual behavior that deviates in various ways from perfect rationality.
In our first experiment, we consider first-price sealed-bid auctions. Auctions are one of the most important market allocation mechanisms. Empirical analysis of auction data has been transformed in recent years by structural estimation of auction models based on games of incomplete information\footnote{See paarsch_introduction_2006,athey_nonparametric_2007,hickman_structural_2012,perrigne_econometrics_2019 for surveys on econometric analysis of auction data}. Structural analysis of auction data views the observed bids as equilibrium outcomes and attempts to recover the distribution of bidders' private values by estimating relationships derived directly from equilibrium bid functions. This approach, while offering a tight integration of theory and observations, relies on a set of strong assumptions on the information structure and rationality of bidders bajari_are_2005.
In this exercise, we conduct three experiments by simulating auction data with varying number of participants under three scenarios. The first scenario features rational bidders with independent private values drawn from a uniform distribution. The second scenario features rational bidders whose values are drawn from a beta distribution. The third scenario features boundedly-rational bidders whose bids deviate from optimal bidding strategies. In each experiment, we're interested in the effect of the number of bidders $n$ on the winning bid $b^{*}$, $\mathbb{E}\left[\left.b^{*}\right|n\right]$. We estimate this target function using (a) a statistical model, (b) a structural model, (c) the DRSS estimator, (d) the ESS estimators (ESS-LN, ESS-NP), and compare their performances. For all experiments, we use a structural model that assumes rational bidders with uniformly distributed values. The model is thus correctly specified for experiment 1, but is misspecified in experiment 2 and 3. Table (ref) summarizes this setup. Below we detail the data-generating models of the three experiments.
\paragraph{Setup}
Consider a first-price sealed-bid auction with $n$ risk-neutral bidders with independent private value $v_{i}\sim^{i.i.d.}F(v)$. Each bidder submits a bid $b_{i}$ to maximize her expected return
, where $b_{-i}$ denotes the other submitted bids. In Bayesian-Nash equilibrium, each bidder's bidding strategy is given by
For experiment 1 and 3, we let $F$ be $U\left(0,1\right)$. In this case the equilibrium bid function simplifies to:
For experiment 2, we let $F$ be $\text{Beta}\left(2,5\right)$. In each experiment, we simulate repeated auctions with varying number of bidders\footnote{Assuming the same object is being repeatedly auctioned.}. For experiment 1 and 2, the observed bids $b_{i}$ are the equilibrium outcomes, i.e. $b_{i}=b\left(v_{i}\right)$. For experiment 3, we let $b_{i}=\eta_{i}\cdot b\left(v_{i}\right)$, where $\eta_{i}$ follows a normal distribution left-truncated at $0$, $\eta_{i}\overset{\text{i.i.d.}}{\sim}\text{TN}\left(0,0.25,0,\infty\right)$. Bidders in experiment 3 thus “overbid” relative to the Bayesian-Nash equilibrium.
\paragraph*{Simulation}
For each experiment, we simulate $M=500$ auctions with number of bidders $n_{m}$ varying between $5$ and $25$. The observed data thus consist of $\text{\ensuremath{\mathcal{D}}}=\left\{ \left\{ b_{i}^{m}\right\} _{i=1}^{n_{m}}\right\} _{m=1}^{M}$. In this exercise, our goal is to learn $\mathbb{E}\left[\left.b^{*}\right|n\right]$, the relationship between the number of bidders and the winning bid. To assess the performance of various estimators, we use the true data-generating models to compute $\mathbb{E}\left[\left.b^{*}\right|n\right]$ for $n\in\left[5,50\right]$, so that we can compare the predictions of each method with the true values both in-domain and out-of-domain.
\paragraph*{Statistical Model}
To estimate $\mathbb{E}\left[\left.b^{*}\right|n\right]$ using a statistical model\footnote{Since $n$ is exogenous, $\mathbb{E}\left[\left.b^{*}\right|n\right]$ is also a causal relationship and (ref) can also be thought of a reduced-form model of the effect of the number of bidders on the winning bid.}, the data we need are $\left\{ \left(n_{m},b_{m}^{*}\right)\right\} _{m=1}^{M}$, where $b_{m}^{*}$ is the winning bid of auction $m$. We adopt the following second degree polynomial as the model for $\mathbb{E}\left[\left.b^{*}\right|n\right]$:
\paragraph*{Structural Model}
Our structural model assumes that bidders are rational, risk-neutral, and have independent private values drawn from a $U\left(0,1\right)$ distribution. Under these assumptions, the bidders' private values can be easily identified from the observed bids in each auction by $v_{i}=\frac{n}{n-1}b_{i}$\footnote{In general, if we do not impose the assumption that $v_{i}\overset{\text{i.i.d.}}{\sim}U(0,1)$ and assume instead that $v_{i}\overset{\text{i.i.d.}}{\sim}F\left(v\right)$, with $F$ unknown, then we can identify and estimate $v_{i}$ using the following strategy based on guerre_optimal_2000: let $G\left(b\right)$ and $g\left(b\right)$ be the distribution and density of the bids. (ref) implies \[ v_{i}=b_{i}+\frac{1}{n-1}\frac{G\left(b_{i}\right)}{g\left(b_{i}\right)} \] Thus, by nonparametrically estimating $G\left(b\right)$ and $g\left(b\right)$ from the observed bids, we can obtain an estimate of $v_{i}$.}. The structural model makes it even easier to make predictions on the winning bid. The model implies that:
No estimation is necessary.
\paragraph*{Results}
Figure (ref) and (ref) show the results of the first experiment. In Figure (ref), we plot the number of participants $n$ against the winning bid $b^{*}$, the true relationship $\mathbb{E}\left[\left.b^{*}\right|n\right]$, and the predictions obtained from five models: statistical, structural, DRSS, ESS-LN, and ESS-NP. Since the structural model is the true model in this experiment, it predicts the true expected winning bids. The other four models, however, all fit relatively well. Figure (ref) plots the results of extrapolating the model predictions from $n\in\left[5,25\right]$ to $n\in\left[2,50\right]$. While the structural predictions still hold true, the statistical fit becomes very bad, as can be expected. Because the structural model is correctly specified while the statistical model is not, the DRSS puts most of the weight on the structural model and closely approximates its performance. The two ensemble estimators, ESS-LN and ESS-NP, are also able to significantly outperform the statistical model out-of-domain. In the first panel of Table (ref), we report the bias, variance, and mean squared error of all the estimators for $100$ simulation runs\footnote{Given an estimator $f$, let $f^{\left(r\right)}\left(n\right)$ denote the estimator's prediction of the winning bid in simulation $r$, then
Reported are their empirical estimates.}. In domain, compared to the true structural model, the DRSS provides the best fit, followed by the ESS-LN. Both the statistical and the ESS-NP models fit well as well. Out of domain, the statistical model has by far the worst performance. The three proposed estimators all have similar MSE and achieve significant gains in performance over the statistical model. Out of the three, the DRSS has the smallest bias. Thus, the DRSS estimator appears to work the best in this experiment. This is not surprising as one of its candidate models is correctly specified, satisfying the condition for DRSS consistency.
Figure (ref) $-$ (ref) show the results of experiment 2 and 3. The results tell as similar story. In both experiments, the structural model is misspecified. In experiment 2, it misspecifies the private value distribution. In experiment 3, it assumes that bidders are rational and the observed bids are Bayesian-Nash equilibrium outcomes when they are not. As a consequence, in both cases, the structural fit deviates from the true model significantly. The statistical model, like in experiment 1, is able to fit well in-domain but poorly out-of-domain. Since both of its candidate models are misspecified in these experiments, the DRSS does not perform well. As the statistical model has better in-domain fit relative to the misspecified structural, the DRSS puts the majority of its weight on the statistical model. In comparison, the two ensemble estimators are able to both fit well in-domain and extrapolate better than the statistical, the structural, and the DRSS models. In the second and third panels of Table (ref), we observe the performance of these estimators over $100$ simulation runs. In both experiments, the ESS-LN produces the best in-domain fit, while the ESS-NP produces the best out-of-domain fit. Intuitively, the ensemble methods are able to achieve these performance gains due to a complementarity that exists between the statistical and the structural models in these two experiments: the statistical model fits well in-domain, while the structural model, though misspecified, provides useful guidance on the functional form of $\mathbb{E}\left[\left.b^{*}\right|n\right]$ when we extrapolate beyond the observed domain, as evidenced in Figure (ref), (ref).
Our second application concerns the modeling and estimation of firm entry and exit dynamics. Structural analysis of dynamic firm behavior based on dynamic discrete choice (DDC) and dynamic game models has been an important part of empirical industrial organization\footnote{See aguirregabiria_dynamic_2010,bajari_game_2013 for surveys on structural estimation of dynamic discrete choice and dynamic game models.}. These dynamic structural models capture the path dependence and forward-looking behavior of agents, but pays the price of imposing strong behavioral and parametric assumptions for tractability and computational convenience.
In this exercise, we focus our attention on the rational expectations assumption that has been a key building block of dynamic structural models in macro- and microeconomic analyses. The assumption and its variants state that agents have expectations that do not systematically differ from the realized outcomes\footnote{More precisely, rational expectations are mathematical expectations based on information and probabilities that are model-consistent muth_rational_1961.}. Despite having long been criticized as unrealistic, the rational expectations paradigm has remained dominant due to a lack of tractable alternatives and the fact that economists still know preciously little about belief formation.
We conduct three experiments in the context of the dynamic entry and exit of firms in competitive markets in non-stationary environments. Our data-generating models are DDC models of entry and exit with entry costs and exogenously evolving economic conditions. In our first experiment, agents have rational expectations about future economic conditions. In the second experiment, agents have a simple form of adaptive expectations that assume the future is always like the past. The third experiment features myopic agents who optimize only their current period returns. In all experiments, we are interested in predicting the number of firms that are operating in the market each period. To this end, we estimate (a) a statistical model, (b) a structural model, and combine them using (c) the DRSS estimator, and (d) the ESS estimators (ESS-LN, ESS-NP). The structural model we estimate assumes rational expectations and is thus correctly specified only in experiment 1. Table (ref) summarizes this setup.
\paragraph{Setup}
Consider a market with $N$ firms. In each period, the market structure consists of $n_{t}$ incumbent firms and $N-n_{t}$ potential entrants. The profit to operating in the market at time $t$ is $R_{t}$, which we assume to be exogenous and time-varying. At the beginning of each period, both incumbents and potential entrants observe the current period payoff $R_{t}$ and each draws an idiosyncratic utility shock $\epsilon_{it}$. Incumbent firms then decide whether to remain or exit the market by weighing the expected present values of each option, while potential incumbents decide whether or not to enter the market, which will incur a one-time entry cost $c$. Specifically, let the entry status of a firm be represented by $\left(0,1\right)$. The time-$t$ flow utility of a firm, who is in state $j\in\left\{ 0,1\right\} $ in time $t-1$ and state $k\in\left\{ 0,1\right\} $ in time $t$, is given by
, where
is the deterministic payoff function and $\epsilon_{it}=\left(\epsilon_{it}^{0},\epsilon_{it}^{1}\right)$ are idiosyncratic shocks, which we assume are i.i.d. type-I extreme value distributed. The parameter $\alpha$ measures the importance of operating profits to entry-exit decisions relative to the idiosyncratic utility shocks.
The ex-ante value function of a firm at the beginning of a period is given by
, where $j$ is the firm's state in $t-1$, $\beta$ is the discount factor, $\overline{V}_{t}^{j}\coloneqq\mathbb{E}_{\epsilon}\left[V_{t}^{j}\left(\epsilon_{it}\right)\right]$ is the expected value integrated over idiosyncratic shocks, and $\mathcal{V}_{t}^{jk}\coloneqq\pi_{t}^{jk}+\beta\cdot\mathbb{E}_{t}\left[\overline{V}_{t+1}^{k}\right]$ is the choice-specific conditional value function.
At the beginning of each period, after idiosyncratic shocks are realized, each firm thus chooses its action, $a_{it}\in\left\{ 0,1\right\} $, by solving the following problem:
, which gives rise to the conditional choice probability (CCP) function:
, which follows from the extreme value distribution assumption.
Since the value function involves the continuation values $\mathbb{E}_{t}\left[\overline{V}_{t+1}^{k}\right]$, which requires expectations of the future profits $\left(R_{t+1},R_{t+2},\ldots\right)$, its solution requires us to specify how such expectations are formed. In experiment 1, we assume firms have perfect foresight on $R_{t}$. This is a stronger form of rational expectations that assumes individuals knows the future realized values. Firms can then compute $\overline{V}_{t}^{j}=\mathbb{E}_{\epsilon}\left[V_{t}^{j}\left(\epsilon_{it}\right)\right],\ j\in\left\{ 0,1\right\} $ in a model-consistent way, i.e. based on the distributional assumption of $\epsilon_{it}$. In experiment 2, we assume firms have a form of adaptive expectations, according to which beliefs about the future are formed based on past values. Here for simplicity, we assume that firms expect future profits to be always the same as in current period, i.e. $R_{t}=R_{t+1}=R_{t+2}=\cdots$. Finally, in experiment 3, we allow firms to be myopic, so that they do not care about the future and only maximize current payoffs.
\paragraph*{Simulation}
For each experiment, we simulate $N=10,000$ firms for $\mathcal{T}=1000$ periods. The first $T=500$ periods are used for training and the last $\mathcal{T}-T=500$ periods are used to assess the out-of-domain performance of our estimators. The training data thus consist of $\text{\ensuremath{\mathcal{D}}}=\left\{ \left\{ a_{it}\right\} _{i=1}^{N},R_{t}\right\} _{t=1}^{T}$. We simulate $R_{t}$ to follow an autoregressive process with a time trend so that the environment is non-stationary. Figure (ref) shows a realized path of $R_{t}$. A different $R_{t}$ process is chosen for each experiment so that the entry and exit dynamics over the first $T$ periods are significantly different from the last $\mathcal{T}-T$ periods, allowing us to better distinguish the performance of the estimators. Appendix \hyperlink{APP}{B.1} reports the parameter values we use as well as other details of the simulation.
\paragraph*{Statistical Model}
To predict the number of firms operating in the market each period, $n_{t}$, based on observed exogenous operating profits, $R_{t}$, we adopt the following ARX model:
\paragraph*{Structural Model}
We estimate the DDC model given by (ref)--(ref) assuming rational expectations. Our estimation strategy builds on arcidiacono_conditional_2011 and estimates an Euler-type equation constructed out of CCPs. Here we sketch the strategy while presenting its details in Appendix \hyperlink{APP}{B.1}\footnote{See arcidiacono_practical_2011 for a review of related CCP estimators. For empirical implementations, see, e.g. artuc_trade_2010,scott_dynamic_2014.}. A key to our strategy is the assumption that because agents have rational expectations, their expected continuation values do not deviate systematically from the realized values, i.e. $\overline{V}_{t+1}^{j}=\mathbb{E}_{t}\left[\overline{V}_{t+1}^{j}\right]+\xi_{t}^{j}$, where $\xi_{t}^{j}$ is a time-$t$ expectational error with $\mathbb{E}\left(\xi_{t}^{j}\right)=0$. Given this assumption, and since our model has the finite dependence property of arcidiacono_conditional_2011, solution to (ref) can be written in the form of the following Euler equation:
, where $\epsilon_{t}^{j,k}=\beta\left(\xi_{t}^{k}-\xi_{t}^{j}\right)$.
Replacing the CCPs with their sample analogues, i.e. let $\widehat{p}_{t}\left(k|j\right)=$ observed percentage of firms that are in state $j$ in $t-1$ and state $k$ in time $t$, we obtain the following estimating equations: for all $j\ne k$,
, where $e_{t}=\left(e_{t}^{01},e_{t}^{10}\right)$ is an error term that captures both the expectational errors in $\epsilon_{t}^{j,k}$ and the approximation errors in $\widehat{p}_{t}\left(k|j\right)$.
We assume that the value of the discount factor $\beta$ is known. Estimating (ref) gives us an estimate of the model parameters $\left(\mu,\alpha,c\right)$. These estimates are consistent for a model that assumes rational expectations. Our structural model is therefore correctly specified for experiment 1, but misspecified in experiment 2 and 3.
\paragraph*{Results}
Figure (ref) shows the results of the first experiment. Figure (ref) plots the expected percentage of firms in the market, $\mathbb{E}\left[n_{t}\right]$, for entire periods of $t=1-1000$, including both the in-domain periods of $t=1-500$ and the out-of-domain periods of $t=501-1000$, together with the predictions of the five estimators. Predictions are made using one-step ahead forecasting\footnote{Given an estimated model, in each period $t$, we predict $n_{t}$ based on $\left\{ \left(n_{t-1},n_{t-2},\ldots\right),\left(R_{t},R_{t-1},\ldots\right)\right\} $. To generate predictions for the structural model, we also assume agents have perfect foresight regarding $\left(R_{t+1},R_{t+2},\ldots\right)$ .}. A closer look at in-domain and out-of-domain results are presented in Figure (ref) and (ref) for chosen periods.
All estimators fit relatively well in-domain. However, out-of-domain, the time series model is unable to capture the rising number of firms as $R_{t}$ increases. This is partly by design: as we have discussed, we intentionally choose parameter values so that out-of-domain dynamics differ markedly from those in-domain. A statistical model that fits to the in-domain data is unable to extrapolate well in this case. On the other hand, the structural model, which is correctly specified in this experiment, extrapolates very well, as expected. Since one of its candidate models is correctly specified, the DRSS is also expected to perform well. Here, the DRSS model successfully allocates most of its weight on the structural model. However, because some weight is still put on the statistical model, it systematically underestimates the number of firms in out-of-domain periods as well. This is also expected as inability to distinguish between competing models based on limited data is what motivates doubly robust and model averaging approaches in the first place. Like the DRSS, the two ensemble estimators are able to largely capture the rising number of firms in out-of-domain periods, offering significantly better predictions than the statistical model. Out of the two ensemble models, the ESS-LN performs particularly well, matching the true model closely.
In Table (ref) Panel 1, we report the bias, variance, and mean squared error of all the estimators with respect to the true $\mathbb{E}\left[n_{t}\right]$ over $100$ trials. Somewhat surprisingly, the structural model, albeit correctly specified, performs the worst in terms of MSE out of the five estimators in-domain. This is perhaps due to a loss of efficiency associated with our Euler-equation approach in estimating the model aguirregabiria_euler_2013. Out of domain, though, it predictably delivers the best performance. Out of the remaining four estimators, the ESS-NP produces the best in-domain fit, while the ESS-LN produces the best out-of-domain fit.
Figure (ref) shows the results of the second experiment. In Experiment 2, agents have adaptive expectations in the sense that they always assume $R_{t'}=R_{t}\ \forall t'>t$. Since in our simulations, $R_{t}$ follows a rising trend, this means that agents systematically underestimate future profits. The realized dynamics show that for most of the in-domain periods, there are few firms in the market. Number of firms increases significantly during the out-of-domain periods. This marked difference between in-domain and out-of-domain dynamics pose significant challenges. Looking at the model fits, the time series model again fits relatively well in-domain but is completely unable to extrapolate out-of-domain. The structural model, being misspecified, is able to capture the rising entries, but tends to have larger fluctuations than the true model. This can be explained by the fact that agents in the structural model assumes that future profits will be the same as current profits, thus reacting more dramatically to any changes in $R_{t}$. As both the statistical and the structural model are misspecified, the DRSS does not perform well. It puts most of the weight on the statistical model, leading to a bad extrapolation performance. The ensemble models, ESS-LN and ESS-NP, are both able to fit well in-domain and capture some part of the rising trend out-of-domain. Compared to the structural model, they tend to underfit rather than overfit the true expected number of firms in out-of-domain periods.
Looking at Panel 2 of Table (ref), we see that the ESS-NP achieves the smallest MSE both in-domain and out-of-domain, making it the winner in this experiment. The structural model is a close second in out-of-domain performance but is by far the worst in-domain. Indeed, the DRSS, the ESS-LN, and the ESS-NP all achieve significantly smaller MSEs in-domain. This experiment serves to illustrate a scenario in which the complementarity between the structural and the statistical model is especially pronounced, with the former fitting relatively badly in-domain and the latter completely unable to extrapolate. By combining the two, our ensemble models mainly rely on the former to guide out-of-domain prediction and on the latter to regulate in-domain fit.
Figure (ref) shows the results of the third experiment. In this experiment, agents are myopic in that they only care about current period returns when making entry and exit decisions. The data-generating model is therefore static in nature. Looking at estimator performances, the story is broadly similar to that of experiment 2, with the difference being that, in this experiment, the true model exhibits less dramatic difference between its in-domain and out-of-domain dynamics and the misspecified structural model tends to more significantly overestimate the number of firms in the market. As a consequence, according to Panel 3 of Table (ref), the structural model is the worst performer both in-domain and out-of-domain in this experiment. On the other hand, both ensemble estimators perform better than the other estimators both in-domain and out-of-domain, with the ESS-NP the clear winner. Thus, as in the auction experiments, our ensemble methods are able to consistently outperform the other estimators when both the structural and the statistical model are misspecified.
In our final application, we revisit the demand estimation problem under a different setting. Suppose now that instead of observing consumer demand under exogenously varying prices, the prices we observe are set by a monopolist. In this case, changes in prices are endogenous and the relationship between price and quantity sold is confounded. We are interested in learning the true demand curve. To this end, if we have access to a variable that shifts the cost of production for the monopoly firm but does not affect demand directly, then it can be used as an instrumental variable to help identify the demand curve. This is the reduced-form approach. Alternatively, we can estimate a structural model that fully specifies monopoly pricing behavior. This is the structural approach. Finally, we can combine the two using the DRSS and the ESS-LN\footnote{For instrumental variable estimation, we do not offer an ESS-NP estimator.}.
In this exercise, we conduct four experiments. In all four experiments, we assume that we have access to a valid instrument so that the demand curve is identified. However, the functional form of the reduced-form model may still be misspecified. On the other hand, using the structural approach, we estimate a model that assumes the observed prices are optimally set by a profit-maximizing monopoly firm. When this assumption is violated, as when for example the firm's pricing is not optimal or it does not have monopoly power, the structural model will also be misspecified. The four experiments we conduct are thus arranged as follows: in the first experiment, both the reduced-form and the structural models are correctly specified. In experiment 2 and 3, only one of the two is correctly specified. In experiment 4, both are misspecified. Table (ref) summarizes this setup. For each experiment, we also simulate both a slightly confounded data set, in which the relationship between price and quantity does not deviate too much from the demand curve, and a highly confounded data set, in which they look nothing alike.
In contrast to the previous two exercises, in this exercise, we focus on comparisons of in-domain performance. We show that when either the reduced-form or the structural model is misspecified, the DRSS and the ESS-LN will have better in-domain performance -- more internal validity -- than the misspecified model. When both are misspecified, the ESS-LN outperforms them both.
\paragraph*{Setup}
Consider $M$ geographical markets in which a product is sold. The equilibrium price and quantity sold in market $m$ are $\left(p_{m},q_{m}\right)$. Assume that all markets share the same aggregate demand function $Q^{d}\left(p\right)$:
In experiment 1 and 3, we assume the product is sold by a monopoly firm who sets the prices in each market to maximize its profit. The firm has different marginal costs $c_{m}$ for operating in different markets. Hence it sets
Assume that we also observe a cost-shifter $z_{m}$, e.g. transportation costs, such that
, then $z_{m}$ can serve as an instrument for $p_{m}$ for identifying the demand curve.
In experiment 2 and 4, we assume the monopoly firm fails to set optimal prices or does not have complete monopoly power. Its pricing decisions are given by
, where $\lambda\in\left(0,1\right)$. The firm thus earns a lower markup than a optimal price-setting monopoly.
\paragraph*{Simulation }
For each experiment, we simulate two data sets. Each data set consists of prices, quantities, and cost shifters in $M=1000$ markets, i.e. $\text{\ensuremath{\mathcal{D}}}=\left\{ \left(p_{m},q_{m},z_{m}\right)\right\} _{m=1}^{M}$. One data set is only slightly confounded, so that $\mathbb{E}\left[\left.q_{m}\right|p_{m}\right]$ is close to the demand relation (ref). The other is highly confounded, so that they are completely different. See Appendix \hyperlink{APP}{B.2} for the parameter values we use in simulation.
\paragraph*{Reduced-Form Model}
Because $p_{m}$ is now endogenous -- $p_{m}$ and $\epsilon_{m}$ are correlated through (ref) -- the statistical relation between $p_{m}$ and $q_{m}$ is confounded and no longer represents the demand function. To estimate the demand curve using the reduced-form approach, we avail of the instrumental variable $z_{m}$ and estimate $Q^{d}\left(p\right)$ by two-stage least squares (2SLS). In experiment 1 and 2, our reduced-form model is correctly specified, i.e. we fit (ref) to the data by 2SLS. In experiment 3 an 4, however, we assume the demand function takes on a log-log form:
, and is therefore misspecified in these two experiments.
\paragraph*{Structural Model}
We fit a structural model featuring linear demand function (ref) and price-setting function (ref). This structural model is correctly specified for experiment 1 and 3, but misspecified for experiment 2 and 4. The structural parameters are $\left(\alpha,\beta,a,b\right)$ and can be estimated as follows: from (ref) and (ref), we obtain
If our model is correct, (ref) is a deterministic linear equation system from which we can solve directly for $\left(\widehat{a},\widehat{b},\widehat{\beta}\right)$. Substituting $\widehat{\beta}$ into (ref), we then obtain $\widehat{\alpha}=\frac{1}{M}\sum_{m=1}^{M}\left(q_{m}+\widehat{\beta}p_{m}\right)$.
\paragraph*{Results}
In Figure (ref) and (ref), we plot the results of the four experiments respectively for the slightly and highly confounded scenarios. In the latter case, the observed data $\left(p_{m},q_{m}\right)$ are significantly confounded such that fitting a least squares model to the data would produce an upward-sloping curve. Regardless of the level of confounding, however, the two groups of plots tell a similar story. When correctly specified, both reduced-form and structural estimation are able to identify the true demand curve (Figure (ref), (ref))\footnote{This is because both use correctly specified models and $z$ is a valid instrument.}. When only one of them is correctly specified, the misspecified model produces fits that, while still managing to capture the downward-sloping nature of the demand curve, can deviate significantly from the true relationship (Figure (ref), (ref), (ref), (ref)). In this case, the ESS-LN generally still performs well, while the DRSS is able to fit the demand curve well in Figure (ref) and (ref) but not in (ref) and (ref). Finally, when both the reduced-form and the structural models are misspecified, the ESS-LN becomes the only method that is able to fit the true demand curve well (Figure (ref), (ref)).
Table (ref) reports the bias, variance, and mean squared error of the estimators with respect to the true demand curve over 100 trials. In both the slightly and highly confounded scenarios, when they are correctly specified, the reduced-form and the structural models exhibit low biases. The structural model, by virtue of imposing more structure on the data, attains a lower variance. When misspecified, both types of models exhibit large biases and MSEs. The DRSS is able to outperform the misspecified model in experiment 2 and 3, while the ESS-LN consistently achieves the lowest MSE -- often significantly lower than those of the other estimators, regardless of which model -- the reduced-form or the structural or even both -- is misspecified. Note, however, for all experiments, the DRSS and the ESS-LN perform better on the slightly confounded data. This is not surprising. In particular, as Figure (ref) reveals, when the data are highly confounded, the structural and the reduced-form models can behave similarly on the observed data, even when their predicted demand curves are actually very different due to one or both of them being misspecified, making it difficult for the DRSS method to distinguish between them and for the ESS-LN to leverage their differences in functional form. More confounding thus presents more challenges for our methods to work well.
In this paper, we propose a set of methods for combining statistical and structural models for improved prediction and causal inference. We demonstrate the effectiveness of our methods in a number of economic applications including first-price auctions, dynamic models of entry and exit, and demand estimation with instrumental variables. Our methods offer a way to bridge the gap between the (reduced-form) statistical approach and the structural approach in economic analysis and have potentially wide applications in addressing problems for which significant concerns about model misspecification exist.