Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
111,121 characters · 22 sections · 53 citation commands
Testing for Endogeneity: A Moment-Based Bayesian Approach
Keywords: Bayesian inference; Causal inference; Exponentially tilted empirical likelihood; Endogeneity; Exogeneity; Instrumental variables; Marginal likelihood; Posterior consistency. \thispagestyle{empty}
\thispagestyle{empty}
\doublespacing
\setcounter{page}{2}
Consider the semiparametric linear regression model \[ y = x^{\prime}{\Greekmath 010C} + z_{1}^{\prime}{\Greekmath 010D} + {\Greekmath 0122}, \] where $y \in \mathbb{R}$ is the outcome variable, $x \in \mathbb{R}^{d_x}$ is the treatment vector of interest, $z_{1} \in \mathbb{R}^{d_{z_{1}}}$ is a vector of controls, and ${\Greekmath 0122}$ is an unobserved disturbance. A common assumption in Bayesian analysis is that the regressors $x$ are exogenous, meaning that they are uncorrelated with the error term ${\Greekmath 0122}$. In many empirical settings this assumption is questionable. If one has access to a set of valid instruments $z_{2} \in \mathbb{R}^{d_{z_{2}}}$, with dimension at least as large as that of $x$, it becomes possible to conduct a Bayesian analysis that correctly accounts for endogeneity. Such analysis can be formulated within both parametric and semiparametric frameworks, as in the early contributions of Dreze1976 and the subsequent developments in KleibergenVanDijk1998, CHAOPhillips1998, KleibergenZivot2003, and Schennach2005, among many others. Recent work, including HoogerheideKleibergenVanDijk200763, liao2011, FS2012JoE, florens_simoni_2016, FlorensSimoni2015, kato2013, Shin2014, and ChibShinSimoni2018, has extended these ideas to semiparametric and likelihood-free settings.
An important question that has received little attention in the Bayesian literature concerns the testing of endogeneity. Frequentist methods, such as the classical Durbin-Wu-Hausman test, offer an asymptotic procedure that assesses exogeneity by comparing estimators that are consistent under different assumptions. These procedures, however, do not translate naturally into the Bayesian framework. From a Bayesian standpoint, it is more straightforward to conceptualize the test for endogeneity as a comparison of models, rather than that of parameters. Specifically, one can develop a test that is based on the relative support provided by the data for a model with exogeneity versus a model with endogeneity.
To develop this approach, and to avoid distributional assumptions, we proceed within a Bayesian framework for moment condition models. We consider two competing specifications. The first is a base model defined by the moment conditions \[ \mathbf{E}[{\Greekmath 0122}({\Greekmath 0112})x]=0, \quad \mathbf{E}[{\Greekmath 0122}({\Greekmath 0112})z_1]=0, \quad \mathbf{E}[{\Greekmath 0122}({\Greekmath 0112})z_2]=0, \] where ${\Greekmath 0122}({\Greekmath 0112}) = y - x^{\prime}{\Greekmath 010C} + z_{1}^{\prime}{\Greekmath 010D}$ and ${\Greekmath 0112} := ({\Greekmath 010C},{\Greekmath 010D})$. The second is an extended model that relaxes the exogeneity restriction and allows \[ \mathbf{E}[{\Greekmath 0122}({\Greekmath 0112})x] = v, \] where $v$ captures the covariance between the error term and the endogenous variable $x$. We formulate the prior-posterior analysis of each model through the nonparametric exponentially tilted empirical likelihood (ETEL) and then compare the two models by marginal likelihoods and the Bayes factor. This approach offers several methodological advantages. The Bayes factor measures the strength of evidence for the two models on a continuous scale rather than through a strict accept or reject rule. In addition, the use of ETEL provides robustness to misspecification of the joint distribution of $(y, x, z_1, z_2)$ and allows us to obtain results that remain valid without specifying the distribution of the disturbances.
Using the Chib1995 marginal likelihood identity, we know that the log marginal likelihood decomposes into three parts: the log ETEL, the log prior, and the negative log posterior ordinate. We establish that this expression is asymptotically equal to a term bounded in probability, plus a term proportional to the Kullback-Leibler divergence between the true and the closest probability distribution satisfying the moment restrictions term, plus a penalty that corresponds to the ones of the Bayesian information criterion (BIC). The penalty arises from a change-of-variable transformation of the posterior density evaluated at the true or pseudo-true value of the parameters. The log of the Jacobian of this transformation constitutes the penalty, while the posterior density of the local parameter at zero is bounded in probability as $n$ increases. Accordingly, when $x$ is exogenous, the log-ETELs of the two models are asymptotically the same but the penalties differ. When $x$ is endogenous, the difference in the log-ETELs dominates, which leads to the selection of the extended model. Thus, the test correctly discriminates between the two data-generating processes in large samples. Our construction parallels the logic of the Hausman test, where one compares an estimator that is inconsistent under endogeneity with one that is not. Here, the comparison is between models that differ in the number of overidentifying moment conditions. In this sense, our test may be viewed as the Bayesian analogue of the Hausman specification test.
Compared to ChibShinSimoni2018, our work builds on the same Bayesian ETEL framework but makes several key contributions. First, while ChibShinSimoni2018 explains how to test among different models, it does not address how to construct the specific models required to test hypotheses of interest in practical applications, such as the endogeneity problem we examine here. In this paper, we explicitly construct the models necessary for testing endogeneity. Second, we introduce an assumption that guarantees the existence of the ETEL function, which, to our knowledge, is absent from the existing ETEL literature. This assumption ensures that the ETEL function exists at least in a suitable neighborhood of the true parameter value with probability approaching one. The issue arises because the ETEL function, as the solution to a constrained optimization problem, may have an empty feasible set for certain parameter values ${\Greekmath 0112}$. Without this assumption, derivatives of the ETEL function cannot be defined, posing challenges for both frequentist and Bayesian ETEL approaches. Third, we provide a more direct proof demonstrating that the ETEL function is asymptotically equivalent to a quadratic function. This result underpins our establishment of a Bernstein-von Mises theorem, which we then use to prove the consistency of our testing procedure. The direct proof leverages the linearity in ${\Greekmath 0112}$ of the moment restrictions implied in the instrumental variable (IV) regression problem. Along the same lines, the assumptions in this paper are weaker than those in ChibShinSimoni2018, as they exploit the IV linear regression structure.\\ Finally, as a by-product of proving the consistency of our testing procedure, we derive a new asymptotic representation of the log-marginal ETEL function, defined as the ETEL function integrated with respect to the prior distribution of the model parameter. We show that the log-marginal likelihood of each model can be asymptotically decomposed into a Kullback-Leibler (KL) divergence term (between the true distribution and the closest distribution satisfying the model's moment restrictions) plus a BIC-type penalty. We derive this penalty using a novel approach: by re-expressing the posterior ordinate at the true (or pseudo-true) parameter value via a local parameter change of variables, the resulting log-Jacobian yields the penalty term, while the posterior density of the local parameter evaluated at zero is $\mathcal{O}_{p}(1)$ as $n\rightarrow \infty$. This representation clarifies the mechanics of Bayes-factor testing in this context and leads to a more transparent proof of model-selection consistency than that in ChibShinSimoni2018. Notably, we emphasize that the penalty plays a role in selecting the correct model only when $x$ is exogenous, in which case both models are correctly specified.
The rest of the paper proceeds as follows. Section (ref) summarizes Bayesian estimation and comparison of moment condition models using ETEL. Section (ref) presents the base and extended models and provides a simulated example to illustrate the procedure. Section (ref) develops the test for endogeneity and analyzes the large-sample behavior of the log-marginal likelihood, establishing consistency of the test. Section (ref) presents empirical examples, and Section (ref) concludes. An Appendix contains the proofs of the main results.
In this section we briefly provide the background on Bayesian estimation of moment condition models using the exponentially tilted empirical likelihood (ETEL). Further details can be found in Schennach2005 and ChibShinSimoni2018.
Let $w\in \mathbb{R}^{d_w}$ be a random vector, and let ${\Greekmath 0112} \in \Theta \subset \mathbb{R}^{p}$ denote a generic parameter vector. Let $\mathbb{M}$ denote the set of all probability distributions on $\mathbb{R}^{d_w}$. For a known vector of moment functions $$g(w,{\Greekmath 0112} ):\mathbb{R}^{d_w}\times \Theta\rightarrow \mathbb{R}^{d},$$ the moment restriction is given by
where $Q\in\mathbb{M}$ is a probability distribution under which the restriction is imposed and $\mathbf{E}^Q[\cdot]$ denotes the expectation under $Q$. For each ${\Greekmath 0112} \in \Theta$, define the subset of distributions that satisfy the moment restriction by
Suppose the data $w_{1:n}:=(w_{1},\ldots ,w_{n})$ are independently drawn from the true distribution $P$, which need not belong to $\mathcal{Q}({\Greekmath 0112})$ for some ${\Greekmath 0112} \in \Theta $. The expectation taken with respect to the true distribution $P$ is denoted by $\mathbf{E}[\cdot] \equiv \mathbf{E}^P[\cdot]$.
The empirical counterpart of (ref) is the weighted restriction
where $\{q_i\}_{i=1}^n$ is a discrete distribution supported on $\{w_i\}_{i=1}^n$. Equivalently, any such weight vector induces a discrete probability measure on $\mathbb{R}^{d_w}$ with support $\{w_1,\ldots,w_n\}$
where ${\Greekmath 010E}_{w_i}$ denotes the point mass at $w_i$.\\ Since there might be no ${\Greekmath 0112}\in\Theta$ such that the uniform empirical distribution $q_i\equiv 1/n$ satisfy (ref), we define the ETEL weights as the solution to the following Kullback Leibler (KL)-closest feasible reweighting problem:
which depends on the parameter vector ${\Greekmath 0112}\in\Theta$. Let $H_n\subset \Theta$ be the set of ${\Greekmath 0112}$ values for which the program (ref) is feasible, i.e. the set of ${\Greekmath 0112}$s such that the interior of the convex hull of $\{g(w_{i},{\Greekmath 0112} ): i=1,\ldots,n\}$ contains zero. Assumption (ref) below ensures that $H_n$ is non-empty with probability approaching $1$.\\ For ${\Greekmath 0112}\in H_n$, the ETEL (sample likelihood) is defined as the product of the ETEL weights:
The ETEL arises as the integrated likelihood obtained by integrating out the unknown distribution $Q$ with respect to a particular nonparametric prior that imposes the moment restrictions (ref) conditional on a ${\Greekmath 0112}\in H_n$; see Schennach2005.\\ Given a prior density ${\Greekmath 0119}({\Greekmath 0112})$, the ETEL-based posterior is the truncated posterior
where $I[A]$ denotes the indicator function. Since (ref) is not available in closed form, posterior summaries are obtained via tailored Markov chain Monte Carlo (MCMC) methods. Appendix (ref) describes computational details on MCMC sampling and related calculations.
A convenient way to compute $\{\widehat{q}_{i}({\Greekmath 0112} )\}$ is via the dual representation of (ref). Define the ETEL multiplier as, for every ${\Greekmath 0112}\in H_n$
Then, for every ${\Greekmath 0112}\in H_n$:
It is useful to view (ref) as an exponential tilting of the uniform empirical distribution $P_n(\cdot) := \frac{1}{n}\sum_{i=1}^n {\Greekmath 010E}_{w_i}(\cdot)$, which places mass $1/n$ on each observation. For fixed ${\Greekmath 0112}\in H_n$, the ETEL weights $\{\widehat q_i({\Greekmath 0112})\}$ therefore define a tilted sample distribution
that is absolutely continuous with respect to $P_n$. In particular, for each support point $w_i$,
so $\widehat Q_n(\cdot\mid{\Greekmath 0112})$ is the exponential tilt of $P_n$ that enforces the sample moment restriction (ref). The multiplier $\widehat{\Greekmath 0115}({\Greekmath 0112})$ satisfies the sample first-order condition
which is the sample analogue of (ref) under the ETEL weights. For later use, we record the log-ETEL in a form amenable to expansions. Summing $\log\widehat q_i({\Greekmath 0112})$ yields the exact identity
or equivalently,
The population counterpart of $\{\widehat{q}_{i}({\Greekmath 0112} )\}_{i=1}^n$ is the distribution $Q^*({\Greekmath 0112} )\in \mathcal{Q}({\Greekmath 0112})$ that is the closest to $P$ in the KL divergence. For each ${\Greekmath 0112}$ such that $\mathcal{Q}({\Greekmath 0112}) \neq \emptyset$, define
where $$\mathrm{KL}(Q||P):=\int \log \left( \frac{dQ}{dP}\right) dQ$$ if $Q$ is absolutely continuous with respect to $P$, and $\mathrm{KL}(Q||P) = +\infty $, otherwise. The population counterpart of the ETEL multiplier $\widehat{{\Greekmath 0115}}({\Greekmath 0112})$ is
for every ${\Greekmath 0112}\in\Theta$ such that $\mathcal{Q}({\Greekmath 0112}) \neq \emptyset$. This induces the population exponential tilt
Under mild regularity conditions, the KL projection $Q^*({\Greekmath 0112})$ admits the Radon-Nikodym derivative representation
Thus, if $P$ admits a Lebesgue density $p(w)$, then $Q^*({\Greekmath 0112})$ has density given by an exponential tilt of $p(w)$: \[ q^*(w;{\Greekmath 0112}) = p(w)\, \frac{\exp\{{\Greekmath 0115}_*({\Greekmath 0112})' g(w,{\Greekmath 0112})\}}{\mathbf{E}[\exp\{{\Greekmath 0115}_*({\Greekmath 0112})' g(w,{\Greekmath 0112})\}]} . \] By a change of measure,
The right-hand side is the population tilted moment condition. It is the moment restriction expressed under $P$ rather than under $Q^*({\Greekmath 0112})$. If one or more moment conditions are misspecified, then $Q^*({\Greekmath 0112})\neq P$ for all ${\Greekmath 0112}\in\Theta$, and the pseudo-true value ${\Greekmath 0112}_*$ is defined as the minimizer of the reversed KL divergence
where
whenever $P$ is absolutely continuous with respect to $Q^*({\Greekmath 0112})$. Under correct specification, there exists ${\Greekmath 0112}_\circ\in\Theta$ such that $P\in\mathcal{Q}({\Greekmath 0112}_\circ)$, in which case $Q^*({\Greekmath 0112}_\circ)=P$ and ${\Greekmath 0112}_* = {\Greekmath 0112}_\circ$. Moreover, in that case ${\Greekmath 0115}_*({\Greekmath 0112}_\circ)=0$ and hence $q(w;{\Greekmath 0112}_\circ)\equiv 1$. Finally, when the dual representation holds, the pseudo-true value ${\Greekmath 0112}_*$ may also be expressed in terms of the population tilt as
where the term inside the logarithm is the Radon-Nikodym derivative $[dQ^*({\Greekmath 0112})/dP](w)$ in (ref).
In this section we specialize the generic moment-restriction framework in Section (ref) to the semiparametric linear regression setting introduced in the Introduction.
Let $w:=(y,x,z_{1},z_{2})$ $\in \mathbb{R}^{d+1}$ be distributed according to an unknown probability distribution $P$, where $d := d_{x}+d_{z_{1}}+d_{z_{2}}$ and $d_w = d + 1$. Throughout, $\mathbf{E}[\cdot ]:=\mathbf{E}^{P}[\cdot]$ denotes expectation under $P$. We assume that under $P$, the random vector $w$ follows the regression model
where ${\Greekmath 0112}_{\circ }:=({\Greekmath 010C}_{\circ },{\Greekmath 010D} _{\circ })\in\Theta \subset\mathbb{R}^p$ is the true value of the regression coefficients, viewed as a functional of $P$: ${\Greekmath 0112}_{\circ} \equiv {\Greekmath 0112}_{\circ}(P)$ and $p=d_{x}+d_{z_{1}}$. In model (ref), the vector $z_{1}$ contains exogenous controls (including an intercept), and $z_{2}$ contains instrumental variables. The object of interest is the causal effect of $x$ on $y$, represented by ${\Greekmath 010C}_{\circ}$. For any ${\Greekmath 0112} := ({\Greekmath 010C}',{\Greekmath 010D}')'\in\Theta\subset\mathbb{R}^p$, with $p=d_{x}+d_{z_{1}}$ and $\widetilde{w}_1 := (x',z_1')'$, define the regression residual
If $x$ is endogenous under $P$, then $\mathbf{E}[{\Greekmath 0122}({\Greekmath 0112}_\circ)x]\neq 0$. When $d_{z_2}\geq d_x$, the instruments $z_2$ help identify ${\Greekmath 010C}_\circ$ despite endogeneity.
The base model, denoted by $\mathcal{M}_b$, imposes the moment restrictions
where the base-model moment function is
and the set of distributions satisfying the base-model restrictions is
Here, $\mathbb{M}$ denotes the set of all probability distributions on $\mathbb{R}^{d+1}$. Under exogeneity, $\mathbf{E}[{\Greekmath 0122}({\Greekmath 0112}_\circ)x]=0$ and the true distribution $P$ satisfies the base-model moments at ${\Greekmath 0112}_\circ$, so that $P\in\mathcal{Q}_b({\Greekmath 0112}_\circ)$. Under endogeneity, $\mathbf{E}[{\Greekmath 0122}({\Greekmath 0112}_\circ)x]\neq 0$ and therefore $P\notin\mathcal{Q}_b({\Greekmath 0112})$ for every ${\Greekmath 0112}\in\Theta$; in that case $\mathcal{M}_b$ is misspecified. In this case, the ETEL function, constructed from the sample $w_{1:n}$, is the empirical counterpart of the distribution $Q_{b}^{*}({\Greekmath 0112})$ that for every ${\Greekmath 0112}$ solves the moment conditions:
and that is the closest to $P$ in the KL divergence among all the distribution in the set $\mathcal{Q}_{b}({\Greekmath 0112} )$, that is,
Notice that $\mathrm{KL}(Q\Vert P)$ is set to $+\infty$ if $Q$ is not absolutely continuous with respect to $P$. In addition,
denotes the pseudo-true value in the base model. Assumption (ref) given in Section (ref) below guarantees that this value exists. On the other hand, if $x$ is exogenous, then $Q_{b}^{\ast }({\Greekmath 0112} _{\ast})=P$ and ${\Greekmath 0112}_* = {\Greekmath 0112}_\circ$, where ${\Greekmath 0112}_\circ$ denotes the true value of ${\Greekmath 0112}$ as defined above. In the following we denote the ETEL for the base model by $\widehat{q}(w_{1:n}|{\Greekmath 0112},\mathcal{M}_b) := \prod_{i=1}^n \widehat{q}_i({\Greekmath 0112}|\mathcal{M}_b)$, where $\widehat{q}_i({\Greekmath 0112}|\mathcal{M}_b)$ is constructed as in (ref) with $g(w_i,{\Greekmath 0112})$ replaced by $g_{b}(w_i,{\Greekmath 0112})$.
The extended model, denoted by $\mathcal{M}_e$, augments the base model by explicitly parameterizing the endogeneity component \[ v:=\mathbf{E}[{\Greekmath 0122}({\Greekmath 0112})x]\in\mathbb{R}^{d_x}. \] Let $\mathcal{V}\subset\mathbb{R}^{d_x}$ and define the extended parameter
The extended-model moment function is
and the model imposes the moment restrictions
where
By construction, $\mathcal{M}_e$ is correctly specified under both exogeneity and endogeneity of $x$. Indeed, let
Under (ref), we have $\mathbf{E}[{\Greekmath 0122}({\Greekmath 0112}_\circ)z_1]=0$ and $\mathbf{E}[{\Greekmath 0122}({\Greekmath 0112}_\circ)z_2]=0$, and therefore
Consequently, $P\in\mathcal{Q}_e({\Greekmath 0120}_\circ)$ and the KL projection satisfies
In the extended model, the minimizer, $Q_{e}^{\ast }({\Greekmath 0120} )=\arg\inf_{Q\in \mathcal{Q}_{e}({\Greekmath 0120})} \mathrm{KL}(Q\Vert P)$, is equal to $P$, and the population moment conditions in the extended model are
Moreover,
where ${\Greekmath 0115}_*({\Greekmath 0120} ):=\arg \min_{{\Greekmath 0115} \in \mathbb{R}^{d}}\mathbf{E}[e^{{\Greekmath 0115} ^{\prime }g_{e}(w,{\Greekmath 0120} )}]$ for every ${\Greekmath 0120}\in\mathcal{V}$ such that $\mathcal{Q}_e({\Greekmath 0120})\neq \emptyset$. In the following we denote the ETEL for the extended model by $\widehat{q}(w_{1:n}|{\Greekmath 0112},\mathcal{M}_e) := \prod_{i=1}^n \widehat{q}_i({\Greekmath 0112}|\mathcal{M}_e)$, where $\widehat{q}_i({\Greekmath 0112}|\mathcal{M}_e)$ is constructed as in (ref) with $g(w_i,{\Greekmath 0112})$ replaced by $g_{e}(w_i,{\Greekmath 0120} )$.
To illustrate the fitting of the base and extended models, consider first the base model under endogeneity. Let the data-generating process (DGP) be
for $i=1,\ldots ,n$, where $n\in \{250,500,1000,2000\}$. Suppose that the $ (u_{i},v_{i},{\Greekmath 0121}_{i})$ are marginally Gaussian, that ${\Greekmath 0122} _{i}$ is marginally a skewed Gaussian mixture $0.5\mathcal{N}(0.5,0.5^{2})+0.5 \mathcal{N}(-0.5,1.118^{2})$, that $({\Greekmath 0122} _{i},u_{i},v_{i})$ have a joint distribution induced by a Gaussian copula with covariance matrix $R=
$ and that the covariance of ${\Greekmath 0121}_{i}$ with each of the other errors is zero. Also assume that each parameter is one (except for $ {\Greekmath 010E} _{1}$, which is .5). Under this DGP, $z_{1i}$ is uncorrelated with $ {\Greekmath 0122} _{i}$ and correlated with $x_{i}$ but since ${\Greekmath 0121}_{i}$ is uncorrelated with the other shocks, $z_{2i}$ is a valid instrument that is also relevant for $x_{i} $. For each of the four sample sizes, the posterior of ${\Greekmath 0112} :=({\Greekmath 010C} ,{\Greekmath 010D} _{0},{\Greekmath 010D} _{1})$ is calculated from the four moment conditions
The ETEL posterior is sampled by the tailored one-block Metropolis-Hastings (M-H) algorithm ChibGreenberg1995 for 20,000 iterations beyond a burn-in of 1,000 cycles. The marginal posterior density of ${\Greekmath 010C} $ for each sample size is computed from these MCMC sampled draws. Kernel smoothed versions of the posterior densities are given in Figure (ref). As shown in Figure (ref), as the sample size increases, the posterior density of ${\Greekmath 010C}$ under the base model concentrates on a value quite different from the true value of ${\Greekmath 010C} $, indicating misspecification due to neglected endogeneity.
In the extended (correctly specified) model we have
The parameter of interest is now ${\Greekmath 0120} := ({\Greekmath 010C} ,{\Greekmath 010D} _{0},{\Greekmath 010D} _{1},v)$. We use a default student-t prior on $v$ centered at the Generalized Method of Moments (GMM) estimate and spread given by 4 times the GMM asymptotic variance. The prior of ${\Greekmath 0112} $ is the same as in the base model. The ETEL posterior for each of the four different sample sizes is sampled by the tailored one-block M-H method for 20,000 iterations beyond a burn-in of 1,000 cycles. The marginal posterior densities of ${\Greekmath 010C} $ are given in Figure (ref) and those of $v$ are in Figure (ref). One can see that the posterior of ${\Greekmath 010C} $, even for $n=250$, is close to the true value of ${\Greekmath 010C} $, and, for $n=2,000$, is essentially centered around the true value. In addition, the posterior of $v$, the $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{cov}(x,{\Greekmath 0122})$, tends to concentrate around the true value of 0.6.
Our Bayesian test of endogeneity is given by the Bayes factor of model $\mathcal{M}_{e}$ versus model $\mathcal{M}_{b}$ defined as:
where $m(w_{1:n}|\mathcal{M}_{b}):=\int_{H_n} \widehat{q}(w_{1:n}|{\Greekmath 0112} ,\mathcal{M}_{b}){\Greekmath 0119} ({\Greekmath 0112} )d{\Greekmath 0112} $ and $m(w_{1:n}|\mathcal{M}_{e}) := \int_{H_n} \widehat{q}(w_{1:n}|{\Greekmath 0120},\mathcal{M}_{e}){\Greekmath 0119} ({\Greekmath 0120} )d{\Greekmath 0120} $ are the model marginal likelihoods arising from the ETEL functions (also called marginal ETEL functions later on). We compute these by the method of Chib1995, as extended to general M-H chains in ChibJeliazkov2001. We select $ \mathcal{M}_{e}$ over $\mathcal{M}_{b}$ if $\log ($BF$_{eb})>0$, and select $ \mathcal{M}_{b}$ otherwise.
According to the theory in ChibShinSimoni2018, for valid comparisons of moment condition models, the contending models must arise from a common encompassing model and should have the same number of moment conditions. We have ensured that this condition is met by including the $\mathbf{E} [{\Greekmath 0122} _{i}({\Greekmath 0112} )z_{2,i}]=0$ restriction in the base model, and not excluding the $\mathbf{E}[{\Greekmath 0122} _{i}({\Greekmath 0112} )x_{i}]=v$ condition from the extended model.
Intuitively, the Bayes factor picks the correct model because $\mathcal{M} _{b}$ is correctly specified when $x$ is exogenous and misspecified when $x$ is endogenous; however, $\mathcal{M}_{e}$ is correctly specified in both the cases. Therefore, from ChibShinSimoni2018, it follows that $\mathcal{M }_{b}$, which has $(d-p)$ overidentifying restrictions, rather than $M_{e}$, which has $(d-p-d_{x})$ overidentifying restrictions, would be preferred by the Bayes factor when $x$ is exogenous (because it has more overidentifying restrictions than $\mathcal{M}_{e}$), whereas $\mathcal{M}_{e}$ would be preferred when $x$ is endogenous (because $\mathcal{M}_{b}$ in that case would be misspecified).
In this section we explain the rationale behind our testing procedure. The hypothesis that we want to test is the following: $$H_{miss}: \quad P \textrm{ is such that }\nexists {\Greekmath 0112}\in\Theta\subset\mathbb{R}^p \textrm{ such that }\mathbf{E}[{\Greekmath 0122}_{i}({\Greekmath 0112})x_{i}]= 0\qquad (endogeneity)$$ against $$H_{cs}: \quad P \textrm{ is such that }\exists {\Greekmath 0112}\in\Theta\subset\mathbb{R}^p \textrm{ such that }\mathbf{E}[{\Greekmath 0122}_{i}({\Greekmath 0112})x_{i}]= 0\qquad (exogeneity).$$ Here, the subscripts $miss$ and $cs$ are for misspecification and correct specification, respectively. The previous hypothesis can equivalently be written as $H_{miss}': v\neq 0$ and $H_{cs}': v = 0$. Our approach based on $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{BF}_{eb}$ is equivalent to a Bayes test for $H_{miss}'$ versus $H_{cs}'$ based on a mixture prior on $v$ of the type ${\Greekmath 0119}_0{\Greekmath 010E}_0(v) + (1 - {\Greekmath 0119}_0){\Greekmath 0119}(v)$, where ${\Greekmath 010E}_0(\cdot)$ denotes a Dirac mass on zero, ${\Greekmath 0119}_0\in[0,1]$, and ${\Greekmath 0119}(\cdot)$ is a continuous distribution. The two Bayes factors for these two approaches are numerically the same. The testing procedure works as follows: if $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{BF}_{eb} \geq 1$, we conclude that $x$ is endogenous (i.e. accept $H_{miss}$); if $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{BF}_{eb} < 1$, we conclude that $x$ is exogenous (i.e. accept $H_{cs}$).\\ The next theorem shows that $H_{miss}$ and $H_{cs}$ can be expressed in terms of Kullback-Leibler divergences between $P$ and the set $\mathcal{Q}_{b}({\Greekmath 0112})$ of distributions that satisfy the moment restriction that we want to test as well as additional moment restrictions that are known to hold for $P$.
This theorem makes clear that to test $H_{miss}$ and $H_{cs}$ one can equivalently focus on the Kullback-Leibler divergence $\mathrm{KL}(P||Q_b^*({\Greekmath 0112}_*))$. Our Bayes test is based on Bayes factor and comparison of marginal likelihoods. There is a strict link between marginal likelihood and the Kullback-Leibler divergence: log-marginal likelihood of the base model behaves asymptotically as $- n \mathrm{KL}(P||Q_b^*({\Greekmath 0112}_*))$ plus a penalty term, where the penalty depends on the number of parameters to estimate, and similarly for the log-marginal likelihood of the extended model. We are going to demonstrate this fact in the rest of this section.\\ From the Chib1995 identity, we have for the base model: $\forall {\Greekmath 0112}\in H_n\subset\mathbb{R}^p$, $$\log m(w_{1:n}|\mathcal{M}_b) = \log {\Greekmath 0119}({\Greekmath 0112}|\mathcal{M}_b) + \log \widehat{q}(w_{1:n}|{\Greekmath 0112},\mathcal{M}_b) - \log {\Greekmath 0119}^n({\Greekmath 0112}|w_{1:n},\mathcal{M}_b),$$ and similarly for the extended model. Because this identity is true for every ${\Greekmath 0112}\in H_n$, it is true for ${\Greekmath 0112} = {\Greekmath 0112}_*$ under Assumptions (ref) and (ref): $\log m(w_{1:n}|\mathcal{M}_b) = \log {\Greekmath 0119}({\Greekmath 0112}_*|\mathcal{M}_b) + \log \widehat{q}(w_{1:n}|{\Greekmath 0112}_*,\mathcal{M}_b) - \log {\Greekmath 0119}^n({\Greekmath 0112}_*|w_{1:n},\mathcal{M}_b)$. Next, let us introduce the local parameters $h_{{\Greekmath 0112}}:=\sqrt{n}({\Greekmath 0112} - {\Greekmath 0112}_*)$ and $h_{{\Greekmath 0120}}:=\sqrt{n}({\Greekmath 0120} - {\Greekmath 0120}_\circ)$, so that by the formula for transformations of random variables: ${\Greekmath 0119}^n({\Greekmath 0112}|w_{1:n},\mathcal{M}_b) = {\Greekmath 0119}_{h_{{\Greekmath 0112}}}^n(\sqrt{n}({\Greekmath 0112} - {\Greekmath 0112}_{*})|w_{1:n},\mathcal{M}_b)n^{p/2}$ and ${\Greekmath 0119}^n({\Greekmath 0120}|w_{1:n},\mathcal{M}_e) = {\Greekmath 0119}_{h_{{\Greekmath 0120}}}^n(\sqrt{n}({\Greekmath 0120} - {\Greekmath 0120}_{\circ})|w_{1:n},\mathcal{M}_e)n^{(p+d_x)/2}$, where ${\Greekmath 0119}_{h_{{\Greekmath 0112}}}^n(\cdot|w_{1:n},\mathcal{M}_b)$ and ${\Greekmath 0119}_{h_{{\Greekmath 0120}}}^n(\cdot|w_{1:n},\mathcal{M}_e)$ denote the posterior density of $h_{{\Greekmath 0112}}$ and $h_{{\Greekmath 0120}}$, respectively. By replacing this in the expression of the marginal likelihoods we obtain: $\forall {\Greekmath 0112}\in H_n\subset\mathbb{R}^p$,
and, $\forall {\Greekmath 0120}\in H_n\subset\mathbb{R}^{p+d_x}$,
The intuition for expressing the posterior of ${\Greekmath 0112}$ in terms of the posterior of the local parameter is that the Jacobian of the transformation makes explicit the role played by the dimension of the model, while the local parameter has a posterior distribution that is approximately Gaussian. This is true in both cases (i) and (iii) of Theorem (ref). Hence, the Jacobian induces an explicit dimension-dependent penalty through posterior concentration.\\ Therefore, the log-marginal likelihood decomposes into two terms that are bounded in probability as $n\rightarrow\infty$ (the prior ordinate evaluated at the pseudo-true value and the posterior ordinate of the local parameter) and two terms that diverge with $n$: the log-ETEL term and a model-dimension penalty of order $\tfrac{1}{2}\log n$ per parameter. Asymptotically, the marginal likelihood behaves like a penalized log-ETEL criterion, where the penalty arises endogenously from posterior concentration via the local reparameterization, rather than being imposed ad hoc.\\
Of course, for a testing procedure based on marginal likelihoods to be valid, it is necessary to establish that ${\Greekmath 0119}_{h_{{\Greekmath 0112}}}^n(\sqrt{n}({\Greekmath 0112} - {\Greekmath 0112}_{*}) \mid w_{1:n},\mathcal{M}_b)$ and ${\Greekmath 0119}_{h_{{\Greekmath 0120}}}^n(\sqrt{n}({\Greekmath 0120} - {\Greekmath 0120}_{\circ}) \mid w_{1:n},\mathcal{M}_e)$ are bounded in probability as $n \rightarrow \infty$. This requirement can be quite challenging to verify, particularly in non-standard settings such as the one considered here, where there is no parametric likelihood and the models may be misspecified. We establish these results in Theorems (ref) and (ref) in the Online Appendix, which refine Theorems 1 and 2 of ChibShinSimoni2018.
A critical step in proving these results is to show that the log-ETEL function satisfies a stochastic local asymptotic normality (LAN) property. While the remainder of the Bernstein--von Mises argument follows standard lines, establishing stochastic LAN is challenging because the ETEL function is itself a random likelihood. In this paper, we provide a new and more direct proof of the LAN property for the log-ETEL function (see Theorems (ref), (ref), and (ref) in the Online Appendix). Our proof leverages the specific structure of the IV regression problem: due to linearity, each term in the Mean Value Theorem expansion of the log-ETEL function around ${\Greekmath 0112}_*$ can be controlled more directly and uniformly in $h_{{\Greekmath 0112}}$ over compact sets. This allows us to avoid the empirical process theory used in ChibShinSimoni2018.
The final step toward understanding the asymptotic behavior of the marginal likelihood is provided by Theorems (ref) and (ref), which derive stochastic expansions of the log-ETEL function in the base and extended models. These expansions, which were not made explicit in ChibShinSimoni2018, are new to the best of our knowledge.
Our starting point is the exact master identity for the log-ETEL, (ref). Evaluating this identity at ${\Greekmath 0112}={\Greekmath 0112}_*$ and writing $\widehat{\Greekmath 0115}({\Greekmath 0112}_*)={\Greekmath 0115}_*({\Greekmath 0112}_*)+ (\widehat{\Greekmath 0115}({\Greekmath 0112}_*)-{\Greekmath 0115}_*({\Greekmath 0112}_*))$, we obtain the expansion by performing a second-order Taylor expansion in ${\Greekmath 0115}$ around ${\Greekmath 0115}_*({\Greekmath 0112}_*)$. The stochastic LAN property delivers a linear representation for $\sqrt n(\widehat{\Greekmath 0115}({\Greekmath 0112}_*)-{\Greekmath 0115}_*({\Greekmath 0112}_*))$, which, when substituted back into the master identity, yields the quadratic empirical-process terms reported below.
The assumptions under which the results hold are collected in Section (ref). We use the notation $\mathbf{E}_n[\cdot]=n^{-1}\sum_{i=1}^n(\cdot)$ for the empirical mean and $\mathbb G_n f=\sqrt n(\mathbf{E}_n[f]-\mathbf{E}[f])$ for the centered empirical process.
This theorem establishes a decomposition of the log-ETEL function for the base model $\mathcal{M}_b$, characterizing its asymptotic behaviour. This decomposition is used to prove the consistency of our Bayes factor testing procedure. Theorem (ref) states that $\log \widehat{q}(w_{1:n}|{\Greekmath 0112}_*,\mathcal{M}_{b}) + n\log n$ can be expressed, up to an $o_p(1)$ term, as the sum of four random components. The third and fifth terms on the right-hand side of equation (ref) are both of order $\mathcal{O}_p(1)$, while the second and fourth terms are of order $\mathcal{O}_p(n)$ and $\mathcal{O}_p(\sqrt{n})$, respectively, when $x_i$ is endogenous, and equal to zero when $x_i$ is exogenous. The fourth term is linear, with its rate $\mathcal{O}_p(\sqrt{n})$ following from the last part of the theorem, whereas the fifth term in (ref) is quadratic. The term $n\log n$ does not affect the comparison, as it cancels out with the corresponding term in the log-ETEL function of the extended model, as shown in the next theorem.\\ For the extended model, we recall that ${\Greekmath 0120}_\circ = ({\Greekmath 0112}_\circ',v_\circ')'$ denotes the true value of the parameter in the extended model with $v_{\circ} = \mathbf{E}[{\Greekmath 0122}({\Greekmath 0112}_\circ)x]$.
Theorem (ref) establishes the asymptotic behaviour of the log-ETEL function of the extended model $\mathcal{M}_e$. Unlike the base model, $\log \widehat{q}(w_{1:n}|{\Greekmath 0120}_\circ,\mathcal{M}_{e}) + n\log n$ is equal, up to an asymptotically negligible term, to a quadratic random term that remains bounded in probability as $n\rightarrow \infty$.\\ If $\mathbf{E}[{\Greekmath 0122}_i({\Greekmath 0112}_\circ) x_i] = 0$ (exogenous case), so that the assumptions in Theorem (ref) hold with ${\Greekmath 0112}_*$ replaced by ${\Greekmath 0112}_\circ$ and ${\Greekmath 0115}_*({\Greekmath 0112}_*) = {\Greekmath 0115}_*({\Greekmath 0112}_\circ) = 0$, then the log-ETEL function in the base model simplifies as
where $\Omega_\circ := \mathbf{E}[{\Greekmath 0122}_i({\Greekmath 0112}_\circ)^2 \widetilde{w}_i\widetilde{w}_i']$, $\mathbb{G}_n\left[g_b(w_i,{\Greekmath 0112}_\circ)'\right] \Omega_\circ^{-1} \mathbb{G}_n\left[g_b(w_i,{\Greekmath 0112}_\circ)\right]\xrightarrow{d}{\Greekmath 011F}_{d}^2$, and ${\Greekmath 011F}_{d}^2$ denotes a chi square distribution with $d$ degrees of freedom. For the extended model, if $\mathbf{E}[{\Greekmath 0122}_i({\Greekmath 0112}_\circ) x_i] = 0$ then $g_e(w_i,{\Greekmath 0120}_\circ) = g_b(w_i,{\Greekmath 0112}_\circ)$ and the log-ETEL function slightly simplifies as:
where $\mathbb{G}_n\left[g_b(w_i,{\Greekmath 0120}_\circ)'\right] \Omega_{{\Greekmath 0120}_\circ}^{-1} \mathbb{G}_n\left[g_b(w_i,{\Greekmath 0120}_\circ)\right]\xrightarrow{d}{\Greekmath 011F}_{d}^2$. Hence, when $x$ is exogenous, $\log \widehat{q}(w_{1:n}|{\Greekmath 0112}_\circ,\mathcal{M}_{b}) $ and $\log \widehat{q}(w_{1:n}|{\Greekmath 0120}_\circ,\mathcal{M}_{e})$ are equal asymptotically and they cancel in the comparison of the marginal likelihoods.\\ In case of endogeneity, instead, $\log \widehat{q}(w_{1:n}|{\Greekmath 0112}_\circ,\mathcal{M}_{b}) $ and $\log \widehat{q}(w_{1:n}|{\Greekmath 0120}_\circ,\mathcal{M}_{e})$ are different and they play a central role in the comparison of marginal likelihoods. In this case, it is important to consider the behaviour of the average log-ETEL function $\frac{1}{n}\log \widehat{q}(w_{1:n}|{\Greekmath 0112}_*,\mathcal{M}_b)$ which stays bounded asymptotically. The following two corollaries demonstrates that asymptotically the average log-ETEL functions behave as the Kullback-Leibler divergence, up to a $\log(n)$ term. While these results are implicit in the definition of the ETEL, we provide here a formal and explicit statement and in the Appendix their proof. This allows us to understand the behaviour of the marginal likelihood.
Notice that $\mathbf{E}\left[\log(dQ_e^*({\Greekmath 0120}_\circ)/dP)\right] = 0$ since the extended model is correctly specified and so $dQ_e^*({\Greekmath 0120}_\circ)/dP = 1$. \\
From Theorems (ref) and (ref), and Theorems (ref) and (ref) in the Online Appendix and from (ref)-(ref), there exists an $N$ such that for every $n>N$:
These log-marginal likelihoods quantify the overall validity of the model. In fact, from these expressions one sees that when $x_i$ is exogenous, that is, $\mathbf{E}[{\Greekmath 0122}_i({\Greekmath 0112}_\circ) x_i] = 0$, then ${\Greekmath 0115}_*({\Greekmath 0112}_*) = 0$ and $\sum_{i=1}^n \log\left(\frac{e^{{\Greekmath 0115}_*({\Greekmath 0112}_*)'g_b(w_i,{\Greekmath 0112}_*)}}{\mathbf{E}_n[e^{{\Greekmath 0115}_*({\Greekmath 0112}_*)' g_b(w_j,{\Greekmath 0112}_*)}]}\right) = 0$ for every $n\in\mathbb{N}$. Therefore, it is clear that asymptotically $\log m(w_{1:n}|\mathcal{M}_b)$ is larger than $\log m(w_{1:n}|\mathcal{M}_e)$.\\ On the other hand, when there is no ${\Greekmath 0112}\in\Theta$ such that $\mathbf{E}[{\Greekmath 0122}_i({\Greekmath 0112}) x_i] = 0$, then ${\Greekmath 0115}_*({\Greekmath 0112}_*) \neq 0$ and $\sum_{i=1}^n \log\left(\frac{e^{{\Greekmath 0115}_*({\Greekmath 0112}_*)'g_b(w_i,{\Greekmath 0112}_*)}}{\mathbf{E}_n[e^{{\Greekmath 0115}_*({\Greekmath 0112}_*)' g_b(w_j,{\Greekmath 0112}_*)}]}\right) - n\log n$ diverges to $-\infty$ faster than the last two terms in (ref), so that asymptotically $\log m(w_{1:n}|\mathcal{M}_b)$ is smaller than $\log m(w_{1:n}|\mathcal{M}_e)$. This is the main intuition of the consistency results in Theorems (ref) and (ref) in the next section. The proof of these theorems, which is provided in the Appendix, is more involved than this argument because the theorems provide an `if and only if' statement, which is stronger than consistency.
We now use the preceding theory to establish consistency of our testing procedure based on the Bayes factor constructed from the marginal ETEL functions. The theorems below establish that, as the sample size increases, $BF_{eb}$ selects $\mathcal{M}_{b}$ if and only if $x$ is exogenous, and selects $\mathcal{M}_{e}$ if and only if $x$ is endogenous, with probability approaching one.
As we show in the proof, the failure of the necessary and sufficient condition $\mathbf{E}[{\Greekmath 0122} _{i}({\Greekmath 0112} )x_{i}]=0$ for any ${\Greekmath 0112} \in\Theta$, is equivalent to the inequality $\mathrm{KL}(P||Q_{e}^{\ast }({\Greekmath 0120} ))<\mathrm{KL}(P||Q_{b}^{\ast }({\Greekmath 0112} ))$, where $\mathrm{KL}(P||Q_{e}^{\ast }({\Greekmath 0120}_\circ)) = 0$. In this case, the log-ETEL function dominates the other components of the log-marginal ETEL function so that the build-in penalty does not play any role. Thus, as in the general result in ChibShinSimoni2018 for moment condition models, comparing the log-marginal likelihoods of the base and extended models, and selecting the one with the higher value, in the limit, selects the model that is closest in the KL divergence to the true model. In the framework of the present paper, this means that the comparison of marginal likelihoods allows to correctly conclude that $x_i$ is endogenous.\\ Next, we show what happens when the variables $x_i$ are exogenous so that the moment restriction $\mathbf{E}[{\Greekmath 0122} _{i}({\Greekmath 0112} )x_{i}]=0$ holds for a particular value ${\Greekmath 0112}_\circ$ and the two models under comparison are correctly specified. The next theorem states that in this case the base model is selected. This is understandable through an argument of parsimony: the base model has the smaller number of parameters to estimate and so it is the preferred one when it is correctly specified.
When $x$ is exogenous, as in the previous theorem, both the log-ETEL function and the build-in penalty plays a role in selecting the correct model.
\paragraph{Discussion.} In this and the previous subsection, we demonstrate that our model selection criteria favor a model with a smaller Kullback-Leibler Information Criterion (KLIC). When two models share the same KLIC, our procedure opts for the model with a greater number of overidentifying restrictions, i.e., a more parsimonious or less flexible model. Interestingly, this aligns with the goal of sin1996information's penalized likelihood criteria for a parametric model. Consequently, our proposed model selection procedure in this paper and ChibShinSimoni2018 can be viewed as a fully Bayesian semi-parametric version of consistent model selection criteria, applied specifically to an endogeneity testing problem. Unlike other frequentist procedures, the `penalty' term required for consistency is inherently built into our Bayesian calculation. This point was not stressed in ChibShinSimoni2018 and it is a contribution of this paper.
Andrews1999, AndrewsLu2001, and HongPrestonShum2003 have proposed and studied model selection criteria for moment condition models, even though a formal likelihood function is not defined. These criteria involve a penalization term that is attached to the Generalized Method of Moments (GMM) and, more broadly, the Generalized Empirical Likelihood (GEL) objective function, rather than the likelihood function. Examples of such frequentist model selection approaches based on GMM estimation can be found in Online Appendix (ref). However, the relationship between these model selection criteria and the KLIC minimization principle of sin1996information for potentially misspecified parametric models is not immediately apparent.
It is noteworthy that our procedure exhibits the same asymptotic behavior as hong2012bayesian's generalized empirical likelihood Bayes factor. They impose a separate prior on the Lagrangian multiplier that is independent of ${\Greekmath 0112}$, which does not guarantee the imposition of moment restrictions. In contrast, we introduce an additional parameter $v$ to the `inactive' moment restriction, ensuring that our prior on ${\Greekmath 0112}$ and $v$ respects the moment restrictions.
Our testing procedure can be extended to settings in which more than two models are compared. Consider the case in which only a subset of the variables in $x$ is endogenous. To start, suppose that $d_x = 2$ and that only $x_1$ is endogenous, whereas $x_2$ is exogenous. That is, $\mathbf{E}[{\Greekmath 0122}({\Greekmath 0112}_{\circ})(x_2,z_1')']=0$, while there exists no ${\Greekmath 0112}\in\mathbb{R}^p$ such that $\mathbf{E}[{\Greekmath 0122}({\Greekmath 0112})x_1]=0$. If we compare only the base and extended models, $\mathcal{M}_b$ and $\mathcal{M}_e$, we could erroneously conclude that $x_1$ and $x_2$ are both endogenous. Instead, it is more appropriate to consider the following models: $\mathcal{M}_b$, $\mathcal{M}_e$, $\mathcal{M}_{e_1}$ and $\mathcal{M}_{e_2}$, where, for $i=1,2$, $\mathcal{M}_{e_i}$ is the model defined by the moment condition
where
${\Greekmath 0120}_i :=({\Greekmath 0112} ,v^{(i)})\in \Psi $, $\Psi :=\Theta \times \mathcal{V}$, $\mathcal{V}\subset \mathbb{R}^{2}$ with $v^{(i)} = (v_1^{(i)}, v_2^{(i)})$ and $v_j^{(i)}\in\mathbb{R}$ for $j=1,2$, and $\mathcal{Q}_{e_i}({\Greekmath 0120}_i) := \left\{Q\in \mathbb{M};\,\mathbf{E}^{Q}[g_{e_i}(w,{\Greekmath 0120}_i)] = 0\right\} $. We enforce the restriction that one component of $x$ is treated as exogenous by setting the corresponding element of $v^{(i)}$ to zero. Specifically, define $v^{(1)} := (0,v_{2}^{(1)})$ and $v^{(2)} := (v_{1}^{(2)},0)$. Model $\mathcal{M}_{e_{1}}$ treats $x_{1}$ as exogenous and allows $x_2$ to be endogenous, while model $\mathcal{M}_{e_{2}}$ treats $x_{2}$ as exogenous and allows $x_1$ to be endogenous.
This construction allows a direct application of our baseline-versus-extended comparison. Treat $\mathcal{M}_e$ as the common reference extended model and compare $\mathcal{M}_{e_i}$ against $\mathcal{M}_e$. If the marginal likelihood of $\mathcal{M}_{e_2}$ exceeds that of $\mathcal{M}_e$, then the data support the restriction $v_2=0$, suggesting that $x_2$ is exogenous while $x_1$ is treated as endogenous. Similarly, if the marginal likelihood of $\mathcal{M}_{e_1}$ exceeds that of $\mathcal{M}_e$, then the data support $v_1=0$, suggesting that $x_1$ is exogenous while $x_2$ is treated as endogenous.
More generally, when more than two models are under consideration, one can compare them via their marginal likelihoods. In our context, these candidate models are obtained from the extended model by setting a subset of the elements of $v$ to zero. In total, there are $2^{d_x}$ models, including the base and extended models. Each model corresponds to a configuration of which elements of $x$ are treated as endogenous: $x_i$ is treated as endogenous when the associated $v_i$ is unrestricted, whereas $v_i=0$ corresponds to $x_i$ being treated as exogenous. Marginal-likelihood comparison over this finite model set selects the configuration most supported by the data.
Moreover, our endogeneity testing can be enriched by comparing different model specifications. For example, suppose $x$ is scalar and consider linear versus quadratic specifications, \[ y_{i} = {\Greekmath 010D}_{0} + {\Greekmath 010C}_{1} x_{i} + {\Greekmath 010C}_{2} x_{i}^{2} + {\Greekmath 010D}_{1} z_{1i} + {\Greekmath 0122}_{i}. \] Then endogeneity of $x_i$ can be assessed under each specification, leading to four candidate models, {linear-exogenous, linear-endogenous, quadratic-exogenous, quadratic-endogenous}. A marginal likelihood comparison can be used to select the best model among these candidates. We apply this idea in our real data example (BLP model). In that setting, we consider four candidate models that jointly vary the functional form and the endogeneity status of price, namely {linear-exogenous, linear-endogenous, nonlinear-exogenous, nonlinear-endogenous}. Marginal likelihood comparison over these four candidates simultaneously assesses endogeneity within each specification and delivers a unified ranking across specifications.
Appendix (ref) reports Monte Carlo experiments for these two use cases; see Sections (ref) and (ref).
We provide the assumptions that we use to prove the results in the previous sections. The first assumption ensures that the set of distributions satisfying the moment conditions is non-empty, which is necessary for the ETEL to be well-defined. It guarantees that the dual representation of the optimization problem (ref) holds even when $P\notin\mathcal{Q}_{b,{\Greekmath 0112}}$ for every ${\Greekmath 0112}\in\Theta$. In fact, in the latter case it is possible that $Q_b^*({\Greekmath 0112} )$ and $P$ do not have a common support for any ${\Greekmath 0112} $, in which case, the equality in (ref) does not hold; see Sueishi2013 for a discussion on this point.
This assumption implies that there is a ${\Greekmath 0112} $ for which $\mathcal{Q}_{b,{\Greekmath 0112} }$ is non-empty, that $dQ_b^{*}({\Greekmath 0112} )/dP = \Bigl(\frac{e^{{\Greekmath 0115}_{*}({\Greekmath 0112} )'g(w,{\Greekmath 0112})}}{\mathbf{E}[e^{{\Greekmath 0115}_*({\Greekmath 0112} )'g(w,{\Greekmath 0112} )}]}\Bigr)$ and that ${\Greekmath 0112}_{\ast }$ is identified by (ref). In addition, by assuming mutual absolute continuity, it ensures that both $\mathrm{KL}(Q_b^{*}({\Greekmath 0112})||P)$ and $\mathrm{KL}(P||Q_b^{*}({\Greekmath 0112}))$ are well defined. We then assume that ${\Greekmath 0112}_*$ is unique.
Since under Assumption (ref), ${\Greekmath 0112}_*$ coincides with the minimizer in (ref), then the previous assumption implies uniqueness also of the latter.\\ As we have pointed out in Section (ref), the ETEL is not defined at the ${\Greekmath 0112}$s for which the optimization problem (ref) does not have a feasible solution, that is, at the ${\Greekmath 0112}$s that do not belong to the set $H_n$ in (ref). Assumption (ref) guarantees that asymptotically the optimization problem (ref) is feasible at ${\Greekmath 0112}_*$. Similarly, if the model is correctly specified, then the optimization problem (ref) is feasible at ${\Greekmath 0112}_{\circ}$ (or ${\Greekmath 0120}_{\circ}$ depending on which model we consider). However, this is not enough for our asymptotic analysis. Instead, we need to assume that (ref) has a solution for every ${\Greekmath 0112}$ that is sufficiently close to ${\Greekmath 0112}_*$ with probability approaching $1$. The required notion of how close is specified in the next assumption, for which we introduce the following ball. For any sequence $M_n\rightarrow \infty$ as $n\rightarrow \infty$, and any $\widetilde{\Greekmath 0112}\in\Theta$, define the $n^{-1/2}$-ball around $\widetilde{\Greekmath 0112}$ as $B(\widetilde{\Greekmath 0112},M_n n^{-1/2}) := \{{\Greekmath 0112}\in\Theta:\,\|{\Greekmath 0112} - \widetilde{\Greekmath 0112}\| \leq M_n n^{-1/2}\}.$ The ball $B(\widetilde{\Greekmath 0112},M_n n^{-1/2})$ shrinks to $0$ slightly slower than $n^{-1/2}$ depending on how fast $M_n$ goes to $\infty$. The $n^{-1/2}$-ball $B(\widetilde{\Greekmath 0120},M_n n^{-1/2})$ around some $\widetilde{\Greekmath 0120}\in\Psi$ is defined similarly. In addition, denote by $int \Delta_n := \{q := (q_1,\ldots,q_n)'; \sum_{i=1}^n q_i = 1, \; q_1>0,\ldots,q_n>0\}$ the interior of the $n-1$ simplex. Finally, for any sequence $M_n\rightarrow \infty$: $B_{*,n} := \bigcap_{\{M_n; M_n\rightarrow \infty\} }B({\Greekmath 0112}_*, M_n n^{-1/2})$ and $B_{\circ,n} := \bigcap_{\{M_n; M_n\rightarrow \infty\} }B({\Greekmath 0120}_{\circ},M_n n^{-1/2})$.
This assumption is much weaker than requiring that (ref) has a solution for every ${\Greekmath 0112}\in\Theta$ with probability approaching $1$. Even if not explicitly stated in the literature, a similar assumption is necessary for the frequentist ETEL, EL and ET estimators.
The next fourth assumptions concern the model. Compared to the assumptions in ChibShinSimoni2018, our assumptions are weaker due to the linearity in ${\Greekmath 0112}$ of the moment functions $g_b(w,{\Greekmath 0112})$ and $g_e(w,{\Greekmath 0112})$. Consequently, assumptions regarding the continuity of the moment functions and their derivatives are automatically satisfied. Recall the notation $w_i:= (y_i, x_i', z_i')'$, $z_i := (z_{1,i}',z_{2,i}')'$ and $\tilde w_{1,i} := (x_i', z_{1,i}')'$. Moreover, we denote $\tilde w_i := (x_i', z_i')'$, $\|\cdot\|_2$ the Euclidean norm and $\|\cdot\|_{op}$ the operator norm.
Assumptions (ref) and (ref) are standard in the literature, see e.g. Schennach2007. The following assumption instead is new and it is used to prove the approximation for the marginal likelihood. We denote by $\widetilde{w}_{i,k}$ the $k$-th component of $\widetilde{w}_i$. Moreover, for any ${\Greekmath 010E}>0$ and for some constant $C>0$, we denote by $B_{{\Greekmath 010E}}({\Greekmath 0115}_*) := \{{\Greekmath 0115}\in\mathbb{R}^d; \|{\Greekmath 0115} - {\Greekmath 0115}_*({\Greekmath 0112}_*)\|_2 \leq C{\Greekmath 010E}\}$ (resp. $B_{{\Greekmath 010E}}({\Greekmath 0112}_*) := \{{\Greekmath 0112}\in\mathbb{R}^p; \|{\Greekmath 0112} - {\Greekmath 0112}_*\|_2 \leq C{\Greekmath 010E}\}$) a closed ball in $\mathbb{R}^d$ (resp. in $\mathbb{R}^p$) centered around ${\Greekmath 0115}_* := {\Greekmath 0115}_*({\Greekmath 0112}_*)$ (resp. ${\Greekmath 0112}_*$) with radius ${\Greekmath 010E}$, where $\|\cdot\|_2$ denotes the Euclidean norm. For any ${\Greekmath 010E} >0$, $q\in\mathbb{N}_+$, a set $B$, and a function ${\Greekmath 010D}:\mathbb{R}^{d+1}\times \mathbb{R}^p\rightarrow \mathbb{R}$, we introduce the following set of functions: $\mathcal{F}_{q,{\Greekmath 010E}}(B;{\Greekmath 010D}) := \{f:B\times \mathbb{R}^{d+1}\times B_{{\Greekmath 010E}}({\Greekmath 0112}_*) \rightarrow \mathbb{R}^q; \|g(u,v,{\Greekmath 0112})\|_2 \leq {\Greekmath 010D}(v,{\Greekmath 0112}),\, \forall u\in B\}$. For a random vector $w$, we denote with $w_k$ its $k$-th component. Finally, $K_n\subset\mathbb{R}^p$ denotes a compact set centred on $0$.
This regularity assumption is a weak moment condition, requiring the existence of an upper bound -- dependent on the data and on ${\Greekmath 0112}$ -- with finite expectation for every ${\Greekmath 0112}\in B_{{\Greekmath 010E}}({\Greekmath 0112}_*)$. This condition ensures that a uniform Law of Large Number holds (see e.g. NeweyMcFadden1994). In particular, we use this assumption in our proofs to establish the uniform convergence in probability of several terms, including: $\mathbf{E}_n\left[e^{\widehat{\Greekmath 0115}'g(w_i,{\Greekmath 0112}_*)}{\Greekmath 0122}_i({\Greekmath 0112}_*)\widetilde{w}_{i}\right]$, $\mathbf{E}_n\left[e^{\widehat{\Greekmath 0115}'g(w_i,{\Greekmath 0112}_*)}\widetilde{w}_{1,i}\widetilde{w}_{i}\right]$, $h'\mathbf{E}_n\left[e^{\widehat{\Greekmath 0115}'g(w_i,{\Greekmath 0112}_*)} h'\widetilde{w}_{1,i}\widetilde{w}_{i}'\widehat{\Greekmath 0115}{\Greekmath 0122}_i({\Greekmath 0112}_*)\widetilde{w}_i\right]$, and $\mathbf{E}_n\left[e^{\widehat{\Greekmath 0115}'g(w_i,{\Greekmath 0112}_*)}{\Greekmath 0122}_i({\Greekmath 0112}_*)\widetilde{w}_{1,i}\widetilde{w}_{i}\right]$. The uniform convergence must hold uniformly over $\widehat{\Greekmath 0115}$ in a closed ball $B_{{\Greekmath 010E}}({\Greekmath 0115}_*({\Greekmath 0112}_*))$. Additionally, this assumption also guarantees that the dominance condition required for the application of the Dominated Convergence Theorem is satisfied, which is another result used in the proofs.
Part (e) of the previous assumption implies that $$ \mathbf{E}\left[\sup_{{\Greekmath 0115}\in B_{{\Greekmath 010E}}({\Greekmath 0115}_*({\Greekmath 0112}_*))}\left\|e^{\ell {\Greekmath 0115}'\widetilde w {\Greekmath 0122}({\Greekmath 0112}_*)}{\Greekmath 0122}({\Greekmath 0112}_{*})^{\ell '}\widetilde w_{k}^{i} \widetilde w\tilde w'\right\|_{op}\right] $$ is bounded away from infinity for every $k'=\{1\,\ldots,d\}$ which is what we need in the proof because this operator norm is upper bounded by $d\,\max_{\{j,k\in\1,\ldots,d\}}|e^{\ell {\Greekmath 0115}'g(w,{\Greekmath 0112}_*)}\widetilde w_{j} \widetilde w_{k}{\Greekmath 0122}({\Greekmath 0112}_*)^{\ell} \widetilde w_{k'}^{i} |$ for every $k'\in\{1,\ldots, d\}$.\\ An assumption similar to Assumption (ref) is necessary to establish uniform convergences of $\mathbf{E}_n\left[{\Greekmath 0122}_i({\Greekmath 0112} )\widetilde{w}_{i}\right]$ and $\mathbf{E}_n\left[{\Greekmath 0122}_i({\Greekmath 0112})\widetilde{w}_{i} \widetilde{w_i}\right]$ uniformly over ${\Greekmath 0112}\in B_{{\Greekmath 010E}}({\Greekmath 0112}_*)$. For this, we introduce the class
For the next assumption we denote by $\Theta_{n}:=\{\Vert {\Greekmath 0112} -{\Greekmath 0112} _{\ast }\Vert \leq M_{n}/\sqrt{n}\}$ a ball around ${\Greekmath 0112}_{*}$ with the radius at most $M_{n}/\sqrt{n}$, where $M_{n}$ is any sequence of positive constants diverging to $+\infty $. We denote by $\ell_{n,{\Greekmath 0112}}(w_i)$ the log-likelihood function for one observation $w_i$: $\ell_{n,{\Greekmath 0112}}(w_i) := \log \widehat{q}_i({\Greekmath 0112}|\mathcal{M}_b)$, and by $\ell_{n,{\Greekmath 0112}}(w_{1:n}) := \sum_{i=1}^n \ell_{n,{\Greekmath 0112}}(w_i) = \log \widehat{q}(w_{1:n}|{\Greekmath 0112},\mathcal{M}_b)$ the log-ETEL function. We recall that both $\ell_{n,{\Greekmath 0112}}(w_i)$ and $\ell_{n,{\Greekmath 0112}}(w_{1:n})$ are defined only for ${\Greekmath 0112}\in H_n$ and that under Assumption (ref) they are at least defined on $B_{*,n}$. The next assumption controls the behaviour of the ETEL function ${\Greekmath 0112}\mapsto \ell_{n,{\Greekmath 0112} }(w_{i})$ at a distance from ${\Greekmath 0112}_*$ and it ensures that ${\Greekmath 0112}_*$ is well-separated from the ${\Greekmath 0112}$s that are at a certain distance from it.
A condition similar to Assumption (ref) is in kleijn2012 and it is also related to the classical condition in e.g. LehmanCasella1998 and ChernozhukovHong2003. However, in our case the supremum in the assumption is taken over a smaller set, which is $H_n\cap\Theta_n^c$, instead of over $\Theta_n^c$ as in the mentioned literature. To better understand the meaning of this assumption, note that asymptotically the log-ETEL function is maximized at the pseudo-true value ${\Greekmath 0112}_*$. Hence, Assumption (ref) requires that if the parameter ${\Greekmath 0112} $ is far from the pseudo-true value ${\Greekmath 0112}_{*}$, that is $\|{\Greekmath 0112} -{\Greekmath 0112}_* \| > M_{n}/\sqrt{n}$, then the sum $\sum_{i=1}^{n} \ell_{n,{\Greekmath 0112} }(w_{i})$ evaluated at such a ${\Greekmath 0112} $ must be small relative to the sum $\sum_{i=1}^{n}\ell _{n,{\Greekmath 0112}_*}(w_{i})$, which is the maximum value. Controlling this behavior is important because the posterior involves integration over the whole support of ${\Greekmath 0112}$. Subsets of $\Theta$ that can be distinguished from ${\Greekmath 0112}_*$ uniformly (with probability approaching $1$ as $n\rightarrow \infty$) based on the ETEL function will receive a posterior probability that is asymptotically negligible. An alternative to this condition would be to require the existence of asymptotically consistent tests ${\Greekmath 011E}_n$ that are able to distinguish from the true distribution $P$ in a uniform way, that is, for every ${\Greekmath 010F}>0$, there exists a sequence of tests $\{{\Greekmath 011E}_n\}$ such that as $n\rightarrow$ 0,
Similarly, for the extended model we denote by $\ell_{n,{\Greekmath 0120}}(w_i)$ the log-likelihood function for one observation $w_i$: $\ell_{n,{\Greekmath 0120}}(w_i) := \log \widehat{q}_i({\Greekmath 0120}|\mathcal{M}_e)$ and by $\ell_{n,{\Greekmath 0120}}(w_{1:n}) := \sum_{i=1}^n \ell_{n,{\Greekmath 0120}}(w_i) = \log \widehat{q}(w_{1:n}|{\Greekmath 0120},\mathcal{M}_e)$ the log-ETEL function. The next assumption has the same interpretation of Assumption (ref) but for the extended model.
Consider the same generating process as in Section (ref), and suppose that $({\Greekmath 0122} _{i},u_{i},v_{i})$ have a joint distribution induced by a Gaussian copula with covariance matrix $R=
$. The parameter ${\Greekmath 011A} $ controls the degree of endogeneity. We let ${\Greekmath 011A} $ take values in the set from -.5 to .5, in increments of 0.1. For each value of ${\Greekmath 011A} $ in this set, we generate 100 samples of size $n$. For each sample, we compute the base and extended models, and calculate the log-marginal likelihoods. We then count the number of times the log-marginal likelihood of $\mathcal{M}_{e}$ exceeds that of $ \mathcal{M}_{b}$. The results are given Table \ref{tab:sim_MegreaterthanMb}. We can see from this table that even for small values of ${\Greekmath 011A} $, our test of endogeneity correctly concludes that the correct model is $\mathcal{M} _{e} $.
We consider the classic problem of automobile demand studied in berry1995econometrica. This problem has recently been revisited by chernozhukov2015post, henceforth BLP and CHS, respectively. Apart from its intrinsic value, this problem is worth analyzing because it involves a realistically large number of controls and instruments.
To set up the problem, let $y_{ijt}$ denote the log of the ratio of the market share of product $i$ in market $j$ at time $t$, relative to an external option, and let $x_{ijt}$ denote the potentially endogenous automobile price variable. In the sample data, this variable is demeaned. For controls, let $z_{ijt}$ denote the observed characteristics of the product. In BLP these are taken to be a constant, an air conditioning dummy ( $air$), horsepower divided by weight ($hpwt$), miles per dollar ($mpd$), and vehicle size ($space$). In our notation, $y_{ijt} = x_{ijt} \, {\Greekmath 010C} + {z} _{1ijt}^{\prime }{{\Greekmath 010D} }+ {\Greekmath 0122} _{i}$, where ${z}_{1ijt}=(1, mpd_{ijt}, space_{ijt}, hpwt_{ijt}, air_{ijt})$. BLP used ten instruments, five formed by summing the value of these five characteristics over other automobiles produced by the same firm and five formed by summing the above characteristics over automobiles produced by other firms. These form $ z_{2ijt}$. In revisiting this analysis, CHS augment the original controls with quadratics and cubics in $trend$, $mpd$, $space$, $hpwt$, and all first order interactions, and then used sums of these characteristics as potential instruments.
In our analysis, we consider both formulations, but in the augmented variant we introduce nonlinear controls by transforming each of $trend$, $hpwt$, $ mpd $ and $space$ by natural cubic spline basis functions, each centered at five equally spaced quantile knots (the cubic spline basis functions are taken from Chib2010). We opt for this approach to avoid widely different covariate values from parametric quadratic and cubic terms of these covariates. After the imposition of an identification restriction on the basis expansions, which reduces the number of nonlinear terms to four for each continuous covariate, the right hand side of the augmented outcome model is defined by $x$ (price) and $z_1$ (consisting of an intercept, sixteen nonlinear covariates denoted by $trend{_Bj}$, $mpd_{Bj}$, $space_{Bj}$ and $ hpwt_{Bj}$ for $j = 1,\ldots,4$, and the air-conditioning dummy). The set of augmented instruments that form $z_2$ in this augmented model is then constructed as in BLP.
We fit four models to these data: the base and extended models under the controls and instruments in BLP, and the base and extended models under the augmented set of controls and instruments. In the BLP version, the base and extended models contain six and seven parameters, respectively, and ten instruments, while in the augmented variant, the base and extended models have nineteen and twenty parameters, respectively, and $53$ moment restrictions. We assume that the $n = 2217$ observations on $(y_{ijt},x_{ijt},z_{1ijt})$ are a random sample from the population of automobile products across markets and time. Because it is difficult to formulate priors on the parameters by a priori considerations, we randomly select $15\%$ of the sample to make training sample priors. In particular, we used the GMM estimate and its standard error fitted on the training data (model by model) as the prior mean and twice the GMM standard error as the prior standard deviation (sd). The ETEL is constructed from the remaining data and the posterior distribution of each model is sampled by the single block M-H algorithm of ChibGreenberg1995. This algorithm is fast and efficient despite the relatively large numbers of parameters and instruments. The results show that the posterior mean of the coefficient on $price$ is -0.14, and the 95% posterior credibility interval is (-.16,-.13). The posterior mean is larger in magnitude than the OLS estimate originally reported by BLP. Note that the posterior distribution of the covariance parameter, $v$, is concentrated to the right of zero, indicating that the $price$ is likely endogenous.
For confirmation, we turn to our formal test of endogeneity. The results are reported in Table (ref). We can see that the marginal likelihood is larger for the extended models in both the original BLP and the augmented BLP specifications, supporting the conclusion that price is endogenous.
We conclude this analysis by plotting the posterior distributions of the price coefficient from each model. The estimated effect of price on automobile demand is larger (in absolute value) when endogeneity of price is taken into account. Interestingly, the price effect is smaller and more concentrated in the augmented models, suggesting that some of the excess sensitivity to price observed in the original BLP model is due to the omission of the nonlinear controls. In addition, it is worth noting that if we were to only fit the base model (which the marginal likelihood confirms is misspecified in this case), we would miss the fact that incorporating nonlinearities impacts the posterior distribution.
The emphasis of the theory and applications in this paper is on situations with a single outcome variable; however, our framework can be applied more broadly. An important example is clustered longitudinal data. Let $y_{i} = (y_{i1},\ldots,y_{iT})$ denote $T$ potentially correlated and heteroskedastic measurements on subject $i$. The outcome is thus a $T\times 1 $ vector, rather than a scalar. Adjusting the dimensions of the controls and instruments, respectively, suppose that independently across $i$ the clustered outcomes follow the linear model $y_{i}= X_{i} {\Greekmath 010C} +Z_{1,i} {\Greekmath 010D} + {\Greekmath 0122}_{i} $, where $X_{i}$ is $T \times d_{x}$, $Z_{1,i}$ is $T \times d_{z_{1}}$, $Z_{2,i}$ is $T \times d_{z_{2}}$, and ${\Greekmath 0122}_{i}$ is $T \times 1$. Now assume that $Z_{1,i}$ and $Z_{2,i}$ satisfy the clustered data exogeneity restrictions $\mathbf{E}[Z_{j,i}'{\Greekmath 0122}_{i}({\Greekmath 0112})] = 0$, $j = 1,2$, but that the clustered data exogeneity restrictions $\mathbf{E}[X_{i}^{\prime }{\Greekmath 0122}_{i}({\Greekmath 0112})] = 0$ related to $X_i$ are in doubt. We can apply our framework to this problem by defining a base model in which the latter restrictions are imposed and an extended model that contains the inactive restrictions $\mathbf{E}[X_{i}'{\Greekmath 0122}_{i}({\Greekmath 0112})]=v$, where $v$ is now a $d_x \times 1$ vector of unknown parameters. In parallel to the approach developed above, the marginal likelihood comparison of these models is a test for the exogeneity of $X$.
As an illustration of this extended set-up, we consider a $T=4$ balanced longitudinal data set on airfares and passenger traffic for the years 1997, 1998, 1999, and 2000 from wooldridge2010econometric. For each year $t$ , $t\leq 4$, the data is clustered by route $i$, $i\leq n=1149$. For each flight route defined by the origin and destination cities, one has the log of the average number of passengers per day ($lpassen$), the log of the average one-way fare in dollars ($lfare$), the log of the distance in miles ( $ldist$), and the fraction of the market corralled by the biggest carrier ($ concen$). The model of interest is $lpassen_{it}={\Greekmath 010C}\, lfare_{it} +{\Greekmath 010D} _{1}trend_{t}+{\Greekmath 010D} _{2}ldist_{it}+{\Greekmath 0122} _{it}$, where $trend$ is a trend variable taking values $1,2,3,$ and $4$, and each of the variables in this regression is mean centered. The goal is to estimate the price elasticity parameter ${\Greekmath 010C}$, but one is concerned that $lfare$ is possibly endogenous. In the estimation we assume that $concen$ is a valid instrument (it does not directly appear in the outcome model and it affects $lfare$, both reasonable assumptions).
Clustered by route $i$, we have
or compactly as $y_{i}=\widetilde{W}_{1,i}{\Greekmath 0112} +{\Greekmath 0122} _{i}$, $ i=1,2,\ldots,1149$, where ${\Greekmath 0112}:7\times 1$ is the unknown parameter of interest. In this model, the distribution of ${\Greekmath 0122} _{i}$ is not specified. Moreover, the elements of ${\Greekmath 0122} _{i}$ can be serially correlated and heteroskedastic in an arbitrary, unknown way.
Now let $Z_{i} := \left(\widetilde{W}_{1,i},1,concen_{i}\right)$, $i \leq n$ , be a $4\times 5$ matrix, where $1$ is a vector of ones, and $ concen_{i}=\left(concen_{i1},\ldots,concen_{i4}\right) ^{\prime }:4\times 1$ is the vector of $concen$ values for route $i$. In the base model, $lfare$ is exogenous. The model is defined by the five moments
In the extended model, the $lfare$ moment condition is inactive. Specifically,
The ETEL-based estimation of these two models makes no assumption about the joint distribution of the cluster-level errors.
We specify the prior from a training sample. We randomly split the sample into a training sample (of say 115 clusters, equal to 10% of the total clusters) and an estimation sample (consisting of the remaining 1,034 clusters). We then estimate the base mode on the training sample with a student-t prior centered on the system wide 2SLS estimate from the training data, sd of 10 and 2.5 degrees of freedom. The posterior mean and sd is calculated from these training data under this prior. We then take the posterior mean and twice the sd from the training sample fit as the mean and sd of the prior. This determination of the prior from the training sample is helpful in the fitting, but, due to the thick tails of the prior, the information brought in by the prior pales in comparison with the information from the estimation sample.
We sample the posterior in each model by the one-block tailored MCMC algorithm. In the base model, from 10,000 MCMC draws beyond a burn-in of 1,000, we find that the posterior mean of ${\Greekmath 010C}$ is -0.551 and its 95% posterior credibility interval is (-0.683, -0.419). Moreover, computation shows that $\log (m(w_{1:n}|\mathcal{M}_{b})=-7190.222$ and $\log (m(w_{1:n}| \mathcal{M}_{e})=-7191.06$, signaling that $lprice$ in this problem can be viewed as exogenous.
This paper develops a Bayesian test for exogeneity/endogeneity of the treatment vector of interest in a linear mean regression model. This endogeneity problem is generally assumed away in the Bayesian literature, but this leads to a serious misspecification problem since endogeneity, in practice, is the rule rather than the exception. In order to avoid the risk of distributional misspecification, the framework we have developed relies only on moment restrictions. The analysis in the paper revolves around the study of two models: the base model, where the exogeneity assumption is enforced, and an extended model, where the exogeneity moment is included but is made inactive.\newline The testing procedure for exogeneity/endogeneity is based on Bayes factor where the marginal ETEL of the base and the extended models are compared. The procedure is validated from a frequentist point of view because we establish the large sample consistency of the Bayes factor test. In addition, we provide a comprehensive study of the log-marginal ETEL function and determine which parts of it plays a role in the testing procedure depending on whether the covariates $x$ are exogenous or endogenous.\newline The real data examples discussed in the paper showcase the practical relevance of the methods.
It is important to mention that the approach proposed here can be extended to situations where the controls are assumed to enter the model nonparametrically. While the finite sample analysis of such models, after approximating the unknown functions by (say) spline basis expansion methods, would proceed in much the same way as discussed in this paper, the specification of the prior and the large sample analysis would require new developments to account for a growing number of basis function parameters with sample size. We intend to describe the theory in a future paper.
Another interesting extension would involve relaxing our current strong identification assumption, which is based on a fixed number of relevant instruments, to allow for weak and many-instrumental variables. This would require substantial changes to our asymptotic results, as it would necessitate developing local-asymptotic Bayes factors for ETEL and incorporating instrument-growth penalties or shrinkage to prevent overfitting.