Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
94,047 characters · 20 sections · 34 citation commands
Focused econometric estimation for noisy and small datasets: A Bayesian Minimum Expected Loss estimator approach
\thispagestyle{empty}
\setcounter{page}{1} \pagenumbering{arabic}
Central to many statistical and econometric inferential situations is the estimation of rational functions of the parameters. These functions might be elasticities, forecasts, impulse responses, marginal effects, odds ratios, optimal quantities, or structural parameters, among others. The mainstream in statistics and econometrics estimates these quantities based on a {\it{plug-in}} approach, where parameter estimates are just plugged in to the objective expressions without consideration of the main objective of the inferential situation. The popularity of this approach is based on the asymptotic properties of the delta method. However, this approach suffers from shortcomings, such as infinite moments and unbounded risks, when based on common considerations such as Gaussian likelihoods, and quadratic loss functions.\\
Models have a purpose. So, we should optimally design model's estimation frameworks for their purposes. This article is concerned with the estimation of rational functions of parameters, but deviates from the mainstream in that we focus the estimation process directly on the quantities of interest. This idea has been used by Claeskens2003 and Hansen2005 for model selection, and DiTraglia2016 for selecting moment conditions in the generalized method of moments (GMM).\\
We extend the idea of the Bayesian Minimum Expected Loss (MELO) approach introduced by Zellner1978, whose main theoretical developments and applications were confined to structural econometric models. We introduce it for problems where the main concerns of the inference are rational functions of the parameters. In particular, we follow a decision theoretic framework where the posterior expected value of a generalized quadratic loss function that depends explicitly on the function of interest is minimized.\\
The Bayesian Minimum Expected Loss approach gives point estimates, so the obvious and common answer to define their degree of variability would be to use the posterior density of the function of interest. This is a right answer, if the prior distribution where based on genuine past experience berger06,Efron2015. However, we use diffuse and vague priors, and see the MELO as an estimator, then we follow the idea of Efron2012,Efron2015, who proposes to estimate the frequentist variability of the Bayesian estimates. In addition, this approach avoids sensitivity analysis of the choice of prior or hierarchical prior structures, and as a consequence their extra computational burden.\\
We find that the asymptotic properties of the MELO estimators are similar to those of the plug-in approach. However, simulation exercises suggest that our proposal obtains better outcomes than competing alternatives; especially in settings characterized by noisy models and small sample sizes. In addition, we apply our proposal to real datasets, finding that MELO is more efficient than other alternatives.\\
The MELO approach has its foundation in statistical decision theory Wald1945,Wald1947, which initially was advocated in econometrics by Marschak1960 and Dreze1974. It was introduced in econometrics by Zellner1978, who analyzed reciprocals and ratios of parameters, and structural parameters in econometric models. He showed for these cases that the MELO estimator has, at least, finite first and second moments, and as a consequence finite risk with respect to a generalized quadratic loss function. On the other hand, common estimators like indirect least squares (ILS), two stage least squares (2SLS), limited information maximum likelihood (LIML), three stage least squares (3SLS) and full information maximum likelihood (FIML) have infinite moments and infinite risks using quadratic loss functions. Further, Zellner1979 approximate the small sample moments and risk functions of the MELO estimators, and compared them with other estimators. Zellner1980 found that coefficient estimates of structural parameters using MELO are matrix weighted averages of direct least squares (DLS) and 2SLS. Park1982 showed through simulation exercises that for structural parameters, the MELO estimates have more bias than 2SLS. However, MELO outperforms 2SLS in criteria like mean squared error (MSE) and mean absolute error (MAE). Swamy1983 analyzed the requirements of prior distributions for reduced form parameters associated with the MELO estimator in undersized sample conditions, that is, situations where the number of exogenous variables in simultaneous equations models exceeds the sample size. They found that the conditions for existence of the FIML estimator are more demanding than the conditions to obtain the MELO. Diebold1997 used the MELO approach to do an interesting application related to the response of agricultural supply to movements in expected price. They argued that the large variability of previous estimates associated with this phenomenon is due to the infinite moments and multimodal distributions of common frequentist estimators. In contrast, the MELO estimator has at least finite first and second moments; however, it also may exhibit multimodal distributions. Finally, Zellner1998 introduces the Bayesian Method of Moments, and related it to the MELO, extending the approach to cases where we have only moment conditions for our inferential problem. He presents the resuts of simulation exercises that show that Bayesian estimators perform better than popular frequentist estimators.\\
This paper is structured as follows. The next section develops the theoretical framework. Section (ref) exhibits the outcomes of the simulation exercises. Section (ref) presents the main findings in our applications. Finally, we make some concluding remarks.
Suppose that the main concern of the econometric inference is $\bm{\omega}={\bf{g}}(\bm{\theta}):\bm{\Theta}\subset\mathcal{R}^L\rightarrow \mathcal{R}^K$, $K\leq L$, that is, $\omega=(\omega_1,\omega_2,\ldots\omega_K)^T=(g_1(\bm{\theta}),g_2(\bm{\theta}),\ldots,g_K(\bm{\theta}))^T$, $g_k(\bm{\theta})=\frac{l_k(\bm{\theta})}{m_k(\bm{\theta})}:\mathcal{R}^L\rightarrow \mathcal{R}, k=1,2,\ldots,K$, such that
is a one-to-one continuously differentiable transformation for some nuisance transformation $\bm{\psi}=\bm{q}(\bm{\theta}):\mathcal{R}^{L}\rightarrow \mathcal{R}^{L-K}$.\\
Our view is that such an inferential problem should be directly tackled focusing on the functions of interest. So, we propose for this inferential problem the posterior Bayesian action that minimizes the posterior expected value of a generalized quadratic loss function focused on ${\bf{g}}(\bm{\theta})$, that is,
where $\mathcal{L}({\bf{g}}(\bm{\theta}),\hat{\omega})=({\bf{g}}(\bm{\theta})-\hat{\omega})^T{\bf{Q}}(\bm{\theta})({\bf{g}}(\bm{\theta})-\hat{\omega})$, ${\bf{Q}}(\bm{\theta})=diag\left\{h_k(\bm{\theta})\right\}$, where $h_k(\bm{\theta})$ are case specific weighting functions.\\
Let $Y_1,Y_2,\dots,Y_N$ be iid, each with density $f(y|\bm{\theta})$ with respect to a $\sigma$--finite measure $\mu$, where $\bm{\theta}$ is real--valued, and suppose the following regularity conditions hold.
Under these assumptions is well know that the maximum likelihood estimator satisfies $\sqrt{N}(\hat{\bm{\theta}}-\bm{\theta}_0)\xrightarrow{d} \mathcal{N}\left(0, I(\bm{\theta}_0)^{-1}\right)$, that is, $\hat{\bm{\theta}}$ is consistent for the true value $\bm{\theta}_0$, and asymptotically efficient Lehmann2003. Then $\frac{1}{N}[R_N(\bm{\theta})_{ij}]\xrightarrow{p}\bm{0}$ in the second order Taylor series expansion
where $l(\bm{\theta})=log(f(\bm{y}|\bm{\theta}))$ is the log likelihood, $\bm{y}=[y_1,y_2,\dots,y_N].$ However, Bayesian estimators involve an integral over the whole range of $\bm{\theta}$ values, then it is necessary the following assumption.
This assumption implies that the log likelihood contribution of $\bm{\theta}\notin \mathcal{B}_{\delta}(\bm{\theta}_0)$, where $\mathcal{B}_{\delta}(\bm{\theta}_0)$ is an open ball centered at $\bm{\theta}_0$ with radius $\delta$, is negligible as $N\rightarrow\infty$.
First assumption implies $\pi(\bm{\theta}_0)>0$, so $\bm{\theta}_0$ is not a priori excluded. The second assumption is required to proof that the minimum expected loss estimator is consistent and asymptotically efficient. We assume that there is a proper prior density function. However, result (ref) can be extended to the case $\int\pi(\bm{\theta})d\bm{\theta}=\infty$ when there is $n_0$, such that the posterior density taking information up to this point is a proper density with probability 1, satisfying assumptions D.
It is well known that assuming A to D, if $\pi^*(\bm{u}|\bm{y})$ is the posterior density of $\bm{u}=\sqrt{N}(\bm{\theta}-\hat{\bm{\theta}})$, then
where $\phi(B,\bm{x})$ is the density function of a multivariate normal distribution with mean $\bm{0}$ and covariance matrix $B$ Bickel1969,Lehmann2003.\\
Proof in Appendix (ref).\\
Observe that our MELO estimate is a weighted average of ${\bf{g}}(\bm{\theta})$, whose weights are given by $\left[\int_{\Theta}{\bf{Q}}(\bm{\theta})\pi(\bm{\theta}|{\bf{y}})d\bm{\theta}\right]^{-1}{\bf{Q}}(\bm{\theta})$. These weights implicitly depend on the probability associated with each $\bm{\theta}$ in their parameter space as well as their magnitude. When ${\bf{Q}}$ does not depend on $\bm{\theta}$, which implies equal weight to each $\bm{\theta}$, the Minimum Expected Loss estimate is the posterior mean, that is, $\hat{\omega}^*({\bf{y}})=E_{\pi(\bm{\theta}|{\bf{y}})}{\bf{g}}(\bm{\theta})$.\\
A good advantage of the MELO estimates is that they can be easily calculated from the draws of the posterior distributions, $\bm{\theta}_s\sim \pi(\bm{\theta}|{\bf{y}})$, and given $S\rightarrow \infty$, $\frac{1}{S}\sum_{s=1}^S {\bf{Q}}(\bm{\theta}_s)\xrightarrow{p}E_{\pi(\bm{\theta}|{\bf{y}})}{\bf{Q}}(\bm{\theta})$ and $\frac{1}{S}\sum_{s=1}^S {\bf{Q}}(\bm{\theta}_s){\bf{g}}(\bm{\theta}_s)\xrightarrow{p}E_{\pi(\bm{\theta}|{\bf{y}})}\left[{\bf{Q}}(\bm{\theta}){\bf{g}}(\bm{\theta})\right]$ by the law of the large numbers, then
converges in probability to $\hat{\omega}^*$ by Slutsky's theorem.\\
Observe that $\hat{\omega}^*_{k,S}({\bf{y}})=\sum_{s=1}^Sw_{ks}g_{k}(\bm{\theta}_s)$ where $w_{ks}=\frac{h_k(\bm{\theta}_s)}{\sum_{s=1}^Sh_k(\bm{\theta}_s)}$, that is, the MELO is a weighted average. Casella1998 show that weighted average estimators may perform better than unweighted average estimators, when evaluated under squared error loss functions. Obviously, this depends on the choice of $w_{ks}$. In particular, as our objective function is rational, there are singularities when $m_k(\bm{\theta})=0$, and as a consequence, $g_k(\bm{\theta})$ is not integrable under the posterior distribution. Therefore, if $\hat{\omega}^*_S({\bf{y}})$ is built such that puts less weight on the draws near singularity points, then the MELO estimator gains stability. In particular, setting $\mathcal{L}(g_k(\bm{\theta}),\hat{\omega}_k)=\epsilon^2_k$, where $\epsilon_k=\hat{\omega}_k m_k(\bm{\theta})-l_k(\bm{\theta})$ is an estimation error, such that $\hat{\omega}_k=g_k(\bm{\theta})$ implies $\epsilon_k=0$, then $h_k(\bm{\theta})=m_k(\bm{\theta})^2$.\\
From a frequentist perspective, there are many situations when the posterior Bayesian actions are equal to the Bayes rules, that is, the estimators that minimize the Bayes risk, $r(\pi_{\bm{\theta}},\hat{\omega})=\int_{\Theta}\int_{Y}\mathcal{L}({\bf{g}}(\bm{\theta}),\hat{\omega})f_Y({\bf{y}}|\bm{\theta})dy\ \pi(\bm{\theta})d\bm{\theta}$. However, there are situations where the Bayes risk is infinite, for instance using improper priors in conjunction general quadratic loss functions, and as a consequence, the Bayes rule does not exist. Nevertheless, it is still possible to obtain the posterior Bayesian action.\%
Proof: See Appendix (ref).\\
Proposition (ref) establishes that asymptotically, the MELO has similar characteristics to the maximum likelihood estimator. However, it seems that in noisy finite samples the MELO has better properties.\\
If uninformative priors are used, as in our exercises, it seems convenient to estimate the frequentist variability of the MELO estimator, as they were not based on genuine past experience berger06,Efron2012,Efron2015. In addition, this approach avoids sensitivity analysis of the choice of priors, or hierarchical Bayesian models, both imposing an extra computational burden, at the cost of requiring sufficient statistics.\\
To accomplish this task we have the following result.
Proof: See Appendix (ref).\\
Equality (ref) shows that our MELO estimate can be obtained from the posterior distribution associated with the data or its sufficient statistic. The resulting data reduction helps to estimate the frequentist variability of the MELO.\\
Setting
we have the following useful result.
See the proof in Appendix (ref).\\
See the proof in Appendix (ref).\\
Lemma (ref) allows calculating the frequentist variability of the MELO estimate (ref) through the delta method.\\
See the proof in Appendix (ref).\\
The setting of our formulation establishes Proposition (ref) as an optimal point estimate for functions of parameters. In the case that an analytical solution does not exist, we can use draws of the posterior distributions to obtain the estimates (Equation (ref)). Proposition (ref) allows obtaining the frequentist variance of our Bayesian estimate, provided a sufficient statistic.\\
We consider a very simple problem where a firm is interested in finding the level of input ($x$) that maximizes its profit, where the production function is quadratic, that is, $y = \beta_{1} x + \beta_{2} x^{2}$. So, the problem is
where $p$ is the product's price, $CF$ represents the fixed costs, and $w$ is the input's price.\\
Then the optimal input is given by
This implies that the optimal production and profit are $y^{Opt}=\frac{1}{2\beta_2}\left[\left(\frac{w}{p}\right)^2-\beta_1^2\right]$ and $\Pi^{Opt}=\frac{\beta_1}{2\beta_2}\left[w-\beta_1p\right]-CF$, respectively.\\ Suppose that the decision problem is to find the optimal level of input (Equation (ref)), that is, ${\bf{g}}(\bm{\theta})=\omega(\beta_1,\beta_2)=x^{Opt}$.
We can exploit the variability between $y_i$ and $x_i$ in the product function, the variability between $x_i$ and $w_i/p_i$ or $y_i$ and $w_i/p_i$ in the optimal input or production functions, or the variability between $\Pi_i$, $w_i$ and $p_i$ in the optimal profit function to obtain estimates of $\beta_1$ and $\beta_2$. The choice depends on assumptions regarding the rationality of the firms as well as the availability of the data.\\
We propose to formulate the mean deviation model associated with the production function to obtain the parameter estimates $\beta=\left[\beta_1\ \beta_2\right]'$. In particular, $y_i-\bar{y}=\beta_1(x_i-\bar{x})+\beta_2(x_i^2-\overline{x^2})+u_i$, where $\bar{y}=(1/N)\sum_{i=1}^Ny_i$, $\bar{x}=(1/N)\sum_{i=1}^Nx_i$, $\overline{x^2}=(1/N)\sum_{i=1}^Nx^2_i$ and $u_i\sim\mathcal{N}(0,\sigma^2)$, $i=1,2,\ldots,N$.\\
The likelihood function of this model is
where $X$ is the design matrix, $q=dim\{\beta\}$, $v=N-q$, $\hat{\beta} = (X^TX)^{-1} X^Ty$ and $vs^{2} = (y - X\hat{\beta})^T(y - X\hat{\beta})$. $\hat{\beta}$ and $s^2$ are sufficient independent statistics, such that $\hat{\beta}\sim\mathcal{N}_q(\beta,\sigma^2(X^TX)^{-1})$ and $s^2\sim \left(\frac{\sigma^2}{N-q}\right)\chi^2_{N-q}$. This implies
and
The {\it{plug-in}} estimator for the optimal input would be
In addition, the application of the delta method to estimate the variance would give as a result
On the other hand we can obtain the MELO estimate focusing directly on the inferential problem. We set $\epsilon = -\left(\frac{w}{p} - \beta_{1}\right) - 2\beta_{2} \hat{\omega}$ as the estimation error. Observe that if $\hat{\omega}$ is equal to $x^{Opt}$, the estimation error is equal to 0.\\
The generalized loss function for this problem is given by
where $\omega = {\bf{g}}(\bm{\theta}) = \frac{1}{2\beta_{2}}\left(\frac{w}{p} - \beta_{1}\right)$ and ${\bf{Q}}(\bm{\theta})=4 \beta^{2}_{2}$.\\
Proposition (ref) implies that the MELO estimate is
Using the following diffuse prior $p(\beta,\sigma) \propto 1/\sigma$, $0 < \sigma < \infty$ and $-\infty<\beta_{l}<\infty$, $l=\{1,2\}$, then the marginal posterior pdf for $\beta$ has the form of a multivariate Student-$t$ Zellner1996:
which implies that the mean of $\beta$ is $\hat{\beta}$ and its covariance matrix is $(X^TX)^{-1}vs^2/(v-2)$, $v>2$.\\
We can use the previous expressions to calculate our MELO proposal (Equation {(ref)}), and Equations (ref) and (ref) to obtain the frequentist variance of the MELO estimate.\\
We set the mean deviation problem, $y_i-\bar{y}=1.5(x_i-\bar{x})-0.002(x_i^2-\overline{x^2})+u_i$, where $x_i\sim\mathcal{N}(187.5,70^2)$ and $u_i\sim\mathcal{N}(0,\sigma^2_u)$ such that $\sigma^2_u$ generates different degrees of signal to noise models $\{0.1,1,5,20\}$. In addition, we set the input and output prices equal to \$3,000 and \$4,000, respectively. This implies $x^{Opt}=187.5$.\\
We perform 1,000 simulation exercises using different sample sizes (20, 50 and 500), and calculate the Mean Squared Error (MSE) and the Mean Absolute Error (MAE) for the {\it{plug-in}} approach, and the MELO using the analytical solution (Equation (ref)), which is available in this setting, and the computational strategy of drawing from the posterior distribution (Equation (ref) using 10,000 iterations from a Student's $t$ distribution).\\
We see from Table (ref) that the MELO outperforms the {\it{plug-in}} approach in point estimates of the optimal input; especially in the presence of noisy models and small sample sizes. In addition, we observe that there is no meaningful difference between the analytical and computational solutions.\\
In particular, there is no clear pattern in the MSE and MAE in very noisy models as the sample size increases. However, we do observe that the MELO estimates outperform the {\it{plug-in}} approach in this situation. As the signal of the model improves, the MSE and MAE decrease as the sample size increases. The MSE and MAE from the MELO estimates (analytical and computational) are never worse than the {\it{plug-in}} estimates. However, we basically get the same outcomes using large sample sizes. This outcome follows from the previous asymptotic properties.
Setting $y_{i}$ as a dichotomous variable $\left\{0,1\right\}$ that is distributed as a Bernoulli process with parameter $p$, and assuming that the main interest is the Odds ratio, it follows that
where $p=P(y = 1)$.
The binary probit model can be used to tackle this situation, such that $p=P(y_{i} = 1) = \Phi (x^T_{i}\beta)$, where $\Phi(z)$ is the cumulative distribution function of the standard normal distribution evaluated at $z$.\\
This model can be written with latent variables as follows:
The likelihood function is
Observe that in this setting there are no sufficient statistics Nelder1972.\\
The plug-in estimator for the Odds ratio is
And its variance, calculated by the delta method, is
The loss function is given by
where $\omega = {\bf{g}}(\bm{\theta}) = \frac{\Phi (x^T_{i}\beta)}{1- \Phi (x^T_{i}\beta)}$ and ${\bf{Q}}(\bm{\theta}) = (1 - \Phi (x^T_{i}\beta))^{2}$.
Proposition (ref) implies that the MELO estimate is
Note that if $\hat{p} \rightarrow 1$, then $1-\hat{p} \rightarrow 0$, and so $\hat{\omega}^{plug} \rightarrow \infty$, while $\hat{\omega}^{*}$ can take indeterminate values of the form $0/0$.\\
According to greenberg2012introduction, using the latent variables $y^*_i$, we can write the likelihood function as
where $q=dim\{\beta\}$, $v=N-q$, $\hat{\beta} = (X^TX)^{-1} X^Ty^*$ and $vs^{2} = (y^* - X\hat{\beta})^T(y^* - X\hat{\beta})$. This implies that augmenting the observed binary data $y$ with the latent variable $y^*$, $\hat{\beta}$ and $vs^{2}$ are sufficient statistics, so that
and
Assuming a normally distributed prior for $\beta$, the posterior distributions of $\beta$ and $y^*$ are
Therefore,
Consider the following setting.
We simulate the data set $x_{1}$, $x_{2}$ and the stochastic errors from standard normal distributions, and perform 1,000 simulation exercises using four different sample sizes: 20, 50, 500 and 1,000.\\
Tables (ref) and (ref) show the mean errors of our simulation exercises. In particular, we perform two different evaluations for the Odds ratio, $x=(1,1,1)$ and $x=(1,0,0)$ using Algorithm (ref) while setting $B_0=10,000 \ diag\left\{1,1,1\right\}$ and $\beta_0=\left[0,0,0\right]$ with 25,000 iterations and a burn-in equal to 5,000. We see from these tables that the range of variability of the different measures of the MELO approach is lower than for the {\it{plug-in}} approach. We observe that when the sample size is small, the differences are remarkable, especially when $x=(1,1,1)$, that is, when the data is less informative (noisy) due to regressors not being located in their population means ($x=(1,0,0)$). We obtain similar results for both approaches as the sample sizes increases.
One strategy for active portfolio management looks for finding the asset weights that maximize the Sharpe ratio, that is, the mean portfolio return per unit of risk.
where $\tilde{\mu}$ is the mean vector of the asset's excess returns in the investment period (say $\tau$), $\tilde{\Sigma}$ is its covariance matrix, and ${\bf{1}}$ is a vector of ones.\\
The solution of the previous problem gives the well known tangent portfolio, that is,
As we can see from Equation (ref), the final aim of the inferential problem is a rational function of the parameters of the asset's excess returns.\footnote{Given $A_{d\times d}$ invertible, then there exists a polynomial $p$, such that $A^{-1}=p(A)$.}
The standard financial literature assumes that the asset's excess returns are jointly normally distributed, i.e., $r_t\sim\mathcal{N}_d(\mu, \Sigma)$ for $t=1,2,\ldots,T$, where the excess returns are serially independent.\\
Now put ${\bf{R}}=({\bf{r_1}}, {\bf{r_2}}, \ldots, {\bf{r_L}})$ a $T\times L$ matrix of observations on $L$ asset excess returns. Then we can write the following model for the excess returns:
$${\bf{R}}={\bf{1}}\mu^T+e$$
where $e=(e_1, e_2, \ldots, e_L)$ is an $T\times L$ matrix of unobserved random disturbances. The rows of $e$ are independently distributed, which precludes any auto or serial correlation of disturbance terms, each with an $L$-dimensional normal distribution with zero mean vector and positive definite $L\times L$ covariance matrix $\Sigma$.\\
The likelihood of this model is
$$f(\mu,\Sigma|{\bf{R}})\propto |\Sigma|^{-T/2} exp\left\{-\frac{1}{2}tr(S\Sigma^{-1})-\frac{1}{2T}tr((\mu-\hat{\mu})(\mu-\hat{\mu})^T\Sigma^{-1})\right\}$$
where $\hat{\mu}$ is the sample mean vector and $S=({\bf{R}}-{\bf{1}}\hat{\mu}^T)^T({\bf{R}}-{\bf{1}}\hat{\mu}^T)$. $\hat{\mu}$ and $S$ are sufficient statistics, such that $\hat{\mu}\sim\mathcal{N}_L(\mu,\Sigma)$ and $S\sim \mathcal{W}_L(T-1,\Sigma)$. $\hat{\mu}$ and $S/(T-1)$ are consistent estimators for $\mu$ and $\Sigma$. Then,
where $Var(S_{ij})=(T-1)(\sigma_{ij}^2+\sigma_{ii}\sigma_{jj})$ and $Cov(S_{ij},S_{kl})=(T-1)(\sigma_{ik}\sigma_{jl}+\sigma_{il}\sigma_{jk})$.\\
The {\it{plug-in}} estimator for the tangent portfolio is
On the other hand we can obtain the MELO estimate focusing directly on the inferential problem. We set $\epsilon=({\bf{1}}^T\tilde{\Sigma}^{-1}{\tilde{\mu}}) {\hat{\omega}}-\tilde{\Sigma}^{-1}{{\tilde{\mu}}}$ as the estimation error. Observe that if $\hat{\omega}$ is equal to ${\bf{w}}^{Opt}$, the estimation error is equal to 0.\\
Given ${\bf{g}}(\bm{\theta})=\frac{\tilde{\Sigma}^{-1}{\tilde{\mu}}}{{\bf{1}}^T\tilde{\Sigma}^{-1}{\tilde{\mu}}}$, the generalized loss function for this problem is given by $\mathcal{L}({\bf{g}}(\bm{\theta}),\hat{\omega})=\epsilon^T\epsilon$ and $E(\mathcal{L})=E_{\pi(\mu,\Sigma|{\bf{R}})}\epsilon^T\epsilon$, ${\bf{Q}}(\bm{\theta})=({\bf{1}}^T\tilde{\Sigma}^{-1}{\mu})^2$. Despite the fact that the expected value should be based on information up to the investment period ($T+\tau$), we only have information up to $T$, so the expected value is conditioned on $\textbf{R}$. However, informative priors can be based on experts' views of the investment period.\\
Proposition (ref) implies that the MELO estimate is
Using the diffuse prior $\pi(\mu,\Sigma)=\pi(\mu)\pi(\Sigma)\propto |\Sigma|^{-(L+1)/2}$, the conditional posterior distribution for the mean vector of asset excess returns is $\mu|\Sigma,{\bf{R}}\sim\mathcal{N}_L(\hat{\mu},\Sigma/T)$, and the marginal distribution for the covariance matrix is $\Sigma|{\bf{R}}\sim\mathcal{I}\mathcal{W}_L(T-1,S)$ Zellner1996. Therefore, we can use a Gibbs sampling algorithm to obtain a computational solution for our MELO estimate.\footnote{Observe that following the {concentional} Bayesian portfolio selection, which is based on the predictive distribution of the excess returns in the investment period, we have $\tilde{\mu}=\hat{\mu}$ and $\tilde{\Sigma}=\frac{\left(\tau+\frac{1}{T}\right)(T-1)}{T+\tau-2-L}\hat{\Sigma}$. The term $\frac{\left(\tau+\frac{1}{T}\right)(T-1)}{T+\tau-2-L}$ cancels out in Equation (ref).}
We set $\mu_l\sim\mathcal{U}(-0.2,0.2)$, $l=1,2,\dots,L$, and generate $\Sigma$ such that it is semidefinite positive. We set four different scenarios of portfolio selection: $L=\left\{10, 25, 50, 100\right\}$ assets, and two sample sizes: $T=\left\{120, 240\right\}$ periods. We perform 100 simulations for each of the 8 settings, so that ${\bf{R}}\sim\mathcal{N}(\mu,\Sigma)$.\\
We estimate the sample mean and covariance matrix to calculate the optimal weights using the {\it{plug-in}} approach (Equation (ref)), and the Gibbs sampling algorithm with 1,000 iterations to calculate our MELO proposal (Equation (ref)). Then, we obtain the MSE and MAE using the population parameters (Equation (ref)), and the two estimators. We can see in Tables (ref) and (ref) the outcomes of our simulation exercises. In particular, the mean of the MSE and MAE associated with the MELO is always lower than the {\it{plug-in}} approach; there are remarkable improvements of our proposal when the number of assets in the portfolio selection problem is small. In the latter cases, the range of variability in the MSE and MAE using the {\it{plug-in}} approach is enormous compared with the MELO approach.
Assume the following structural supply--demand model:
where $q^d_i$ and $q^s_i$ are demand and supply functions, $p_i$ is the price, $z_{1i}$ and $z_{2i}$ exogenous regressors, and $\mu_{di}$ and $\mu_{si}$ stochastic errors, $i=1,2,\dots,N$.\\
The equilibrium condition equates demand and supply, that is, $q^s=q^d$. the structural parameters are the main concern of the econometric inferential problem.
Equation (ref) cannot be directly estimated due to endogeneity issues. So, it is necessary to obtain the reduced form system
which can be written as ${\bf{Y}}={\bf{X}}B+{\bf{U}}$, where ${\bf{Y}}=\left[q \quad p\right]$, an $N\times 2$ matrix of observations on quantities and prices, ${\bf{X}}$ is an $N\times 3$ matrix of a vector of ones, and the two independent variables ($z_1$ and $z_2$), with rank $3$, $B=\left[\pi\quad\gamma\right]$ is a $3\times 2$ matrix of regressions parameters from the reduced form, and ${\bf{U}}=\left[e_q\quad e_p\right]$ is a $N\times 2$ matrix of unobserved stochastic errors. We assume that the rows of $U$ are independently distributed, each with a 2-dimensional normal distribution with zero mean vector and positive definite $2\times 2$ covariance matrix $\Sigma$.\\
The likelihood function of this system is $$f(B,\Sigma|{\bf{Y}},{\bf{X}})\propto |\Sigma|^{-T/2} exp\left\{-\frac{1}{2}tr({\bf{S}}\Sigma^{-1})-\frac{1}{2}tr((B-\hat{B})^T({\bf{X}}^T{\bf{X}})(B-\hat{B})\Sigma^{-1})\right\}$$
where $\hat{B}=({\bf{X}}^T{\bf{X}})^{-1}{\bf{X}}^T{\bf{Y}}$ a matrix of least squares quantities, and ${\bf{S}}=({\bf{Y}}-{\bf{X}}\hat{B})^T({\bf{Y}}-{\bf{X}}\hat{B})$. $\hat{B}$ and $S$ are sufficient statistics, such that $vec(\hat{B})=\left[\hat{\pi}^T \quad \hat{\gamma}^T\right]^T\sim\mathcal{N}_6(vec(B),\Sigma\otimes ({\bf{X}}^T{\bf{X}})^{-1})$ and ${\bf{S}}\sim \mathcal{W}_2(N-3,\Sigma)$. $\hat{B}$ and ${\bf{S}}/(N-3)$ are consistent estimators for $B$ and $\Sigma$.\\
Then,
and
where $\beta=vec(B)$, $\hat{\beta}=vec(\hat{B})$, $Var(S_{ij})=(N-3)(\sigma_{ij}^2+\sigma_{ii}\sigma_{jj})$ and $Cov(S_{ij},S_{kl})=(N-3)(\sigma_{ik}\sigma_{jl}+\sigma_{il}\sigma_{jk})$.\\
The relation between the structural parameters, which are the main concern of the econometric inferential problem, and the reduced form parameters is given by the following system of equations:
There are different alternatives for obtaining the structural parameters from the reduced form. In this setting, which is an exactly identified model, the point estimates using the ILS ({\it{plug-in}} approach), 2SLS, or 3SLS, give the same results.\\
We set the vector of errors to be
The loss function is $\mathcal{L}({\bf{g}}(\bm{\theta}),\hat{\omega}) = \epsilon^{T}\epsilon$, which implies ${\bf{Q}}=diag(\gamma_2^2,\gamma_2^2,\gamma_1^2,\gamma_1^2)$. As a consequence, the MELO is given by the following set of simultaneous equations.
Observe that the different components of the MELO estimates are independent. This is due to the structure of the weighting matrix: it is a diagonal matrix. So, we can focus our effort on the specific structural parameters of interest.\\
Using the diffuse prior $\pi(B,\Sigma)=\pi(B)\pi(\Sigma)\propto |\Sigma|^{-(2+1)/2}$, the conditional posterior distribution for the mean vector is $vec(B)|\Sigma,{\bf{Y}},{\bf{X}}\sim\mathcal{N}_6(vec(\hat{B}),\Sigma\otimes ({\bf{X}}^T{\bf{X}})^{-1})$, and the marginal distribution for the covariance matrix is $\Sigma|{\bf{Y}},{\bf{X}}\sim\mathcal{I}\mathcal{W}_2(N-3,S)$ Zellner1996. Therefore, we can use a Gibbs sampling algorithm to obtain a computational solution of our MELO estimate.\\ We consider the following structural supply--demand model:
which implies the following reduced form model:
We simulate $z_{1i}$ and $z_{2i}$ from standard normal distributions, and the stochastic errors from the reduced system as independent variables with mean zero and standard deviation such that the signal to the noise ratio in the reduced equations are simultaneously equal to 0.1, 0.5, 1 and 5. We know that from a theoretical point of view this is a mistake since there is a correlation between the stochastic errors in the reduced form system. However, we follow this setting to have independence between the equations, and as a consequence we know that the marginal posterior distributions of each equation are independent multivariate Student's $t$ distributions. This implies that
and so we can compare the analytical and computational versions of the MELO estimates with the frequentist competing alternative. In particular, we use 50,000 iterations for the Gibbs sampling algorithm used to calculate the computational MELO. \\
We can see in Table (ref) the mean errors associated with 1,000 simulations exercises using five different sample sizes: 20, 50, 100, 1,000 and 20,000. We observe from this table the same pattern as in the previous simulations exercises. The MELO estimates outperform 2SLS in terms of point estimates, especially in situations characterized by noisy models and small sample sizes. However, we always get the same performance with a sample size equal to 20,000 or a signal to noise ratio equal to 5. The performance of the three approaches improves as the sample size increases as well as when the signal to noise increases. In general, the MELO estimates are never worse than those of the 2SLS, and the analytical and computational solutions have the same performance.
This is the broiler input--output example presented by Judge1988. In particular, the average weight of an experimental lot of broilers and their corresponding levels of average feed consumption was tabulated over the time period in which they changed from baby chickens to mature broilers ready for market.\\
Given the setting of the optimal input problem in subsection (ref), the dataset in Table 5.3 from Judge1988, and taking into account that broilers are 30 cents per pound and feed is 6 cents per pound, the optimal level of feed input is 13.74 with a standard deviation equal to 1.89 using the {\it{plug-in}} approach (Equations (ref) and (ref)), whereas the optimal input point estimate using the MELO approach, both analytical (Equation (ref)) and computational (Equation (ref) using 10,000 iterations), is 13.14. The standard deviations is 1.46, that is, reductions of 22%. These figures are calculated using Equation (ref) with Corollary (ref), and Equation (ref), with Proposition (ref) in the case of the analytical and computational approaches. Despite the fact that the coefficient of determination in this example is very high ($R^2=0.98$), we observe differences between the optimal weight estimates. In addition, the frequentist variability of the optimal weight using the MELO estimates are lower than the one using the {\it{plug-in}} approach.
In 1986, the space shuttle Challenger exploded during take off, killing the seven astronauts aboard. The explosion was the result of an O-ring failure, a splitting of a ring of rubber that seals the parts of the ship together, due to the unusually cold weather ($31^{o}F$, i.e., $0^{o}C$) at the time of launch dalal1989risk.\\
We calculated the Odds ratio at $45^{o}F$ and $69.56^{o}F$ (mean sample temperature) for a sample of 23 observations provided by robert2004monte taking into account the theoretical structure of subsection (ref). Using the {\it{plug-in}} approach, the probability of failure is 0.996 at $45^{o}F$, therefore the odds ratio estimate is 283.644 with a standard deviation of 999.596. The Odds ratio using MELO is 2.585 in the case of the computational approach (Algorithm (ref) setting $B_0=10,000 \ diag\left\{1,1,1\right\}$ and $\beta_0=\left[0,0,0\right]$ with 25,000 iterations and a burn-in equal to 5,000). Observe that the implicit probabilities of the Odds ratio in the MELO approach is 0.721. However, if the main objective of the statistical inference is the probability, that is, ${\bf{g}}(\bm{\theta})=\Phi(x^T\beta)$, which implies ${\bf{Q}}(\bm{\theta})=1$, and $\hat{\omega}^{*}=E(\Phi (x^T_{i}\beta))$, we have point estimates equal to 0.964 using the computational MELO. This highlights a remarkable characteristic of our approach; the estimate depends drastically on the main objective of the inferential situation.\\
Regarding the frequentist variability of the MELO, we get 2.917 using the computational approach. Observe that in this case, the components associated with Corollary (ref) depend on the iteration $g$, so we calculate the mean values over all these components to obtain this figure. Observe that there is a huge difference using the delta method (999.596).\\
The failure probability is 0.266 at $69.56^{o}F$ using the {\it{plug-in}} approach. This implies an Odds ratio equal to 0.363 with standard deviation equal to 0.171. The Odds ratio using the computational MELO is 0.345 with a standard deviation equal to 0.258. We get similar point estimates using the central point in the distribution of regressors. In this case, the standard deviation of the {\it{plug-in}} is lower than for the MELO approaches (33%).\\
The message here is that in the case of evaluating a point in the extreme of the distribution of the regressors, that is, when the sample information is not precise (noisy), it is much better to use the MELO approach. On the other hand, it makes sense to use the {\it{plug-in}} approach.
Acemoglu2001 analyze the effect of property rights on economic growth. They exploit the variability in European settlers' mortality rates during the time of colonization to find the causal effect of protection against expropriation on economic performance. They use 2SLS to accomplish this task. We can write their setting in the following structural system,
where $pcGDP$, $PAER$ and $Mort$ are the per capita GDP in 1995, the average index of protection against expropriation between 1985 and 1995 (Political Risk Services), and settler mortality rate during the time of colonization (see Acemoglu2001 for details), respectively. The reduced form model is
The first structural equation is exactly identified provided that $\alpha_{2}\neq 0$, whereas the second structural equation is sub-identified.\\
We define the estimation error as $\epsilon=\gamma_1(\hat{\omega}-\beta_1)$, where $\beta_1=\pi_1/\gamma_1$, then ${\bf{Q}}(\bm{\theta})=\gamma_1^2$.\\
We find the MELO estimates, and their frequentist variability, using the same ideas of subsection (ref). The outcomes can be seen in Table (ref), where we reproduce the outcomes from Acemoglu2001, Table IV (page 1386), columns 1, 3, 5 and 9.\\
We can see from Table (ref) that the standard errors of our approach are always less than the standard errors from 2SLS. We obtain more efficiency gains in noisier models, for instance column (3), where the coefficient of determination is the lowest ($R^2=0.13$). In general, the MELO estimates of the effects of property rights on economic performance are lower than the 2SLS estimates.
Romer1993 analyzes the effect of openness on inflation. In particular, he shows there is a strong and robust negative link between inflation and openness using 2SLS, where he uses the logarithm of the country's land area as an instrument of openness. His model can be written as a structural system of equations:
where $Inf_i$, $Open_i$, $pinc_i$, $D_i$ and $land_i$ are the inflation rate, openness, which is measured as the ratio of imports to GDP, real per capita income, and data dummies for the alternative measures of openness and inflation, and land area, respectively.\\
The first structural equation is exactly identified provided that $\alpha_{3}\neq 0$, whereas the second structural equation is sub-identified. The reduced form of this model is
We define the estimation error as follows:
where $\beta_1=\pi_2/\gamma_2$, $\beta_2=\pi_1-\frac{\pi_2}{\gamma_2}\gamma_1$, then ${\bf{Q}}(\bm{\theta})=diag(\gamma_2^2,\gamma_2^2)$.\\
We can find the MELO estimates, and their frequentist variability, using the same ideas of subsection (ref). In particular, we have that the structural or causal effect of openness on inflation is -1.252 with a standard error equal to 0.407, and the effect of per capita income on inflation is equal to -0.045 with a standard error of 0.061 (using the computational approach with 10,000 iterations). The analogous estimates using 2SLS are -1.260 (0.414) and -0.045 (0.061) for openness and income, respectively.\\
Despite the fact that MELO and 2SLS are based on completely different frameworks, we practically do not get any differences between these estimates in this application. The reason is that the coefficient of determination in the first stage is equal to 0.48, and there are 100 degrees of freedom. This implies that the signal to noise ratio in the first stage is approximately 1, and given 100 d.f., we are basically replicating the outcomes of our simulation exercise in subsection (ref) (see Table (ref), Signal/Noise=0.5 and Sample size=100).\\
Many times the main concern of an econometric inference is associated with rational functions of parameters. Our approach tackles directly this issue based on a Bayesian decision theory framework, which allows thinking about the whole inferential situation. Our proposal seems to improve the econometric inference in situations characterized by small sample sizes or noisy models. So, our MELO proposal can be used in situations where getting observations can be a difficult task due to data limitations, for instance, expensive experimental designs or availability restrictions, and/or situations where the models are very noisy, for instance, very weak instruments. But, if there is a moderate sample size and/or the models are very informative, it is better to use the commonly used alternatives, due to the availability of the appropriate software for them.\\
However, we must acknowledge that our approach is based on rational functions and sufficient statistics. Future research should explore relaxing these assumptions.