Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
39,287 characters · 6 sections · 37 citation commands
Mean-shift least squares model averaging
\thispagestyle{empty}\setcounter{page}0
The goal of model averaging, where several models are averaged over, is to reduce estimation variance while controlling for bias arising from omitted variables (regression error). hansen2007least proposed the Mallows model averaging (MMA) estimator, which is shown to be asymptotically optimal with regard to the fitted estimates achieving the minimum squared error in a class of discrete model averaging estimators. However, while linear averaging (including MMA) can minimize regression error, it cannot simultaneously reduce regression error and location bias; the bias arising from non-zero expectation in the misspecification bias. In this paper, we propose a novel model averaging estimator that controls for location bias while reducing regression error. We call this new estimator the mean-shift Mallows model averaging (MSA) estimator. We relax the condition of discrete estimators to continuous estimators, similar to that of wan2010least, though under weaker conditions. We show that the proposed MSA estimator is asymptotically optimal to the MMA estimator.
Model averaging (or forecast combination/ensemble learning) has been a staple and necessary tool in the field of econometrics, statistics, and machine learning. Since the seminar paper by BatesGranger1969, the potency of model averaging has been proven over multiple applications and contexts. Recently, the field has received a surge of interest across many disciplines due to an increase in usage and availability of full forecast densities arising from more complex models. In the statistical literature, there is a large literature on model averaging, particularly for Bayesian methods. For example, Bayesian model averaging raftery1997bayesian,hoeting1999bayesian has become standard for Bayesians dealing with model uncertainty. In machine learning, ensemble methods, including bagging breiman1996bagging, boosting schapire2003boosting, and stacking dvzeroski2004combining, which are all linear model averaging (often equal weight), have been used extensively to alleviate overfitting. In econometrics, increased usage in density forecasts has stimulated the field for more advanced strategies for model combination Terui2002,HallMitchell2007,Amisano2007,Hooger2010,Kascha2010,Geweke2011,Geweke2012,Billio2012,Billio2013,Aastveit2014,Fawcett2014,Pettenuzzo2015,Negro2016,yao2018using,aastveit2018evolution,Aastveit2015, with success in applications across many subdisciplines of economics.
While practical success of model averaging is abundant, theoretical considerations, in particular with regard to estimation of averaging weights, have been somewhat limiting. In terms of theoretical results for averaging models over model selection, BatesGranger1969 show that a combination of two forecasts can yield improved forecasts over the two individual forecasts, in terms of mean squared forecast error. In the machine learning literature, louppe2014understanding show that, for averaging regression trees (i.e. random forests), ensemble learning lowers variance while maintaining bias. In the Bayesian literature, when the true model is nested in the candidate models ($\mathcal{M}$-closed, as it is referred, following bernardo2009bayesian), BMA is known to asymptotically converge to the true model. However, under an $\mathcal{M}$-open setting, where the true model is not nested, BMA tends to converge to the “wrong" model, which is often not even the “best" model.
In terms of estimating the weights for model averaging, hansen2007least showed that minimizing the Mallows criterion is asymptotically optimal in a class of discrete model average estimators in terms of mean squared error. wan2010least relaxes this assumption of discrete estimators to continuous estimators, and show that the optimality holds. Further theoretical results are developed for more complex settings hansen2012jackknife,zhang2013model. While the MMA estimator is optimal for model averaging, it cannot simultaneously minimize the bias arising from omitted variables and location bias, due to its limited parameter dimension.
Building on the MMA estimator in hansen2007least, we propose a new model averaging estimator that controls for both biases: the mean-shift Mallows model averaging estimator (MSA). Our main contribution is to show that our proposed MSA estimator is asymptotically optimal to the MMA estimator, achieving lower risk. While our MSA estimator is frequentist by nature, our motivation is, nonetheless, Bayesian. In particular, our development is inspired by the recent practical and theoretical success of Bayesian predictive synthesis (BPS); a coherent Bayesian framework to combine information-- in particular, density forecasts-- from multiple sources mcalinn2019dynamic,mcalinn2017multivariate,takanashi2019predictive. Part of the success of BPS is in considering forecasts as “data" to be used in prior-posterior updating, treating them as latent factors. The flexibility of this framework, of which MMA is a special case, has lead to development of more elaborate model averaging approaches. One of which is to control for location bias arising from misspecification by incorporating a common constant, which shifts the mean to counteract the bias. This idea is central to our proposed MSA estimator.
Section. (ref) discusses the setup as well as review the results from hansen2007least. Section. (ref) introduces the mean-shift Mallows model averaging estimator and its sampling properties. Section. (ref) provides finite sample evidence in favor of the proposed MSA estimator.
Let $\left\{y_{i}\right\} _{i\in n}\in\mathbb{R}$ and $\mbox{\boldmath$x$}$ be a countably infinite vector. The data generating process hansen2007least is
Note that the linearity of the DGP is not restrictive, as it includes series expansions. Consider each model, $m=1,2,...,$ uses the first $k_{m}$ elements of $x_{ji}$, where $0<k_1<k_2<...$, to construct an approximating model. This implies that the models are ordered, though this is not problematic when there are series expansions.
The misspecification bias (regression error) is
which is the component in the DGP that is never captured by any of the models, thus, under this setup, all models are misspecified. The predictive mean is,
where $P_{m}=X_{m}\left(X_{m}^{\top}X_{m}\right)^{-1}X_{m}^{\top}$ is the projection matrix and $\hat{\mbox{\boldmath$\theta$}}_{m}$ is the vector of least squares estimates. The full model, $M=M_{n}\leqq n$, is an integer for which $X_{k_{M}}^{\top}X_{k_{M}}$ is invertible.
The weights for model averaging is specified as
with the resulting linear averaged predictive mean being
Then, we have the following properties:
Define the risk (i.e. the conditional squared error), $R\left(W\right)=\mathbb{E}\left[L_{n}\left(W\right)\left|X\right.\right]$, as
where
following hansen2007least (Lemma 2). This shows that the risk is a quadratic function with regard to the weight vector, $W$. From this, we have
The residual from the linear projection can be written as
When $\ell\leqq m$, the residual is
The Mallows criterion for the model average estimator is
where
is the vector of model dimensions and the collection of residuals is
For the MMA estimator, the weight vector is selected by minimizing the Mallows criterion. This has the expectation hansen2007least,
The empirical weight vector is \[ \hat{W}_{n}=\arg\min_{W}C_{n}\left(W\right). \] for which hansen2007least showed that this criterion is optimal, as we summarize below.
Limiting the weight vector to a finite $N$ number of discrete values, where the possible values, $W\in\left[0,1\right]$, are the $N+1$, \[ w_{m}\in\left\{ 0,\frac{1}{N},\frac{2}{N},\cdots,1\right\}. \] The set of possible discrete weight vectors is denoted as $H_{n}\left(N\right)$. When $n\rightarrow\infty$, assume \[ \xi_{n}=\inf_{W\in H_{n}}R_{n}\left(W\right)\rightarrow\infty \] almost surely. Additionally, assume \[ \mathbb{E}\left[\left|e_{i}\right|^{4\left(N+1\right)}\left|x_{i}\right.\right]\leqq\kappa<\infty. \] Then, the following holds hansen2007least: \[ \frac{L_{n}\left(\hat{W}_{N}\right)}{{\displaystyle \inf_{W\in H_{n}\left(W\right)}}L_{n}\left(W\right)}\rightarrow1. \]
The theoretical results from hansen2007least shows that choosing $\hat{W}$ by minimizing $C_{n}\left(W\right)$ minimizes the loss (but not the risk), as well as achieve the optimal rate. However, limiting the weight vector to a finite $N$ number of discrete values is somewhat restrictive, which we will later relax.
While the MMA estimator is optimal for a class of discrete model averaging estimators, it does not mean that there is no room for improvement. This can be easily seen when we decompose the loss as
Since the linear projection, $P_{m}\mu$, will ultimately not be the conditional expectation of $y$ given $X_{m}$ and the averaged $P\left(W\right)\mu$ will not be the conditional expectation of $y$ given all the available information, this will leave $\left(\mu\left(W\right)-\mathbb{E}\left[y\left|\mathcal{F}_{t}\right.\right]\right)^{2}$ to be improved.
This improvement can be expound further by decomposing the DGP and identifying the source of possible improvement. Consider decomposing the DGP into the full model, $M$, and its misspecification bias; \[ y_{i}=\sum_{j=1}^{k_{M}}\theta_{j}x_{ji}+\sum_{j=k_{M}+1}^{\infty}\theta_{j}x_{ji}+\varepsilon_{i}. \] Here, the covariates for $M$, $\left\{ x_{j}\right\} _{j=1,\cdots,K_{M}}$, and the covariates for misspecification bias, $\left\{ x_{j}\right\} _{j=K_{M},\cdots,\infty}$, are considered independent. The conditional expectation of the full model for $y_{i}$, given all observable covariates, $\left\{ x_{j}\right\} _{j=1,\cdots,k_{M}}$, is \[ \mathbb{E}\left[y\left|\left\{ x_{j}\right\} _{j=1,\cdots,k_{M}}\right.\right]=\sum_{j=1}^{k_{M}}\theta_{j}x_{ji}. \] The conditional expectation of the misspecification bias of $M$, due to independence, is \[ \mathbb{E}\left[b_{Mi}\right]=\mathbb{E}\left[\sum_{j=k_{M}+1}^{\infty}\theta_{j}x_{ji}\right]. \]
If each $\left\{ x_{j}\right\} _{j=1,\cdots,\infty}$ are i.i.d., then $\mathbb{E}\left[b_{Mi}\right]$ does not depend on the sampling order $i$, thus we denote it as $\alpha_{M}$. The best predictive value, in terms of MSE, can be written as, \[ \alpha_{M}+\sum_{j=1}^{k_{M}}\theta_{j}x_{ji}, \] which is to say that the DGP can be decomposed into its location bias, $\alpha_{M}$, and the optimal regression, $\sum_{j=1}^{k_{M}}\theta_{j}x_{ji}$. This can be written in matrix form (with elements corresponding to $i$) as \[ \alpha_{M}\mathbf{1}+X_{M}\theta_{M}. \] Note that $\mathbf{1}$ is a vector of ones.
The expectation of the misspecification bias for each model, $h$, using $\alpha_{h}$, is \[ \alpha_{h}=\mathbb{E}\left[b_{i,h}\right]=\mathbb{E}\left[\sum_{j=k_{h}+1}^{\infty}\theta_{j}x_{ji}\right]. \] The intercept, $\hat{\alpha}_{h}$, is the estimate for $\alpha_{h}$ for each model, $h$. The predictive value, $\hat{\mu}_{h}$, for each model is, \[ \hat{\mu}_{hi}=\hat{\alpha}_{h}+\sum_{j=1}^{k_{h}}\hat{\beta}_{j}x_{ji}, \] where we separate the intercept. The matrix form (with elements corresponding to $i$) is \[ \hat{\mu}_{h}=\hat{\alpha}_{h}\mathbf{1}+X_{h}^{-\alpha_{h}}\hat{\boldsymbol{\beta}}_{h}^{-\alpha_{h}}. \] Note that $\mathbf{1}$ is a vector of ones, and $X_{h}^{-\alpha_{h}},\hat{\boldsymbol{\beta}}_{h}^{-\alpha_{h}}$ is $X_{h}\hat{\boldsymbol{\beta}}_{h}$ minus the intercept.
Given the above, linear model averaging (including MMA) can be written as \[ \sum_{h=1}^{k_{M}}w_{h}\left(\hat{\alpha}_{h}\mathbf{1}+X_{h}^{-\alpha_{h}}\hat{\theta}_{h}^{-\alpha_{h}}\right). \] The weights, $w_{h}$ are chosen to be as close to the best predictive value of $y$; $\alpha_{M}+\sum_{j=1}^{k_{M}}\theta_{j}x_{ji}$. The optimal weight vector is $w_{h}^{*}$. Then, except in rare situations, we have,
Since the weights, $w_{h}$, must be selected by simultaneously minimizing the location bias and regression error, the weights on the location bias need not be zero. There might exist a vector of weights, $w$, that satisfies $\alpha_{M}=\sum_{h=1}^{k_{M}}w_{h}\hat{\alpha}_{h}$, but that will likely worsen the regression error. This is due to the dimension for linear averaging missing one dimension with regard to optimizing the averaging weights. Thus, under linear averaging, there is always a tradeoff between the location bias and regression error, where minimizing one leads to an increase in the other, making all linear averaging estimators (MMA included) suboptimal.
To simultaneously minimize the location bias and regression error, we propose a new estimator that includes a common constant (i.e. an intercept) to the set of models, which we call the mean-shift Mallows model averaging (MSA) estimator. Our MSA estimator has the following predictive mean:
where the empirical weights, $\hat{W}$, and intercept, $\hat{\alpha}$, are estimated by minimizing the criterion we develop in this section.
When we include a common constant, the residual of the linear projection is
where
For the extra term $\left(1\right)$, arbitrarily fixing the weight vector, $W$, we have a quadratic function with regard to $\alpha$:
Since $\left(1\right)$ is downward convex and is zero when $\alpha=0$, the extra term can take a negative value. Then, the optimal $\alpha$ is
Adding the mean-shift $\alpha$ to linear averaging, the component for location bias can be written as \[ \alpha_{M}=\sum_{h=1}^{k_{M}}w_{h}^{*}\hat{\alpha}_{h}+\alpha, \] which uniformly improves the total error by fully controlling for location bias without compromising minimizing regression error. Additionally, adding $\alpha$ changes the optimal solution with regard to the weights $w$, which possibly improves the regression error as well.
Since the loss is defined as
the expectation is
Then, the mean-shift Mallows criterion is
which can be written as
Comparing the criterion with and without an intercept, there exists an $\alpha$ that satisfies \[ \min_{W,\alpha}C_{n}\left(W,\alpha\right)\leqq\min_{W}C_{n}\left(W\right). \] This can be seen by setting $\alpha=0$ with regard to $\left(W,\alpha\right)$ that satisfies $C_{n}\left(W\right)\leqq C_{n}\left(W,\alpha\right)$.
Given the extension, we now show that our proposed MSA is optimal to the MMA estimator in hansen2007least. We set the parameter space $K$ for which the intercept can take values, \[ K_{n}=\left\{ \alpha\left|\min_{W}C_{n}\left(W,\alpha\right)\leqq\min_{W}C_{n}\left(W\right)\right.\right\}. \] Here, $\left\{ K_{n}\right\} $ is on $\mathbb{R}$ including zero. Additionally, we assume \[ \mathbb{E}\left[\left|e_{i}\right|^{4\left(N+1\right)}\left|x_{i}\right.\right]\leqq\kappa<\infty. \]
When $n\rightarrow\infty$, we assume, \[ \xi_{n}\triangleq{ \inf_{
}}R_{n}\left(W,\alpha\right)\rightarrow\infty, \] almost surely. We additionally assume, for each point, $w\in H_{n}$,
This condition is weaker than that of wan2010least. Compared to the conditions in hansen2007least, \[ \inf_{w\in H_{n}}R_{n}\left(w\right)\rightarrow\infty, \] this might seem somewhat restrictive. However, we only need to add the following condition, \[ \frac{R_{n}\left(w,\alpha\right)}{{ \inf_{
}}R_{n}\left(w,\alpha\right)}<\infty,^{\forall}w\in H_{n} \] to the above, which is a reasonable condition to add.
The proof roughly follows that of li1987asymptotic and hansen2007least, though the loss, risk, and criterion are dependent on the continuous parameter, $\left(W,\alpha\right)$, which makes the proof challenging. Note that wan2010least provides the conditions to relax the discreteness of the weights. In contrast, we prove optimality using the convexity lemma by utilizing the fact that the loss and criterion is a convex function. This allows us to prove optimality with a weaker condition to that of wan2010least.
We first decompose $C_{n}\left(W,\alpha\right)$ as
Since the $\varepsilon^{\top}\varepsilon$ term does not affect model selection, this is equivalent to minimizing, \[ L_{n}\left(W,\alpha\right)+2\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle +2\sigma^{2}K^{\top}W-2\left\langle \varepsilon,P\left(W\right)\varepsilon\right\rangle. \]
To prove optimality, we must show \[ { \min_{W,\alpha}}\frac{C_{n}\left(W,\alpha\right)}{{ \inf_{
}}L_{n}\left(W,\alpha\right)}\rightarrow1, \] which implies that \[ P\left(\sup_{
}\left|\frac{C_{p}\left(h\right)}{{ \inf_{
}}L_{n}\left(W,\alpha\right)}-1\right|>\delta\right)\rightarrow0 \] must uniformly hold with regard to $W\in H_{n}$ and $\alpha\in K_{n}$.
Since $\frac{C_{n}\left(W,\alpha\right)}{L_{n}\left(W,\alpha\right)}$ can be written as \[ \frac{L_{n}\left(W,\alpha\right)+2\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle +\left\{ 2\sigma^{2}K^{\top}W-2\left\langle \varepsilon,P\left(W\right)\varepsilon\right\rangle \right\} }{R_{n}\left(W,\alpha\right)}\frac{R_{n}\left(W,\alpha\right)}{L_{n}\left(h\right)} \] and each term is
we need to show the following to prove optimality:
To show uniform convergence, it is sufficient to show pointwise convergence, following the below convexity lemma:
{\it Proof of} {\bf 2.2:} From Chebyshev's inequality, we have \[ P\left(\left|\frac{2\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle }{\xi_{n}}\right|>\delta\right)\leqq\frac{4\mathbb{E}\left[\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle ^{2}\right]}{\delta^{2}\xi_{n}^{2}}, \] which we evaluate the right hand side. From the inequality in whittle1960bounds, the following \[ \mathbb{E}\left[\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle ^{2}\right]\leqq2^{2}C\left(s\right)\sigma\left\Vert A_{m}\mu-\alpha1\right\Vert ^{2} \] holds, thus, \[ \frac{4\mathbb{E}\left[\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle ^{2}\right]}{\delta^{2}\xi_{n}^{2}}\leqq C\frac{1}{\delta^{2}}\frac{\left\Vert A_{m}\mu-\alpha1\right\Vert ^{2}}{\xi_{n}^{2}}, \] where $C$ is some constant. Now, since \[ \left\Vert A_{m}\mu-\alpha1\right\Vert ^{2}\leqq R_{n}\left(W,\alpha\right), \] we have \[ C\frac{1}{\delta^{2}}\frac{\left\Vert A_{m}\mu-\alpha1\right\Vert ^{2}}{\xi_{n}^{2}}\leqq C\frac{1}{\delta^{2}}\frac{R_{n}\left(W,\alpha\right)}{\xi_{n}^{2}}. \] Therefore, from assumption (ref), we prove {\bf 2.2}.
{\it Proof of} {\bf 2.3:} Similarly to {\bf 2.2}, applying Chebyshev's inequality, we evaluate the moment of the right hand side. Applying the quadric version of Whittle's inequality, we have \[ \mathbb{E}\left[\left|2\sigma^{2}K^{\top}W-2\left\langle \varepsilon,P\left(W\right)\varepsilon\right\rangle \right|^{2}\right]\leqq C\left(\sigma^{2}\textrm{tr}\left\{ P\left(W\right)P\left(W\right)\right\} \right), \] where, $C$ is a constant. Additionally, since \[ \sigma^{2}\textrm{tr}\left\{ P\left(W\right)P\left(W\right)\right\} \leqq R_{n}\left(W,\alpha\right), \] we prove {\bf 2.3}.
{\it Proof of} {\bf 2.4:} For this, it is sufficient to show convergence for the following:
thus,
Noting that
the first equation is equivalent to {\bf 2.2}, and the second equation, utilizing the quadratic version of the Whittle's inequality to evaluate the expectation, \[ \mathbb{E}\left[\left|\left\Vert P\left(W\right)\varepsilon\right\Vert ^{2}-\sigma^{2}\textrm{tr}\left\{ P\left(W\right)P\left(W\right)\right\} \right|^{2}\right] \] and further noting that
for some constant, $C$, we have the following: \[ \mathbb{E}\left[\left|\left\Vert P\left(W\right)\varepsilon\right\Vert ^{2}-\sigma^{2}\textrm{tr}\left\{ P\left(W\right)P\left(W\right)\right\} \right|^{2}\right]\leqq CR_{n}\left(W,\alpha\right). \] The rest is the same as {\bf 2.3}. Q.E.D.
We investigate the finite sample mean squared error of our proposed MSA estimator to that of the original MMA estimator in a simple simulation setting similar to hansen2007least. Our data generating process is a slight modification of eq. (ref), where the misspecification bias is positive:
where all $x_{ji}$ are independent and identically distributed $N(0,1)$, the error, $\varepsilon_{i}$, is $N(0,1)$ and independent of $x_{ji}$. The parameters are determined by the rule, $\theta_{j}=c\sqrt{2\alpha}j^{-\alpha -1/2}$, where $c=R^2/(1-R^2)$. Though not reported, the results are robust to alternative specifications, except when the bias has expectation zero (i.e. no location bias), at which MSA is equivalent to MMA.
For the finite sample investigation, we vary the sample size by $n=50,150,400, \textrm{ and } 1,000$ and $\alpha=0.5,1, \textrm{ and } 1.5$. The parameter $\alpha$ controls the rate of decline of the coefficients, $\theta_{j}$, where a larger $\alpha$ implies a steeper decline in $j$. The number of models is determined by $M=3n^{1/3}$, which, for our four sample sizes, means $M=11,16,22,\textrm{ and } 30$. We vary $c$ so as $R^2$ varies from 0.1 to 0.9.
Unlike hansen2007least, which compared the MMA estimator to AIC model selection, Mallows model selection, smoothed AIC, and smoothed BIC, we simply compare our MSA estimator to the MMA estimator, as hansen2007least convincingly show that the MMA estimator is superior to these methods. To evaluate the two estimators, we compute the log risk (expected squared error) for each and calculate the difference. The simulation is done for 1,000 iterations and averaged over.
The results of the finite sample investigation are shown in Fig. (ref). The three panels display the difference in log risk between MMA and MSA for the varying $\alpha$, where the difference is displayed on the $y$ axis and the population $R^2$ is displayed on the $x$ axis. The different colored lines are for the different sample sizes, $n$.
For all results, our proposed MSA estimator uniformly outperforms the MMA estimator to varying degrees. In the case of $\alpha=0.5$, the results are almost identical for all sample sizes. The relative performance of MSA increases as $R^2$ increases, which is expected as the increased $R^2$, under a low $\alpha$, implies a more robust bias component. This bias component is captured in MSA, leading to improved performances. As $\alpha$ increases, the results invert, in the sense that, for $\alpha=1$, the difference in log risk is bowed, while, for $\alpha=1.5$, the difference decreases as $R^2$ increases. This is not unexpected as, in hansen2007least, the performance of the MMA estimator decreases in these cases. If anything, for $\alpha=1$, the MSA estimator seems more resilient for higher $R^2$. Over all variations, the MSA estimator is superior to the MMA estimator.