EconBase
← Back to paper

Mean-shift least squares model averaging

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

39,287 characters · 6 sections · 37 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Mean-shift least squares model averaging

\thispagestyle{empty}\setcounter{page}0

abstractThis paper proposes a new estimator for selecting weights to average over least squares estimates obtained from a set of models. Our proposed estimator builds on the Mallows model average (MMA) estimator of hansen2007least, but, unlike MMA, simultaneously controls for location bias and regression error through a common constant. We show that our proposed estimator-- the mean-shift Mallows model average (MSA) estimator-- is asymptotically optimal to the original MMA estimator in terms of mean squared error. A simulation study is presented, where we show that our proposed estimator uniformly outperforms the MMA estimator. {\em Keywords}: Model averaging, Mallows criterion, series estimators, optimality.

Introduction

The goal of model averaging, where several models are averaged over, is to reduce estimation variance while controlling for bias arising from omitted variables (regression error). hansen2007least proposed the Mallows model averaging (MMA) estimator, which is shown to be asymptotically optimal with regard to the fitted estimates achieving the minimum squared error in a class of discrete model averaging estimators. However, while linear averaging (including MMA) can minimize regression error, it cannot simultaneously reduce regression error and location bias; the bias arising from non-zero expectation in the misspecification bias. In this paper, we propose a novel model averaging estimator that controls for location bias while reducing regression error. We call this new estimator the mean-shift Mallows model averaging (MSA) estimator. We relax the condition of discrete estimators to continuous estimators, similar to that of wan2010least, though under weaker conditions. We show that the proposed MSA estimator is asymptotically optimal to the MMA estimator.

Model averaging (or forecast combination/ensemble learning) has been a staple and necessary tool in the field of econometrics, statistics, and machine learning. Since the seminar paper by BatesGranger1969, the potency of model averaging has been proven over multiple applications and contexts. Recently, the field has received a surge of interest across many disciplines due to an increase in usage and availability of full forecast densities arising from more complex models. In the statistical literature, there is a large literature on model averaging, particularly for Bayesian methods. For example, Bayesian model averaging raftery1997bayesian,hoeting1999bayesian has become standard for Bayesians dealing with model uncertainty. In machine learning, ensemble methods, including bagging breiman1996bagging, boosting schapire2003boosting, and stacking dvzeroski2004combining, which are all linear model averaging (often equal weight), have been used extensively to alleviate overfitting. In econometrics, increased usage in density forecasts has stimulated the field for more advanced strategies for model combination Terui2002,HallMitchell2007,Amisano2007,Hooger2010,Kascha2010,Geweke2011,Geweke2012,Billio2012,Billio2013,Aastveit2014,Fawcett2014,Pettenuzzo2015,Negro2016,yao2018using,aastveit2018evolution,Aastveit2015, with success in applications across many subdisciplines of economics.

While practical success of model averaging is abundant, theoretical considerations, in particular with regard to estimation of averaging weights, have been somewhat limiting. In terms of theoretical results for averaging models over model selection, BatesGranger1969 show that a combination of two forecasts can yield improved forecasts over the two individual forecasts, in terms of mean squared forecast error. In the machine learning literature, louppe2014understanding show that, for averaging regression trees (i.e. random forests), ensemble learning lowers variance while maintaining bias. In the Bayesian literature, when the true model is nested in the candidate models ($\mathcal{M}$-closed, as it is referred, following bernardo2009bayesian), BMA is known to asymptotically converge to the true model. However, under an $\mathcal{M}$-open setting, where the true model is not nested, BMA tends to converge to the “wrong" model, which is often not even the “best" model.

In terms of estimating the weights for model averaging, hansen2007least showed that minimizing the Mallows criterion is asymptotically optimal in a class of discrete model average estimators in terms of mean squared error. wan2010least relaxes this assumption of discrete estimators to continuous estimators, and show that the optimality holds. Further theoretical results are developed for more complex settings hansen2012jackknife,zhang2013model. While the MMA estimator is optimal for model averaging, it cannot simultaneously minimize the bias arising from omitted variables and location bias, due to its limited parameter dimension.

Building on the MMA estimator in hansen2007least, we propose a new model averaging estimator that controls for both biases: the mean-shift Mallows model averaging estimator (MSA). Our main contribution is to show that our proposed MSA estimator is asymptotically optimal to the MMA estimator, achieving lower risk. While our MSA estimator is frequentist by nature, our motivation is, nonetheless, Bayesian. In particular, our development is inspired by the recent practical and theoretical success of Bayesian predictive synthesis (BPS); a coherent Bayesian framework to combine information-- in particular, density forecasts-- from multiple sources mcalinn2019dynamic,mcalinn2017multivariate,takanashi2019predictive. Part of the success of BPS is in considering forecasts as “data" to be used in prior-posterior updating, treating them as latent factors. The flexibility of this framework, of which MMA is a special case, has lead to development of more elaborate model averaging approaches. One of which is to control for location bias arising from misspecification by incorporating a common constant, which shifts the mean to counteract the bias. This idea is central to our proposed MSA estimator.

Section. (ref) discusses the setup as well as review the results from hansen2007least. Section. (ref) introduces the mean-shift Mallows model averaging estimator and its sampling properties. Section. (ref) provides finite sample evidence in favor of the proposed MSA estimator.

Model averaging under misspecification

Let $\left\{y_{i}\right\} _{i\in n}\in\mathbb{R}$ and $\mbox{\boldmath$x$}$ be a countably infinite vector. The data generating process hansen2007least is

align[align omitted — 163 chars of source]

Note that the linearity of the DGP is not restrictive, as it includes series expansions. Consider each model, $m=1,2,...,$ uses the first $k_{m}$ elements of $x_{ji}$, where $0<k_1<k_2<...$, to construct an approximating model. This implies that the models are ordered, though this is not problematic when there are series expansions.

The misspecification bias (regression error) is

align[align omitted — 62 chars of source]

which is the component in the DGP that is never captured by any of the models, thus, under this setup, all models are misspecified. The predictive mean is,

align[align omitted — 128 chars of source]

where $P_{m}=X_{m}\left(X_{m}^{\top}X_{m}\right)^{-1}X_{m}^{\top}$ is the projection matrix and $\hat{\mbox{\boldmath$\theta$}}_{m}$ is the vector of least squares estimates. The full model, $M=M_{n}\leqq n$, is an integer for which $X_{k_{M}}^{\top}X_{k_{M}}$ is invertible.

The weights for model averaging is specified as

align[align omitted — 82 chars of source]

with the resulting linear averaged predictive mean being

align[align omitted — 80 chars of source]

Then, we have the following properties:

subequations\begin{align} P\left(W\right)&=\sum_{m=1}^{m}w_{m}P_{m},\\ P\left(W\right) & is symmetric but not idempotent: P^{\top}\left(W\right)=P\left(W\right),\quad P^{2}\left(W\right)\neq P\left(W\right)\\ tr\left\{ P\left(W\right)\right\} &=\sum_{m=1}^{m}w_{m}k_{m}\\ tr\left\{ P\left(W\right)P\left(W\right)\right\} &=\sum_{m=1}^{m}\sum_{l=1}^{m}w_{m}w_{l}\min\left(k_{l},k_{m}\right)=W^{\top}\varGamma_{m}W\\ tr\left\{ P_{m}\right\} & =k_{m}\\ tr\left\{ P_{m}P_{\ell}\right\} & =tr\left\{ P_{\min\left(k_{\ell},k_{m}\right)}\right\} =\min\left(k_{\ell},k_{m}\right)\\ \textrm{tr}\left\{ P\left(W\right)\right\} & =\textrm{tr}\left\{ \sum_{m=1}^{m}w_{m}P_{m}\right\} =\sum_{m=1}^{m}w_{m}k_{m}=k\left(W\right)\\ \textrm{tr}\left\{ P\left(W\right)P\left(W\right)\right\} & =\sum_{m=1}^{m}\sum_{\ell=1}^{m}w_{m}w_{\ell}\min\left(k_{\ell},k_{m}\right)=W^{\top}\varGamma_{m}W. \end{align}

Define the risk (i.e. the conditional squared error), $R\left(W\right)=\mathbb{E}\left[L_{n}\left(W\right)\left|X\right.\right]$, as

alignat*{1} & R\left(W\right)\\ = & \left[\begin{array}{c} w_{1}\\ w_{2}\\ w_{3}\\ \vdots\\ w_{m} \end{array}\right]^{\top}\left(\underbrace{\left[\begin{array}{ccccc} a_{1} & a_{2} & a_{3} & \cdots & a_{m}\\ a_{2} & a_{2} & a_{3} & \cdots & a_{m}\\ a_{3} & a_{3} & a_{3} & \cdots & a_{m}\\ \vdots & \vdots & \vdots & \ddots & a_{m}\\ a_{m} & a_{m} & a_{m} & \cdots & a_{m} \end{array}\right]}_{A_{m}}+\sigma^{2}\underbrace{\left[\begin{array}{ccccc} k_{1} & k_{1} & k_{1} & \cdots & k_{1}\\ k_{1} & k_{2} & k_{2} & \cdots & k_{2}\\ k_{1} & k_{2} & k_{3} & \cdots & k_{3}\\ \vdots & \vdots & \vdots & \ddots & \vdots\\ k_{1} & k_{2} & k_{3} & \cdots & k_{m} \end{array}\right]}_{\varGamma_{m}}\right)\left[\begin{array}{c} w_{1}\\ w_{2}\\ w_{3}\\ \vdots\\ w_{m} \end{array}\right]\\ & =W^{\top}\left(A_{m}+\sigma^{2}\varGamma_{m}\right)W,

where

align*[align* omitted — 115 chars of source]

following hansen2007least (Lemma 2). This shows that the risk is a quadratic function with regard to the weight vector, $W$. From this, we have

align*[align* omitted — 177 chars of source]

The residual from the linear projection can be written as

align*[align* omitted — 372 chars of source]

When $\ell\leqq m$, the residual is

alignat*{1} P_{\ell}P_{m} & =P_{\ell}\\ \left(I-P_{m}\right)b_{\ell} & =\left(I-P_{m}\right)b_{m}.

The Mallows criterion for the model average estimator is

alignat*{1} C_{n}\left(W\right) & =W^{\top}\left(y-x\hat{\beta}\right)^{\top}\left(y-x\hat{\beta}\right)W+2\sigma^{2}K^{\top}W,

where

align*[align* omitted — 61 chars of source]

is the vector of model dimensions and the collection of residuals is

align*[align* omitted — 121 chars of source]

For the MMA estimator, the weight vector is selected by minimizing the Mallows criterion. This has the expectation hansen2007least,

align*[align* omitted — 115 chars of source]

The empirical weight vector is \[ \hat{W}_{n}=\arg\min_{W}C_{n}\left(W\right). \] for which hansen2007least showed that this criterion is optimal, as we summarize below.

Limiting the weight vector to a finite $N$ number of discrete values, where the possible values, $W\in\left[0,1\right]$, are the $N+1$, \[ w_{m}\in\left\{ 0,\frac{1}{N},\frac{2}{N},\cdots,1\right\}. \] The set of possible discrete weight vectors is denoted as $H_{n}\left(N\right)$. When $n\rightarrow\infty$, assume \[ \xi_{n}=\inf_{W\in H_{n}}R_{n}\left(W\right)\rightarrow\infty \] almost surely. Additionally, assume \[ \mathbb{E}\left[\left|e_{i}\right|^{4\left(N+1\right)}\left|x_{i}\right.\right]\leqq\kappa<\infty. \] Then, the following holds hansen2007least: \[ \frac{L_{n}\left(\hat{W}_{N}\right)}{{\displaystyle \inf_{W\in H_{n}\left(W\right)}}L_{n}\left(W\right)}\rightarrow1. \]

The theoretical results from hansen2007least shows that choosing $\hat{W}$ by minimizing $C_{n}\left(W\right)$ minimizes the loss (but not the risk), as well as achieve the optimal rate. However, limiting the weight vector to a finite $N$ number of discrete values is somewhat restrictive, which we will later relax.

While the MMA estimator is optimal for a class of discrete model averaging estimators, it does not mean that there is no room for improvement. This can be easily seen when we decompose the loss as

alignat*{1} \left(\mu\left(W\right)-\mu\right)^{\top}\left(\mu\left(W\right)-\mu\right) & =\left(\mu\left(W\right)-\mathbb{E}\left[y\left|\mathcal{F}_{t}\right.\right]+\mathbb{E}\left[y\left|\mathcal{F}_{t}\right.\right]-\mu\right)^{\top}\left(\mu\left(W\right)-\mathbb{E}\left[y\left|\mathcal{F}_{t}\right.\right]+\mathbb{E}\left[y\left|\mathcal{F}_{t}\right.\right]-\mu\right)\\ & =\left(\mu\left(W\right)-\mathbb{E}\left[y\left|\mathcal{F}_{t}\right.\right]\right)^{2}+\left(\mathbb{E}\left[y\left|\mathcal{F}_{t}\right.\right]-\mu\right)^{2}+2\left(\mu\left(W\right)-\mathbb{E}\left[y\left|\mathcal{F}_{t}\right.\right]\right)\left(\mathbb{E}\left[y\left|\mathcal{F}_{t}\right.\right]-\mu\right)\\ & =\left(\mu\left(W\right)-\mathbb{E}\left[y\left|\mathcal{F}_{t}\right.\right]\right)^{2}+\left(\mathbb{E}\left[y\left|\mathcal{F}_{t}\right.\right]-\mu\right)^{2}.

Since the linear projection, $P_{m}\mu$, will ultimately not be the conditional expectation of $y$ given $X_{m}$ and the averaged $P\left(W\right)\mu$ will not be the conditional expectation of $y$ given all the available information, this will leave $\left(\mu\left(W\right)-\mathbb{E}\left[y\left|\mathcal{F}_{t}\right.\right]\right)^{2}$ to be improved.

This improvement can be expound further by decomposing the DGP and identifying the source of possible improvement. Consider decomposing the DGP into the full model, $M$, and its misspecification bias; \[ y_{i}=\sum_{j=1}^{k_{M}}\theta_{j}x_{ji}+\sum_{j=k_{M}+1}^{\infty}\theta_{j}x_{ji}+\varepsilon_{i}. \] Here, the covariates for $M$, $\left\{ x_{j}\right\} _{j=1,\cdots,K_{M}}$, and the covariates for misspecification bias, $\left\{ x_{j}\right\} _{j=K_{M},\cdots,\infty}$, are considered independent. The conditional expectation of the full model for $y_{i}$, given all observable covariates, $\left\{ x_{j}\right\} _{j=1,\cdots,k_{M}}$, is \[ \mathbb{E}\left[y\left|\left\{ x_{j}\right\} _{j=1,\cdots,k_{M}}\right.\right]=\sum_{j=1}^{k_{M}}\theta_{j}x_{ji}. \] The conditional expectation of the misspecification bias of $M$, due to independence, is \[ \mathbb{E}\left[b_{Mi}\right]=\mathbb{E}\left[\sum_{j=k_{M}+1}^{\infty}\theta_{j}x_{ji}\right]. \]

If each $\left\{ x_{j}\right\} _{j=1,\cdots,\infty}$ are i.i.d., then $\mathbb{E}\left[b_{Mi}\right]$ does not depend on the sampling order $i$, thus we denote it as $\alpha_{M}$. The best predictive value, in terms of MSE, can be written as, \[ \alpha_{M}+\sum_{j=1}^{k_{M}}\theta_{j}x_{ji}, \] which is to say that the DGP can be decomposed into its location bias, $\alpha_{M}$, and the optimal regression, $\sum_{j=1}^{k_{M}}\theta_{j}x_{ji}$. This can be written in matrix form (with elements corresponding to $i$) as \[ \alpha_{M}\mathbf{1}+X_{M}\theta_{M}. \] Note that $\mathbf{1}$ is a vector of ones.

The expectation of the misspecification bias for each model, $h$, using $\alpha_{h}$, is \[ \alpha_{h}=\mathbb{E}\left[b_{i,h}\right]=\mathbb{E}\left[\sum_{j=k_{h}+1}^{\infty}\theta_{j}x_{ji}\right]. \] The intercept, $\hat{\alpha}_{h}$, is the estimate for $\alpha_{h}$ for each model, $h$. The predictive value, $\hat{\mu}_{h}$, for each model is, \[ \hat{\mu}_{hi}=\hat{\alpha}_{h}+\sum_{j=1}^{k_{h}}\hat{\beta}_{j}x_{ji}, \] where we separate the intercept. The matrix form (with elements corresponding to $i$) is \[ \hat{\mu}_{h}=\hat{\alpha}_{h}\mathbf{1}+X_{h}^{-\alpha_{h}}\hat{\boldsymbol{\beta}}_{h}^{-\alpha_{h}}. \] Note that $\mathbf{1}$ is a vector of ones, and $X_{h}^{-\alpha_{h}},\hat{\boldsymbol{\beta}}_{h}^{-\alpha_{h}}$ is $X_{h}\hat{\boldsymbol{\beta}}_{h}$ minus the intercept.

Given the above, linear model averaging (including MMA) can be written as \[ \sum_{h=1}^{k_{M}}w_{h}\left(\hat{\alpha}_{h}\mathbf{1}+X_{h}^{-\alpha_{h}}\hat{\theta}_{h}^{-\alpha_{h}}\right). \] The weights, $w_{h}$ are chosen to be as close to the best predictive value of $y$; $\alpha_{M}+\sum_{j=1}^{k_{M}}\theta_{j}x_{ji}$. The optimal weight vector is $w_{h}^{*}$. Then, except in rare situations, we have,

align*[align* omitted — 199 chars of source]

Since the weights, $w_{h}$, must be selected by simultaneously minimizing the location bias and regression error, the weights on the location bias need not be zero. There might exist a vector of weights, $w$, that satisfies $\alpha_{M}=\sum_{h=1}^{k_{M}}w_{h}\hat{\alpha}_{h}$, but that will likely worsen the regression error. This is due to the dimension for linear averaging missing one dimension with regard to optimizing the averaging weights. Thus, under linear averaging, there is always a tradeoff between the location bias and regression error, where minimizing one leads to an increase in the other, making all linear averaging estimators (MMA included) suboptimal.

Mean-shift least squares averaging

To simultaneously minimize the location bias and regression error, we propose a new estimator that includes a common constant (i.e. an intercept) to the set of models, which we call the mean-shift Mallows model averaging (MSA) estimator. Our MSA estimator has the following predictive mean:

align[align omitted — 77 chars of source]

where the empirical weights, $\hat{W}$, and intercept, $\hat{\alpha}$, are estimated by minimizing the criterion we develop in this section.

When we include a common constant, the residual of the linear projection is

align*[align* omitted — 339 chars of source]

where

align*[align* omitted — 289 chars of source]

For the extra term $\left(1\right)$, arbitrarily fixing the weight vector, $W$, we have a quadratic function with regard to $\alpha$:

alignat*{1} & \alpha^{2}nW^{\top}\left[\begin{array}{cccc} 1 & 1 & \cdots & 1\\ 1 & \ddots & & \vdots\\ \vdots\\ 1 & \cdots & & 1 \end{array}\right]_{M\times M}W-\alpha W^{\top}\left[\begin{array}{cccc} \sum b_{1} & \sum b_{1} & \cdots & \sum b_{1}\\ \sum b_{2} & \sum b_{2} & \cdots & \sum b_{2}\\ \vdots & \vdots & \vdots & \vdots\\ \sum b_{M} & \sum b_{M} & \cdots & \sum b_{M} \end{array}\right]_{M\times M}W\\ & -\alpha W^{\top}\left[\begin{array}{cccc} \sum b_{1} & \sum b_{2} & \cdots & \sum b_{M}\\ \sum b_{1} & \sum b_{2} & \cdots & \sum b_{M}\\ \vdots & \vdots & \vdots & \vdots\\ \sum b_{1} & \sum b_{2} & \cdots & \sum b_{M} \end{array}\right]_{M\times M}W.

Since $\left(1\right)$ is downward convex and is zero when $\alpha=0$, the extra term can take a negative value. Then, the optimal $\alpha$ is

alignat*{1} FOC: & 0=2n\alpha W^{\top}\left[\begin{array}{cccc} 1 & 1 & \cdots & 1\\ 1 & \ddots & & \vdots\\ \vdots\\ 1 & \cdots & & 1 \end{array}\right]_{M\times M}W-W^{\top}\left[\begin{array}{cccc} \sum b_{1} & \sum b_{1} & \cdots & \sum b_{1}\\ \sum b_{2} & \sum b_{2} & \cdots & \sum b_{2}\\ \vdots & \vdots & \vdots & \vdots\\ \sum b_{M} & \sum b_{M} & \cdots & \sum b_{M} \end{array}\right]_{M\times M}W\\ & -W^{\top}\left[\begin{array}{cccc} \sum b_{1} & \sum b_{2} & \cdots & \sum b_{M}\\ \sum b_{1} & \sum b_{2} & \cdots & \sum b_{M}\\ \vdots & \vdots & \vdots & \vdots\\ \sum b_{1} & \sum b_{2} & \cdots & \sum b_{M} \end{array}\right]_{M\times M}W\\ \alpha & =\frac{W^{\top}\left\{ \left[\begin{array}{cccc} \sum b_{1} & \sum b_{1} & \cdots & \sum b_{1}\\ \sum b_{2} & \sum b_{2} & \cdots & \sum b_{2}\\ \vdots & \vdots & \vdots & \vdots\\ \sum b_{M} & \sum b_{M} & \cdots & \sum b_{M} \end{array}\right]_{M\times M}+\left[\begin{array}{cccc} \sum b_{1} & \sum b_{2} & \cdots & \sum b_{M}\\ \sum b_{1} & \sum b_{2} & \cdots & \sum b_{M}\\ \vdots & \vdots & \vdots & \vdots\\ \sum b_{1} & \sum b_{2} & \cdots & \sum b_{M} \end{array}\right]_{M\times M}\right\} W}{2nW^{\top}\left[\begin{array}{cccc} 1 & 1 & \cdots & 1\\ 1 & \ddots & & \vdots\\ \vdots\\ 1 & \cdots & & 1 \end{array}\right]_{M\times M}W}.

Adding the mean-shift $\alpha$ to linear averaging, the component for location bias can be written as \[ \alpha_{M}=\sum_{h=1}^{k_{M}}w_{h}^{*}\hat{\alpha}_{h}+\alpha, \] which uniformly improves the total error by fully controlling for location bias without compromising minimizing regression error. Additionally, adding $\alpha$ changes the optimal solution with regard to the weights $w$, which possibly improves the regression error as well.

Criterion

Since the loss is defined as

align*[align* omitted — 431 chars of source]

the expectation is

align*[align* omitted — 219 chars of source]

Then, the mean-shift Mallows criterion is

align*[align* omitted — 569 chars of source]

which can be written as

align*[align* omitted — 382 chars of source]

Comparing the criterion with and without an intercept, there exists an $\alpha$ that satisfies \[ \min_{W,\alpha}C_{n}\left(W,\alpha\right)\leqq\min_{W}C_{n}\left(W\right). \] This can be seen by setting $\alpha=0$ with regard to $\left(W,\alpha\right)$ that satisfies $C_{n}\left(W\right)\leqq C_{n}\left(W,\alpha\right)$.

Optimality

Given the extension, we now show that our proposed MSA is optimal to the MMA estimator in hansen2007least. We set the parameter space $K$ for which the intercept can take values, \[ K_{n}=\left\{ \alpha\left|\min_{W}C_{n}\left(W,\alpha\right)\leqq\min_{W}C_{n}\left(W\right)\right.\right\}. \] Here, $\left\{ K_{n}\right\} $ is on $\mathbb{R}$ including zero. Additionally, we assume \[ \mathbb{E}\left[\left|e_{i}\right|^{4\left(N+1\right)}\left|x_{i}\right.\right]\leqq\kappa<\infty. \]

When $n\rightarrow\infty$, we assume, \[ \xi_{n}\triangleq{ \inf_{

array[array omitted — 58 chars of source]

}}R_{n}\left(W,\alpha\right)\rightarrow\infty, \] almost surely. We additionally assume, for each point, $w\in H_{n}$,

equation[equation omitted — 111 chars of source]

This condition is weaker than that of wan2010least. Compared to the conditions in hansen2007least, \[ \inf_{w\in H_{n}}R_{n}\left(w\right)\rightarrow\infty, \] this might seem somewhat restrictive. However, we only need to add the following condition, \[ \frac{R_{n}\left(w,\alpha\right)}{{ \inf_{

array[array omitted — 58 chars of source]

}}R_{n}\left(w,\alpha\right)}<\infty,^{\forall}w\in H_{n} \] to the above, which is a reasonable condition to add.

proposition{\it Given the above assumptions, we have,} \[ \frac{L_{n}\left(\hat{W}_{N},\hat{\alpha}\right)}{{\displaystyle \inf_{\begin{array}{c} W\in H_{n}\left(W\right)\\ \alpha\in K_{n} \end{array}}}L_{n}\left(W,\alpha\right)}\rightarrow1. \]

The proof roughly follows that of li1987asymptotic and hansen2007least, though the loss, risk, and criterion are dependent on the continuous parameter, $\left(W,\alpha\right)$, which makes the proof challenging. Note that wan2010least provides the conditions to relax the discreteness of the weights. In contrast, we prove optimality using the convexity lemma by utilizing the fact that the loss and criterion is a convex function. This allows us to prove optimality with a weaker condition to that of wan2010least.

We first decompose $C_{n}\left(W,\alpha\right)$ as

alignat*{1} C_{n}\left(W,\alpha\right)= & \left(y-\hat{\mu}\left(W\right)-\alpha1\right)^{\top}\left(y-\hat{\mu}\left(W\right)-\alpha1\right)+2\sigma^{2}K^{\top}W\\ = & \left(\mu-\hat{\mu}\left(W\right)-\alpha1\right)^{\top}\left(\mu-\hat{\mu}\left(W\right)-\alpha1\right)\\ & +2\sigma^{2}K^{\top}W-2\varepsilon^{\top}\left(\mu-\hat{\mu}\left(W\right)-\alpha1\right)\\ & +\varepsilon^{\top}\varepsilon\\ = & L_{n}\left(W,\alpha\right)+2\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle \\ & +2\sigma^{2}K^{\top}W-2\left\langle \varepsilon,P\left(W\right)\varepsilon\right\rangle +\varepsilon^{\top}\varepsilon.

Since the $\varepsilon^{\top}\varepsilon$ term does not affect model selection, this is equivalent to minimizing, \[ L_{n}\left(W,\alpha\right)+2\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle +2\sigma^{2}K^{\top}W-2\left\langle \varepsilon,P\left(W\right)\varepsilon\right\rangle. \]

To prove optimality, we must show \[ { \min_{W,\alpha}}\frac{C_{n}\left(W,\alpha\right)}{{ \inf_{

array[array omitted — 58 chars of source]

}}L_{n}\left(W,\alpha\right)}\rightarrow1, \] which implies that \[ P\left(\sup_{

array[array omitted — 58 chars of source]

}\left|\frac{C_{p}\left(h\right)}{{ \inf_{

array[array omitted — 58 chars of source]

}}L_{n}\left(W,\alpha\right)}-1\right|>\delta\right)\rightarrow0 \] must uniformly hold with regard to $W\in H_{n}$ and $\alpha\in K_{n}$.

Since $\frac{C_{n}\left(W,\alpha\right)}{L_{n}\left(W,\alpha\right)}$ can be written as \[ \frac{L_{n}\left(W,\alpha\right)+2\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle +\left\{ 2\sigma^{2}K^{\top}W-2\left\langle \varepsilon,P\left(W\right)\varepsilon\right\rangle \right\} }{R_{n}\left(W,\alpha\right)}\frac{R_{n}\left(W,\alpha\right)}{L_{n}\left(h\right)} \] and each term is

alignat*{1} \left|\frac{2\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle }{R_{n}\left(W,\alpha\right)}\right| & \leqq\left|\frac{2\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle }{\xi_{n}}\right|\\ \frac{\left|2\sigma^{2}K^{\top}W-2\left\langle \varepsilon,P\left(W\right)\varepsilon\right\rangle \right|}{R_{n}\left(W,\alpha\right)} & \leqq\frac{\left|2\sigma^{2}K^{\top}W-2\left\langle \varepsilon,P\left(W\right)\varepsilon\right\rangle \right|}{\xi_{n}}\\ \left|\frac{L_{n}\left(W,\alpha\right)}{R_{n}\left(W,\alpha\right)}-1\right| & \leqq\frac{\left|L_{n}\left(W,\alpha\right)-R_{n}\left(W,\alpha\right)\right|}{\xi_{n}},

we need to show the following to prove optimality:

description\[ \sup_{\begin{array}{c} W\in H_{n}\left(W\right)\\ \alpha\in K_{n} \end{array}}\left|\frac{2\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle }{\xi_{n}}\right|\rightarrow0, \]\[ \sup_{\begin{array}{c} W\in H_{n}\left(W\right)\\ \alpha\in K_{n} \end{array}}\frac{\left|2\sigma^{2}K^{\top}W-2\left\langle \varepsilon,P\left(W\right)\varepsilon\right\rangle \right|}{\xi_{n}}\rightarrow0, \]\[ \sup_{\begin{array}{c} W\in H_{n}\left(W\right)\\ \alpha\in K_{n} \end{array}}\frac{\left|L_{n}\left(W,\alpha\right)-R_{n}\left(W,\alpha\right)\right|}{\xi_{n}}\rightarrow0. \]

To show uniform convergence, it is sufficient to show pointwise convergence, following the below convexity lemma:

theoremrockafellar1970convex,pollard1991asymptotics. {\it Consider the real convex function $\left\{ \rho_{n}\left(\theta\right)\left|\theta\in\Theta\right.\right\} $ with regard to $\theta$, then $\Theta\subset\mathbb{R}^{d}$ is a convex open set. If $\rho_{n}\left(\theta\right)$ pointwise converges to $\rho$, then $\rho_{n}\left(\theta\right)$ compact converges on $\Theta$: \[ \sup_{\theta\in K\subset\Theta}\left|\rho_{n}\left(\theta\right)-\rho\left(\theta\right)\right|\rightarrow0,\ prob. \] Note, it is necessary that $\rho$ is a convex function.}

{\it Proof of} {\bf 2.2:} From Chebyshev's inequality, we have \[ P\left(\left|\frac{2\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle }{\xi_{n}}\right|>\delta\right)\leqq\frac{4\mathbb{E}\left[\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle ^{2}\right]}{\delta^{2}\xi_{n}^{2}}, \] which we evaluate the right hand side. From the inequality in whittle1960bounds, the following \[ \mathbb{E}\left[\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle ^{2}\right]\leqq2^{2}C\left(s\right)\sigma\left\Vert A_{m}\mu-\alpha1\right\Vert ^{2} \] holds, thus, \[ \frac{4\mathbb{E}\left[\left\langle \varepsilon,A_{m}\mu-\alpha1\right\rangle ^{2}\right]}{\delta^{2}\xi_{n}^{2}}\leqq C\frac{1}{\delta^{2}}\frac{\left\Vert A_{m}\mu-\alpha1\right\Vert ^{2}}{\xi_{n}^{2}}, \] where $C$ is some constant. Now, since \[ \left\Vert A_{m}\mu-\alpha1\right\Vert ^{2}\leqq R_{n}\left(W,\alpha\right), \] we have \[ C\frac{1}{\delta^{2}}\frac{\left\Vert A_{m}\mu-\alpha1\right\Vert ^{2}}{\xi_{n}^{2}}\leqq C\frac{1}{\delta^{2}}\frac{R_{n}\left(W,\alpha\right)}{\xi_{n}^{2}}. \] Therefore, from assumption (ref), we prove {\bf 2.2}.

{\it Proof of} {\bf 2.3:} Similarly to {\bf 2.2}, applying Chebyshev's inequality, we evaluate the moment of the right hand side. Applying the quadric version of Whittle's inequality, we have \[ \mathbb{E}\left[\left|2\sigma^{2}K^{\top}W-2\left\langle \varepsilon,P\left(W\right)\varepsilon\right\rangle \right|^{2}\right]\leqq C\left(\sigma^{2}\textrm{tr}\left\{ P\left(W\right)P\left(W\right)\right\} \right), \] where, $C$ is a constant. Additionally, since \[ \sigma^{2}\textrm{tr}\left\{ P\left(W\right)P\left(W\right)\right\} \leqq R_{n}\left(W,\alpha\right), \] we prove {\bf 2.3}.

{\it Proof of} {\bf 2.4:} For this, it is sufficient to show convergence for the following:

alignat*{1} \frac{\left|L_{n}\left(W,\alpha\right)-R_{n}\left(W,\alpha\right)\right|}{\xi_{n}} & =\frac{\left|L_{n}\left(W,\alpha\right)-R_{n}\left(W,\alpha\right)\right|}{\xi_{n}}\\ L_{n}\left(h\right) & =\left\Vert \mu-\mu\left(W\right)-\alpha1\right\Vert ^{2}+2\left\langle \mu-\alpha1^{\top},P\left(W\right)\varepsilon\right\rangle +\left\Vert P\left(W\right)\varepsilon\right\Vert ^{2}\\ R_{n}\left(h\right) & =\left\Vert \mu-\mu\left(W\right)-\alpha1\right\Vert ^{2}+\sigma^{2}tr\left\{ P\left(W\right)P\left(W\right)\right\},

thus,

alignat*{1} \sup_{\left(W,\alpha\right)\in H_{n}}\frac{\left|\left\langle \mu-\alpha1^{\top},P\left(W\right)\varepsilon\right\rangle \right|}{\xi_{n}} & \rightarrow0\\ \sup_{\left(W,\alpha\right)\in H_{n}}\frac{\left|\left\Vert P\left(W\right)\varepsilon\right\Vert ^{2}-\sigma^{2}tr\left\{ P\left(W\right)P\left(W\right)\right\} \right|}{\xi_{n}} & \rightarrow0.

Noting that

alignat*{1} \left\langle \mu-\alpha1^{\top},P\left(W\right)\varepsilon\right\rangle & =\left\langle P\left(W\right)\left(\mu-\alpha1^{\top}\right),\varepsilon\right\rangle \\ \left\Vert P\left(W\right)\left(\mu-\alpha1^{\top}\right)\right\Vert ^{2} & \leqq\lambda\left\{ P\left(W\right)\right\} ^{2}\left\Vert \mu-\alpha1^{\top}\right\Vert ^{2},

the first equation is equivalent to {\bf 2.2}, and the second equation, utilizing the quadratic version of the Whittle's inequality to evaluate the expectation, \[ \mathbb{E}\left[\left|\left\Vert P\left(W\right)\varepsilon\right\Vert ^{2}-\sigma^{2}\textrm{tr}\left\{ P\left(W\right)P\left(W\right)\right\} \right|^{2}\right] \] and further noting that

alignat*{1} \left\Vert P\left(W\right)\varepsilon\right\Vert ^{2} & =\left\langle P\left(W\right)P\left(W\right)\varepsilon,\varepsilon\right\rangle \\ tr\left\{ P\left(W\right)P\left(W\right)\right\} ^{2} & \leqq\lambda\left\{ P\left(W\right)\right\} ^{2}tr\left\{ P\left(W\right)P\left(W\right)\right\},

for some constant, $C$, we have the following: \[ \mathbb{E}\left[\left|\left\Vert P\left(W\right)\varepsilon\right\Vert ^{2}-\sigma^{2}\textrm{tr}\left\{ P\left(W\right)P\left(W\right)\right\} \right|^{2}\right]\leqq CR_{n}\left(W,\alpha\right). \] The rest is the same as {\bf 2.3}. Q.E.D.

Finite sample investigation

We investigate the finite sample mean squared error of our proposed MSA estimator to that of the original MMA estimator in a simple simulation setting similar to hansen2007least. Our data generating process is a slight modification of eq. (ref), where the misspecification bias is positive:

align*[align* omitted — 127 chars of source]

where all $x_{ji}$ are independent and identically distributed $N(0,1)$, the error, $\varepsilon_{i}$, is $N(0,1)$ and independent of $x_{ji}$. The parameters are determined by the rule, $\theta_{j}=c\sqrt{2\alpha}j^{-\alpha -1/2}$, where $c=R^2/(1-R^2)$. Though not reported, the results are robust to alternative specifications, except when the bias has expectation zero (i.e. no location bias), at which MSA is equivalent to MMA.

For the finite sample investigation, we vary the sample size by $n=50,150,400, \textrm{ and } 1,000$ and $\alpha=0.5,1, \textrm{ and } 1.5$. The parameter $\alpha$ controls the rate of decline of the coefficients, $\theta_{j}$, where a larger $\alpha$ implies a steeper decline in $j$. The number of models is determined by $M=3n^{1/3}$, which, for our four sample sizes, means $M=11,16,22,\textrm{ and } 30$. We vary $c$ so as $R^2$ varies from 0.1 to 0.9.

Unlike hansen2007least, which compared the MMA estimator to AIC model selection, Mallows model selection, smoothed AIC, and smoothed BIC, we simply compare our MSA estimator to the MMA estimator, as hansen2007least convincingly show that the MMA estimator is superior to these methods. To evaluate the two estimators, we compute the log risk (expected squared error) for each and calculate the difference. The simulation is done for 1,000 iterations and averaged over.

The results of the finite sample investigation are shown in Fig. (ref). The three panels display the difference in log risk between MMA and MSA for the varying $\alpha$, where the difference is displayed on the $y$ axis and the population $R^2$ is displayed on the $x$ axis. The different colored lines are for the different sample sizes, $n$.

For all results, our proposed MSA estimator uniformly outperforms the MMA estimator to varying degrees. In the case of $\alpha=0.5$, the results are almost identical for all sample sizes. The relative performance of MSA increases as $R^2$ increases, which is expected as the increased $R^2$, under a low $\alpha$, implies a more robust bias component. This bias component is captured in MSA, leading to improved performances. As $\alpha$ increases, the results invert, in the sense that, for $\alpha=1$, the difference in log risk is bowed, while, for $\alpha=1.5$, the difference decreases as $R^2$ increases. This is not unexpected as, in hansen2007least, the performance of the MMA estimator decreases in these cases. If anything, for $\alpha=1$, the MSA estimator seems more resilient for higher $R^2$. Over all variations, the MSA estimator is superior to the MMA estimator.

figure[figure omitted — 310 chars of source]