EconBase
← Back to paper

Simultaneous Mean-Variance Regression

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

79,216 characters · 20 sections · 46 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Simultaneous Mean-Variance Regression

abstractWe propose simultaneous mean-variance regression for the linear estimation and approximation of conditional mean functions. In the presence of heteroskedasticity of unknown form, our method accounts for varying dispersion in the regression outcome across the support of conditioning variables by using weights that are jointly determined with the mean regression parameters. Simultaneity generates outcome predictions that are guaranteed to improve over ordinary least-squares prediction error, with corresponding parameter standard errors that are automatically valid. Under shape misspecification of the conditional mean and variance functions, we establish existence and uniqueness of the resulting approximations and characterize their formal interpretation and robustness properties. In particular, we show that the corresponding mean-variance regression location-scale model weakly dominates the ordinary least-squares location model under a Kullback-Leibler measure of divergence, with strict improvement in the presence of heteroskedasticity. The simultaneous mean-variance regression loss function is globally convex and the corresponding estimator is easy to implement. We establish its consistency and asymptotic normality under misspecification, provide robust inference methods, and present numerical simulations that show large improvements over ordinary and weighted least-squares in terms of estimation and inference in finite samples. We further illustrate our method with two empirical applications to the estimation of the relationship between economic prosperity in 1500 and today, and demand for gasoline in the United States.

Keywords: Conditional mean and variance functions, simultaneous approximation, heteroskedasticity, robust inference, misspecification, convexity, influence function, ordinary least-squares, linear regression, dual regression.

Introduction

Ordinary least-squares (OLS) is the method of choice for the linear estimation and approximation of the conditional mean function (CMF). However, in the presence of heteroskedasticity the standard errors of OLS are inconsistent, and subsequent inference is therefore unreliable. As a way of achieving valid inference, practitioners instead often use the heteroskedasticity-corrected standard errors of Eicker (Eicker:1963, Eicker:1967), Huber:1967 and White:1980a. Although valid asymptotically, numerous limitations of this approach have been highlighted in the literature such as bias and sensitivity to outliers, incorrect size and low power of robust tests in finite samples (MKW:1985; CJ:1987; Chesher:1989; CA1991). These findings in turn generated a large number of proposals in order to reconcile the large-sample validity of the approach and its observed finite-sample limitations, surveyed in MacKinnon:2013.

The finite-sample limitations of OLS-based inference essentially originate from the fact that OLS assigns a constant weight to each observation in fitting the best linear predictor for the regression outcome. Hence the least-squares criterion does not account for the varying accuracy of the information available about the outcome across the covariate space. This yields point estimates and linear approximations that are sensitive to high-leverage points and outliers, which in turn generate biased estimates of the residuals' second moments used in the calculation of the robust variance-covariance matrix of OLS parameters. In finite samples, uniform weighting not only compromises the validity of OLS-based statistical inference in the presence of heteroskedasticity, but also the reliability of OLS point estimates.

In this paper, we propose simultaneous mean-variance regression (MVR) as an alternative to OLS for the linear estimation and approximation of CMFs. MVR characterizes the conditional mean and variance functions jointly, thereby providing a solution to the problems of estimation, approximation and inference in the presence of heteroskedasticity of unknown form with five main features. First, it incorporates information from the second conditional moment in the determination of the first conditional moment parameters. Second, simultaneity generates approximations with improved robustness properties relative to OLS. Third, the resulting approximations have a formal interpretation under shape misspecification of the conditional mean and variance functions. Fourth, MVR solutions are well-defined, i.e., exist and are unique. Fifth, standard errors of the corresponding estimator are automatically valid in the presence of heteroskedasticity of unknown form, and reduce to those of OLS under homoskedasticity.

The MVR criterion can be interpreted as a penalized weighted least-squares (WLS) loss function. The presence and the form of the penalty ensure global convexity of the objective function, so that MVR conditional mean and variance approximations are jointly well-defined. This differs from the usual WLS approach where a sequential procedure is followed, obtaining the weights first, and then implementing a weighted regression to determine the parameters of the linear specification. Our simultaneous approach allows us to give theoretical guarantees on the relative approximation properties of MVR and OLS. We use MVR to construct and estimate a new class of approximations of the conditional mean and variance functions, with improved robustness and precision in finite samples. We establish the interpretation of MVR approximations, we derive the asymptotic properties of the corresponding MVR estimator, and we give tools for robust inference. We also illustrate the practical benefits of the MVR estimator with extensive numerical simulations, and find very large finite-sample improvements over both OLS and WLS in terms of estimation performance and heteroskedasticity-robust inference.

This paper generalizes the results of SS:2018 for the primal problem of the dual regression estimator of linear location-scale models. We provide a unified theory allowing for a large class of scale functions. This paper is also related to the interpretation of OLS under misspecification of the shape of the CMF. OLS gives the minimum mean squared error linear approximation to the CMF, an important motivation for its use in empirical work (White:1980b; Chamberlain:1984; Angrist:Krueger:1999; AP:2008). MVR introduces a class of WLS approximations accounting for potential variation in the outcome across the support of conditioning variables, and with weights that have a clear interpretation under misspecification. Our approach thus complements the textbook WLS proposal of CT:2005 and Wooldridge (Wooldridge:2010, Wooldridge:2012) (see also RW:2017), who advocate the reweighting of OLS with generalized least-squares (GLS) weights and further correcting the standard errors for heteroskedasticity.

This paper makes three main contributions. First, we show existence and uniqueness of MVR solutions under general misspecification, thereby introducing a new class of location-scale models corresponding to MVR approximations. The results in SS:2018 did not cover the case of misspecified conditional mean and variance functions. Second, we establish favorable approximation and robustness properties of MVR relative to OLS. We show that MVR is a minimum WLS linear approximation to the CMF, with weights determined such that the MVR approximation improves over OLS in the presence of heteroskedasticity under the MVR loss. For our main specifications of the scale function, we further show that OLS root mean squared prediction error is an upper bound for the MVR weighted mean squared prediction error. We then extend this result to show that under a Kullback-Leibler information criterion (KLIC) the proposed MVR location-scale models weakly dominate the OLS location model, with strict improvement in the presence of heteroskedasticity. These results provide theoretical guarantees motivating the use of MVR over OLS, and are not shared by alternative WLS proposals. Third, we derive the asymptotic distribution of the MVR estimator under misspecification and provide robust inference methods. In particular we propose a robust one-step heteroskedasticity test that complements existing OLS-based tests (e.g., BP:1979; White:1980a; Koenker:1981).

The rest of the paper is organized as follows. Section (ref) introduces MVR under correct specification of conditional mean and variance functions. Section (ref) establishes the main approximation properties of MVR under misspecification, including existence and uniqueness. Section (ref) gives asymptotic theory. Section (ref) reports the results of an empirical application to the relationship between economic prosperity in 1500 and today, and illustrates the finite-sample performance of MVR with numerical simulations. All proofs of the main results are given in the Appendix. The online Appendix Spady and Stouli (2018) contains supplemental material, including an additional empirical application to demand for gasoline in the United States.

Simultaneous Mean-Variance Regression

The Mean-Variance Regression Problem

Given a scalar random variable $Y$ and a random $k\times1$ vector $X$ that includes an intercept, i.e., has first component $1$, denote the mean and standard deviation functions of $Y$ conditional on $X$ by $\mu(X):=E[Y\mid X]$ and $\sigma(X):=E[(Y-E[Y\mid X])^{2}\mid X]^{1/2}$, respectively. We start with a simplified setting where the conditional mean and variance functions take the parametric forms

equation[equation omitted — 95 chars of source]

for some positive scale function $t\mapsto s(t)$, and where the parameters $\beta_{0}$ and $\gamma_{0}$ belong to the parameter space $\Theta=\mathbb{R}^{k}\times\Theta_{\gamma}$, with $\Theta_{\gamma}=\{\gamma\in\mathbb{R}^{k}\,:\,\Pr[s(X'\gamma)>0]=1\}$. Two leading examples for the scale function are the linear and exponential specifications $s(t)=t$ and $s(t)=\exp(t)$, with domains $(0,\infty)$ and $\mathbb{R}$, respectively.

The parameter vector $\theta_{0}:=(\beta_{0},\gamma_{0})'$ is uniquely determined as the solution to the globally convex MVR population problem

equation[equation omitted — 154 chars of source]

When the functions $x\mapsto\mu(x)$ and $x\mapsto\sigma(x)^{2}$ satisfy model ((ref)), they are simultaneously characterized by problem ((ref)). As a consequence, MVR incorporates information on the dispersion of $Y$ across the support of $X$ in the determination of the mean parameter $\beta$. We show below that problem ((ref)) is formally equivalent to an infeasible sequential least-squares estimator of the conditional mean and variance functions for model ((ref)). Problem ((ref)) is a generalization of the dual regression primal problem introduced in SS:2018, for which the scale function is linear. Considering scale functions with domain the real line, such as the exponential function, allows the transformation of the dual regression primal problem into an unconstrained convex problem over $\Theta=\mathbb{R}^{2\times k}$.

Inspection of the first-order conditions confirms that $\theta_{0}$ is indeed a valid solution to problem ((ref)). Denoting the derivative of the scale function by $s_{1}(t):=\partial s(t)/\partial t$ and letting $e\left(Y,X,\theta\right):=(Y-X'\beta)/s(X'\gamma)$, the first-order conditions of ((ref)) are

align[align omitted — 160 chars of source]

These conditions are satisfied by $\theta_{0}$ since specification ((ref)) is equivalent to the location-scale model

equation[equation omitted — 145 chars of source]

Therefore, the parameter vector $\theta_{0}$ also satisfies the relations

align*[align* omitted — 178 chars of source]

which imply that $E[h(X)e(Y,X,\theta_{0})]=0$ and $E[h(X)\{e(Y,X,\theta_{0})^{2}-1\}]=0$ hold for any measurable function $x\mapsto h(x)$, and in particular for $h(X)=X$ and $h(X)=Xs_{1}(X'\gamma)$.

Formal Framework

Let $\mathcal{X}$ denote the support of $X$, and for a vector $u=(u_{1},\ldots,u_{k})'\in\mathbb{R}^{k}$, let $||\cdot||$ denote the Euclidean norm, i.e., $||u||=(u_{1}^{2}+\ldots+u_{k}^{2})^{1/2}$; we define a compact subset $\Theta^{c}\subset\Theta$ as \[ \Theta^{c}:=\left\{ \theta\in\Theta\,:\,||\theta||\leq C_{\theta}\:\textrm{and}\:\inf_{x\in\mathcal{X}}s(x'\gamma)\geq C_{s}\right\} , \] for some finite constant $C_{\theta}$ and some constant $C_{s}>0$, with interior set denoted $\textrm{int}(\Theta^{c})$. The second and third derivatives of the scale function $t\mapsto s(t)$ are denoted by $s_{j}(t):=\partial^{j}s(t)/\partial t^{j}$, $j=2,3$. We also denote the MVR objective function in ((ref)) by $Q(\theta):=E[\{e\left(Y,X,\theta\right)^{2}+1\}s(X'\gamma)/2]$.

Our first assumption specifies the class of scale functions we consider.

assumptionFor $a=0$ or $-\infty$, the scale function $s:(a,\infty)\rightarrow(0,\infty)$ is a three times differentiable strictly increasing convex function that satisfies $\lim_{t\rightarrow a}s(t)=0$ and $\lim_{t\rightarrow\infty}s(t)=\infty$.

Assumption (ref) encompasses several types of scale functions such as polynomial specifications $s(x'\gamma)=(x'\gamma)^{\alpha}$ with $a=0$ and $\Pr[X'\gamma>0]=1$, or exponential-polynomial specifications $s(x'\gamma)=\exp(x'\gamma)^{\alpha}$ with $a=-\infty$, for some $\alpha>0$. For $\alpha=1$, we recover the linear and exponential scale leading cases.

The next assumptions complete our formal framework.

assumptionThe conditional variance function $x\mapsto\sigma(x)^{2}$ is bounded away from $0$ uniformly in $\mathcal{X}$.
assumptionWe have (i) $E[Y^{4}]<\infty$ and $E||X||^{4}<\infty$, and, (ii) for all $\gamma\in\Theta_{\gamma}$, $E[\left\Vert X\right\Vert ^{4}s_{2}(X'\gamma)^{2}]<\infty$, $E[\left\Vert X\right\Vert ^{6}s_{3}(X'\gamma)^{2}]<\infty$ and $E[\left\Vert X\right\Vert ^{6}s_{1}(X'\gamma)^{2}s_{2}(X'\gamma)^{2}]<\infty$.
assumptionFor all $\gamma\in\Theta_{\gamma}$, $E[XX'/s(X'\gamma)]$ is nonsingular.

Assumptions (ref)-(ref) are sufficient conditions for global convexity of the MVR criterion over the parameter space $\Theta$, and therefore for problem ((ref)) to have a unique solution.

thmIf Assumptions (ref)-(ref) hold, and the conditional mean and variance functions of $Y$ given $X$ satisfy model ((ref)) a.s. with $\theta_{0}\in\textrm{int}(\Theta^{c})$, then $\theta_{0}$ is the unique minimizer of $Q(\theta)$ over $\Theta$.

Theorem (ref) applies when the conditional mean and variance functions are well-specified, and thus provides primitive conditions for identification of $\theta_{0}$ in the location-scale model ((ref)). This extends the uniqueness result in SS:2018 for objective ((ref)) with a linear scale function to the class of scale functions defined in Assumption (ref).

remIn the linear scale case, $s_{1}(t)=1$ and $s_{j}(t)=0$, $j=2,3$, so that Assumption (ref) reduces to Assumption (ref)(i). In the exponential scale case, $s_{j}(t)=\exp(t)$, $j=1,2,3$, so that Assumption (ref)(ii) reduces to the requirement that $E[\left\Vert X\right\Vert ^{6}\exp(X'\gamma)^{4}]$ be finite. This is satisfied for instance if $X$ is bounded.\qed

Simultaneous Mean-Variance Regression Interpretation

Problem ((ref)) is equivalent to an infeasible sequential least-squares estimator of conditional mean and variance functions. The first-order conditions of ((ref)) can also be written as

eqnarray[eqnarray omitted — 229 chars of source]

Given knowledge of $\gamma_{0}$, WLS regression of $Y$ on $X$ with weights $1/s(X'\gamma_{0})$ has first-order conditions ((ref)), with solution $\beta_{0}$. Moreover, given knowledge of $\beta_{0}$, nonlinear WLS regression of $(Y-X'\beta_{0})^{2}$ on $X$ with weights $1/s(X'\gamma_{0})^{3}$ and quadratic link function has first-order conditions ((ref)), and therefore solution $\gamma_{0}$.

propIf Assumptions (ref)-(ref) hold, and (i) $E[Y^{4}]<\infty$, $E[\left\Vert X\right\Vert ^{4}]<\infty$ and $E[s(X'\gamma)^{4}]<\infty$ for all $\gamma\in\Theta_{\gamma}$, and (ii) the conditional mean and variance functions of $Y$ given $X$ satisfy model ((ref)) a.s., then the MVR population problem ((ref)) is equivalent to the infeasible sequential estimator with first step \begin{equation} \beta_{0}=\arg\min_{\beta\in\Theta_{\beta}}E\left[\frac{1}{\sigma(X)}(Y-X'\beta)^{2}\right], \end{equation} and second step \begin{equation} \gamma_{0}=\arg\min_{\gamma\in\Theta_{\gamma}}E\left[\frac{1}{\sigma(X)^{3}}\left\{ (Y-X'\beta_{0})^{2}-s(X'\gamma)^{2}\right\} ^{2}\right]. \end{equation}

An immediate implication of the Law of Iterated Expectations and Proposition (ref) is that MVR implements simultaneous weighted linear regression of $\mu(X)$ on $X$ and weighted nonlinear regression of $\sigma(X)^{2}$ on $X$ by solving for $\beta$ and $\gamma$ such that the weighted residuals $(\mu(X)-X'\beta)/s(X'\gamma)$ and $\{\sigma(X)^{2}-s(X'\gamma)^{2}\}/s(X'\gamma)^{2}$ are simultaneously orthogonal to $X$ and $Xs_{1}(X'\gamma)$, respectively. Proposition (ref) thus establishes the simultaneous mean and variance regression interpretation of problem ((ref)).

Approximation Properties of MVR under Misspecification

Under misspecification, OLS provides the minimum mean squared error linear approximation to the CMF. For the proposed MVR criterion, existence of an approximating solution and the nature of the approximation are nontrivial when the shapes of the conditional mean and variance functions are misspecified. In this section, we first establish existence and uniqueness of a solution to the MVR problem under misspecification, and then characterize the interpretation and properties of the corresponding MVR approximations.

Existence and Uniqueness of an MVR Solution

Assumptions (ref)-(ref) are sufficient for characterizing the smoothness properties, shape, and behaviour on the boundaries of the parameter space of the MVR criterion $Q(\theta)$. Under these assumptions $\theta\mapsto Q(\theta)$ is continuous and its level sets are compact. Compactness of the level sets is a sufficient condition for existence of a minimizer in $\Theta$, and is a consequence of the explosive behaviour of the objective function at the boundaries of the parameter space. The objective $Q(\theta)$ is a coercive function over the open set $\Theta$, i.e., it satisfies \[ \lim_{||\theta||\rightarrow\infty}Q(\theta)=\infty,\quad\quad\lim_{\theta\rightarrow\partial\Theta}Q(\theta)=\infty, \] where $\partial\Theta$ is the boundary set of $\Theta$. Thus the MVR criterion is infinity at infinity, and for any sequence of parameter values in $\Theta$ approaching the boundary set $\partial\Theta$, the value of the objective is also driven towards infinity. Therefore, the level sets of the objective function have no limit point on their boundary, ruling out existence of a boundary solution, and continuity of $\theta\mapsto Q(\theta)$ is then sufficient to conclude that it admits a minimizer. Continuity and coercivity of the objective function are the two properties that guarantee existence of at least one minimizer in $\Theta$.\footnote{The boundary set of $\Theta$ may be empty, for instance for the exponential scale specification. In that case the coercivity property reduces to $\lim_{||\theta||\rightarrow\infty}Q(\theta)=\infty$.} Assumptions (ref)-(ref) are also sufficient for $\theta\mapsto Q(\theta)$ to be strictly convex, and therefore further ensure that $Q(\theta)$ admits at most one minimizer in $\Theta$.

thmIf Assumptions (ref)-(ref) hold, then there exists a unique solution $\theta^{*}\in\Theta$ to the MVR population problem ((ref)).

Theorem (ref) is the second main result of the paper. It establishes that the MVR problem ((ref)) has a well-defined solution, and an immediate corollary is the existence and uniqueness of the MVR location-scale representation \[ Y=X'\beta^{*}+s(X'\gamma^{*})e,\quad E[Xe]=0,\quad E[Xs_{1}(X'\gamma^{*})(e^{2}-1)]=0. \] This result clarifies further how MVR generalizes OLS by establishing the existence and the form of the MVR location-scale model when no shape restrictions are imposed on the conditional mean and variance functions. The OLS location model is a particular case with the scale function restricted to be a constant function.

sloppyAlthough a unique MVR approximation exists irrespective of the nature of the misspecification, the interpretation of the MVR approximating functions $x\mapsto(x'\beta^{*},s(x'\gamma^{*})^{2})$ depends on which of the conditional moment functions is misspecified. We distinguish two types of shape misspecification: \begin{enumerate} • Mean misspecification: the CMF $x\mapsto\mu(x)$ is misspecified. • Variance misspecification: only the conditional variance function $x\mapsto\sigma(x)^{2}$ is misspecified. \end{enumerate} The case when both the conditional mean and variance functions are misspecified is a particular case of mean misspecification.

Interpretation Under Mean Misspecification

The location-scale representation

equation[equation omitted — 132 chars of source]

provides a general expression for $Y$ in terms of its conditional mean and standard deviation functions, and is always valid, as long as first and second conditional moments exist. Substituting expression ((ref)) for $Y$ into the MVR objective function $Q(\theta)$ gives rise to a criterion for the joint approximation of $x\mapsto(\mu(x),\sigma(x)^{2})$.

The criterion $Q(\theta)$ can also be appropriately restricted in order to define the corresponding OLS approximations. Letting $\Theta_{\gamma,\textrm{LS}}=\{\gamma\in\mathbb{R}\,:\,s(\gamma)>0\}$, define $\Theta_{\textrm{LS}}=\mathbb{R}^{k}\times\Theta_{\gamma,\textrm{LS}}$. Upon setting $s(X'\gamma)=s(\gamma)$ in the MVR problem ((ref)), \[ (\beta_{\textrm{LS}},\gamma_{\textrm{LS}}):=\arg\min_{(\beta,\gamma)\in\Theta_{\textrm{LS}}}E\left[\frac{1}{2}\left\{ \left(\frac{Y-X'\beta}{s(\gamma)}\right)^{2}+1\right\} s(\gamma)\right] \] is a particular case of MVR. Since the OLS solution $\theta_{\textrm{LS}}:=(\beta_{\textrm{LS}},\gamma_{\textrm{LS}},0_{k-1})'$ belongs to the parameter space $\Theta$, uniqueness of $\theta^{*}$ implies that the OLS approximation of the conditional moment functions $x\mapsto(\mu(x),\sigma(x)^{2})$ cannot improve upon the MVR approximation, according to the MVR loss.

thmIf Assumptions (ref)-(ref) hold, then the MVR population problem ((ref)) has the following properties. (i) Problem ((ref)) is equivalent to the infeasible problem \begin{align} \min_{\theta\in\Theta} & \;\frac{1}{2}E\left[\left\{ \left(\frac{\mu(X)-X'\beta}{s(X'\gamma)}\right)^{2}+1\right\} s(X'\gamma)\right]+\frac{1}{2}E\left[\frac{\sigma(X)^{2}}{s(X'\gamma)}\right], \end{align} with first-order conditions \begin{eqnarray} E\left[X\left(\frac{\mu(X)-X'\beta}{s(X'\gamma)}\right)\right] & = & 0\\ E\left[Xs_{1}(X'\gamma)\left\{ \left(\frac{\mu(X)-X'\beta}{s(X'\gamma)}\right)^{2}+\left(\frac{\sigma(X)^{2}}{s(X'\gamma)^{2}}-1\right)\right\} \right] & = & 0. \end{eqnarray} (ii) The optimal value of problem ((ref)) satisfies $Q(\theta^{*})\leq Q(\theta_{\textrm{LS}})$, with equality if and only if $\theta^{*}=\theta_{\textrm{LS}}$.

Theorem (ref)(i) shows that under misspecification the function $x\mapsto x'\beta^{*}$ is an infeasible MVR approximation of the true CMF penalized by the mean ratio of the true variance over its standard deviation approximation. An equivalent formulation is

equation[equation omitted — 217 chars of source]

the penalized WLS interpretation of the MVR problem ((ref)).

The penalty term in ((ref)) is a functional of a weighted mean variance ratio of the true variance over its approximation. The first-order conditions ((ref))-((ref)) shed additional light on how the weights are determined as well as on the form of the penalty, by characterizing the optimality properties of MVR approximations. Because $X$ includes an intercept, when both functions $x\mapsto\mu(x)$ and $x\mapsto\sigma(x)^{2}$ are misspecified, $\beta^{*}$ and $\gamma^{*}$ are chosen such that the sum of the weighted mean squared error for the conditional mean and the mean variance ratio error is zero, balancing the two approximation errors. When the scale function is linear the two types of approximation error are equalized. For the exponential specification, the two types of approximation error weighted by $\exp(X'\gamma)$ are equalized. The MVR solution is thus determined by minimizing the weighted mean squared error for the conditional mean, while simultaneously setting the weighted mean variance ratio as close as possible to one.

Theorem (ref)(ii) formalizes the approximation guarantee of MVR in terms of the MVR criterion. For the linear and exponential scale function specifications, the improvement of the MVR solution relative to the OLS solution in MVR loss further guarantees that optimal weights are selected such that the weighted mean squared MVR prediction error for $Y$ is not larger than the root mean squared OLS prediction error.

corIf the scale function $t\mapsto s(t)$ is specified as $s(t)=t$ or $s(t)=\exp(t)$, then \[ E\left[\frac{1}{s(X'\gamma^{*})}(Y-X'\beta^{*})^{2}\right]\leq E\left[(Y-X'\beta_{\textrm{LS}})^{2}\right]^{\frac{1}{2}}, \] with equality if and only if $\theta^{*}=\theta_{\textrm{LS}}$.

Compared to OLS, improvement in MVR loss is also related to a key robustness property of MVR under a Kullback-Leibler measure of divergence. Define the scaled Gaussian density function

equation[equation omitted — 131 chars of source]

where $\phi(z)=\left(2\pi\right)^{-1/2}\exp(-z^{2}/2)$. The OLS solution maximizes a restricted version, $E[\log f_{\theta}\left(Y,X\right)]|_{\gamma_{-1}=0}$, of the expected log-likelihood $E[\log f_{\theta}\left(Y,X\right)]$ over $\Theta_{\textrm{LS}}$, where the components of $\gamma$ except the first are set to zero. The corresponding expected log-likelihood value $E[\log f_{\theta_{\textrm{LS}}}\left(Y,X\right)]$ is no greater than the value of the expected log-likelihood at the MVR solution:

equation[equation omitted — 146 chars of source]

The MVR solution $\theta^{*}$ formally corresponds to an improvement of the expected log-likelihood value over OLS, and therefore corresponds to a probability distribution that is KLIC closer to the true data probability distribution $f_{Y\mid X}(Y\mid X)$.

Define the quantity $\epsilon:=E\left[\log\left(\sigma_{\textrm{LS}}^{2}/s(X'\gamma^{*})^{2}\right)\right]$, which is positive as shown in Appendix (ref). Our next result summarises the key implication of ((ref)).

thmSuppose that $E[|\log f_{Y\mid X}(Y\mid X)|]<\infty$, $E\left[e(Y,X,\theta^{*})^{2}\right]\leq1+\epsilon$, and the scale function $t\mapsto s(t)$ is specified as $s(t)=t$ or $s(t)=\exp(t)$. Then the probability distribution $f_{\theta^{*}}$ corresponding to $\theta^{*}$ satisfies \begin{equation} E\left[\log\left(\frac{f_{Y\mid X}(Y\mid X)}{f_{\theta^{*}}(Y,X)}\right)\right]\leq E\left[\log\left(\frac{f_{Y\mid X}(Y\mid X)}{f_{\theta_{LS}}(Y,X)}\right)\right], \end{equation} with equality if and only if $\theta^{*}=\theta_{\textrm{LS}}$.

For the linear scale specification, the MVR first-order conditions include the constraint $E\left[e(Y,X,\theta^{*})^{2}-1\right]=0$ so that the bound on $E\left[e(Y,X,\theta^{*})^{2}\right]$ is satisfied by construction. For the exponential specification, the corresponding constraint is $E\left[\exp(X'\gamma^{*})\left\{ e(Y,X,\theta^{*})^{2}-1\right\} \right]=0$ which can result in a value for $E\left[e(Y,X,\theta^{*})^{2}\right]$ that differs from one under misspecification. The bound then characterizes the deviations from unit variance that preserve the validity of ((ref)).\footnote{For the numerical simulations in Section (ref) the typical sample mean of an estimate $e(Y,X,\hat{\theta})^{2}$ we observe is smaller than or equal to one, when the scale function is specified as $s(t)=\exp(t)$. It is an open question whether the bound $E\left[e(Y,X,\theta^{*})^{2}\right]\leq1+\epsilon$ can be binding.}

The approximation guarantee ((ref)) is a general result that holds under misspecification of the conditional mean and/or variance functions. When the mean is misspecified, it formally establishes that the MVR approximation for the mean corresponds to a better model than the OLS location model according to the classical KLIC for model selection (e.g., Akaike:1973; Sawa:1978). Similarly to the classical argument motivating the use of maximum likelihood (ML) under misspecification, Theorem (ref) thus provides an information-theoretic justification for the use of MVR and a formal characterization of the robustness to misspecification of MVR relative to OLS.

Interpretation Under Variance Misspecification

If the CMF is linear, Theorem (ref) has important additional implications for the robustness and optimality properties of MVR solutions. The $k$ orthogonality conditions ((ref)) are then sufficient to determine the scale parameter $\gamma^{*}$ since condition ((ref)) is uniquely satisfied by $\beta=\beta_{0}$. Thus in the classical particular case of the linear conditional mean model, the MVR solution for $\beta$ is fully robust to misspecification of the scale function. Consequently, when the CMF is correctly specified the OLS and MVR solutions for $\beta$ coincide. In the special case of linear scale specification, $Xs_{1}(X'\gamma)$ reduces to $X$. Because $X$ includes an intercept, the scale parameter $\gamma^{*}$ is then chosen such that the MVR conditional variance approximation also satisfies the remarkable property of zero mean variance ratio error.

corIf Assumptions (ref)-(ref) hold and $\mu(X)=X'\beta_{0}$ a.s., then $\beta^{*}=\beta_{0}$ and $\gamma^{*}$ is solely determined by the $k$ orthogonality conditions \begin{equation} E\left[Xs_{1}(X'\gamma)\left\{ \frac{\sigma(X)^{2}}{s(X'\gamma)^{2}}-1\right\} \right]=0. \end{equation} In particular, for the linear specification $s(t)=t$, the conditional variance approximating function $x\mapsto(x'\gamma^{*}){}^{2}$ satisfies the optimality property $E[\{\sigma(X)^{2}/(X'\gamma^{*}){}^{2}\}-1]=0.$

When the CMF is correctly specified an optimal characterization of $\beta_{0}$ that will lead to an efficient estimator can be formulated by GLS. Define \[ f_{\beta}^{\dagger}\left(Y,X\right):=\frac{1}{\sigma(X)}\phi\left(\frac{Y-X'\beta}{\sigma(X)}\right). \] In the population, GLS maximizes the expected log-likelihood $E[\log f_{\beta}^{\dagger}\left(Y,X\right)]$ with respect to $\beta$, with solution $\beta_{0}$. Then we further have

equation[equation omitted — 145 chars of source]

and inequalities ((ref)) and ((ref)) together imply that, compared to OLS, the MVR solution $\theta^{*}$ formally corresponds to a probability distribution that is KLIC closer to the reference probability distribution $f_{\beta_{0}}^{\dagger}\left(Y,X\right)$ associated to the GLS model.

thmSuppose that $E[|\log f_{Y\mid X}(Y\mid X)|]<\infty$, $E\left[e(Y,X,\theta^{*})^{2}\right]\leq1+\epsilon$, and the scale function $t\mapsto s(t)$ is specified as $s(t)=t$ or $s(t)=\exp(t)$. If Assumptions (ref)-(ref) hold and $\mu(X)=X'\beta_{0}$ a.s., then $f_{\theta^{*}}$ also satisfies \begin{equation} E\left[\log\left(\frac{f_{\beta_{0}}^{\dagger}(Y,X)}{f_{\theta^{*}}(Y,X)}\right)\right]\leq E\left[\log\left(\frac{f_{\beta_{0}}^{\dagger}(Y,X)}{f_{\theta_{LS}}(Y,X)}\right)\right], \end{equation} with equality if and only if $\theta^{*}=\theta_{\textrm{LS}}$.

When the mean is correctly specified, all of the likelihood improvement comes from selecting a better approximation for the standard deviation function than OLS. Relative to the efficient GLS model for the mean, inequality ((ref)) formally establishes that the OLS location model is rejected against the MVR location-scale model according to a likelihood ratio criterion (e.g., V:1989; SW:2017). If the true conditional variance is not constant then the improvement in ((ref)) is strict and the MVR model is closer to the efficient GLS model than the OLS location model.

In the presence of heteroskedasticity, MVR optimality and approximation properties ((ref)) and ((ref)) under correct mean specification provide a theoretical justification for the largely improved MVR-based inference relative to OLS-based inference in the numerical simulations of Section (ref) and the Supplementary Material. In view of its interpretation and since it always admits a well-defined minimizer, the MVR criterion thus offers a natural generalization of OLS for the estimation of linear models.

Connection with Gaussian Maximum Likelihood

MVR provides one criterion for the simultaneous approximation of conditional mean and variance functions. A related criterion is the KLIC of the scaled Gaussian density $f_{\theta}\left(Y,X\right)$ defined in ((ref)) from the true conditional density function $f_{Y\mid X}(Y\mid X)$, which is minimized at a ML pseudo-true value (White:1982). Define for $\theta\in\Theta$,

equation[equation omitted — 212 chars of source]

with first-order conditions

equation[equation omitted — 201 chars of source]

In general the MVR solution $\theta^{*}$ need not satisfy equations ((ref)), and therefore cannot be interpreted as a ML pseudo-true value. Compared with the MVR criterion, an important limitation of criterion ((ref)) is its lack of convexity. The second-order derivative of $\mathcal{L}\left(\theta\right)$ with respect to the first component $\gamma_{1}$ of $\gamma$, i.e., for fixed $\beta$, $\gamma_{-1}$, is \[ \frac{\partial^{\textrm{2}}\mathcal{L}\left(\theta\right)}{\partial\gamma_{1}^{2}}=E\left[\frac{1}{s(X'\gamma)^{2}}\left\{ 3e(Y,X,\theta)^{2}-1\right\} \right], \] which is strictly negative for all $\theta\in\Theta$ such that $e\left(Y,X,\theta\right)^{2}\leq1/3$ a.s. The non convexity of ((ref)) in $\gamma_{1}$ implies that $\mathcal{L}\left(\theta\right)$ is not jointly convex\footnote{Owen:2007 also noted the lack of joint convexity of the negative Gaussian log-likelihood when the scale function is specified to a constant, i.e., for the case $s(X'\gamma)=\sigma\in(0,\infty)$ in ((ref)).}, and that a ML pseudo-true value might not exist; even if there exists one, it need not be unique. In the latter case, some solutions may only be local minima of ((ref)), not endowed with a KLIC-closest interpretation and thus no longer guaranteed to improve over OLS and MVR in a meaningful way.

These observations together with Theorems (ref) and (ref) clarify the relationship between ML, OLS and MVR approximating properties. The ML pseudo-true value is the parameter value associated with the distribution which is KLIC closest to the true data generating process, but is not well-defined due to the objective's lack of convexity. The OLS pseudo-true value is well-defined, but it maximizes a restricted version of the Gaussian expected log-likelihood resulting in a relatively lower likelihood. The MVR loss function strikes a compromise by providing a well-defined convex alternative to Gaussian ML, and relative to OLS by selecting a pseudo-true value that corresponds to a distribution which is KLIC closer to the true data generating process, and KLIC closer to the efficient GLS model under correct mean specification.

Estimation and Inference

We use the sample analog of the MVR population problem ((ref)) for estimation of its solution $\theta^{*}$ in finite samples. We establish existence, uniqueness and consistency of the MVR estimator. We also derive its asymptotic distribution allowing for misspecification of the shapes of the conditional mean and variance functions, and discuss the robustness properties of its influence function. Finally, we provide corresponding tools for robust inference and introduce a one-step MVR-based test for heteroskedasticity.

We assume that we observe a sample of $n$ independent and identically distributed realizations $\{(y_{i},x_{i})\}_{i=1}^{n}$ of the random vector $(Y,X)$. We denote the $n\times k$ matrix of explanatory variables values by $X_{n}$. We define $\Theta_{n}=\mathbb{R}^{k}\times\Theta_{\gamma,n}$, with $\Theta_{\gamma,n}=\{\gamma\in\mathbb{R}^{k}:s(x_{i}'\gamma)>0,i=1,\ldots,n\}$, the sample analog of the parameter space $\Theta$. For $\gamma\in\Theta_{\gamma,n}$, we let $\Omega_{n}(\gamma)=\textrm{diag}(s(x_{i}'\gamma))$, an $n\times n$ diagonal matrix with diagonal elements $s(x_{1}'\gamma),\dots,s(x_{n}'\gamma)$. We also define the MVR moment functions \[ m_{1}(y_{i},x_{i},\theta):=x_{i}e(y_{i},x_{i},\theta),\quad m_{2}(y_{i},x_{i},\theta):=\frac{1}{2}x_{i}s_{1}(x_{i}'\gamma)\{e(y_{i},x_{i},\theta)^{2}-1\}, \] and the corresponding vector $m(y_{i},x_{i},\theta):=(m_{1}(y_{i},x_{i},\theta),m_{2}(y_{i},x_{i},\theta))'$.

The MVR Estimator

The solution to the finite-sample analog of problem ((ref)) is the MVR estimator

equation[equation omitted — 213 chars of source]

For $a=0$ in Assumption (ref), the sample objective in ((ref)) is minimized subject to the $n$ inequality constraints $s(x_{i}'\gamma)>0$, $i=1,\ldots,n$. For $a=-\infty$, the parameter space simplifies to $\Theta_{n}=\mathbb{R}^{2\times k}$ and problem ((ref)) is unconstrained. In terms of implementation, this constitutes an attractive feature of the exponential scale specification.

We derive the asymptotic properties of $\hat{\theta}$ under the following assumptions stated for a scale function in the class defined by Assumption (ref).

assumption(i) $\{(y_{i},x_{i})\}_{i=1}^{n}$ are identically and independently distributed, and (ii) for all $\gamma\in\Theta_{\gamma,n}$, the matrix $X_{n}'\Omega_{n}^{-1}(\gamma)X_{n}$ is finite and positive definite.
assumption$E[Y^{6}]<\infty$, $E[||X||^{6}]<\infty$, and for all $\gamma\in\Theta_{\gamma}$, $E[\left\Vert X\right\Vert ^{6}s_{2}(X'\gamma)^{6}]<\infty$.

Assumption (ref)(i) can be replaced with the condition that $\{(y_{i},x_{i})\}_{i=1}^{n}$ is stationary and ergodic Newey:McFadden:1994. Assumption (ref) is needed for asymptotic normality of estimates of $\theta^{*}$. When the scale function $t\mapsto s(t)$ is specified to be linear, this assumption simplifies to the requirement that $E[Y^{6}]$ and $E[\left\Vert X\right\Vert ^{6}]$ be finite.

Letting $e=e(Y,X,\theta^{*})$, the variance-covariance matrix of the MVR estimator $\hat{\theta}$ is $G^{-1}SG^{-1}/n$, where \[ G:=\left[

array[array omitted — 50 chars of source]

\right]:=E\left[

array[array omitted — 258 chars of source]

\right] \] and \[ S:=\left[

array[array omitted — 50 chars of source]

\right]:=E\left[

array[array omitted — 170 chars of source]

\right]. \] The exact form of each component of matrices $G$ and $S$ depends on the specification of the conditional mean and variance functions, and simplifications of the variance-covariance matrix occur according to the type of misspecification. Under mean misspecification, the form of the variance-covariance matrix of the MVR estimator is not affected by the specification of the conditional variance function.

sloppyDefine estimates of $G$ and $S$ by $\hat{G}:=n^{-1}\sum_{i=1}^{n}\partial m(y_{i},x_{i},\hat{\theta})/\partial\theta$ and $\hat{S}:=n^{-1}\sum_{i=1}^{n}m(y_{i},x_{i},\hat{\theta})m(y_{i},x_{i},\hat{\theta})'$, respectively. The next theorem states the asymptotic properties of the MVR estimator.
thmIf Assumptions (ref)-(ref) hold, then (i) there exists $\hat{\theta}$ in $\Theta$ with probability approaching one; (ii) $\hat{\theta}\rightarrow^{p}\theta^{*}$; and (iii) \begin{equation} n^{1/2}(\hat{\theta}-\theta^{*})\rightarrow_{d}\mathcal{N}(0,G^{-1}SG^{-1}). \end{equation} If $\mu(X)=X'\beta^{*}$ a.s., then the following simplifications occur \begin{equation} G_{12}=G_{21}=0_{k\times k},\quad S_{12}=S_{21}=\frac{1}{2}E[XX's_{1}(X'\gamma^{*})e^{3}]. \end{equation} If $\mu(X)=X'\beta^{*}$ a.s. and $\sigma(X)^{2}=s(X'\gamma^{*})^{2}$ a.s., then the following additional simplifications occur \begin{equation} G_{22}=E\left[\frac{XX'}{s(X'\gamma^{*})}s_{1}(X'\gamma^{*})\right],\quad S_{11}=E[XX'],\quad S_{22}=\frac{1}{4}E[XX's_{1}(X'\gamma^{*})^{2}(e^{4}-1)]. \end{equation} Moreover, $\hat{G}^{-1}\hat{S}\hat{G}^{-1}\rightarrow^{p}G^{-1}SG^{-1}$.

Theorem (ref) allows the construction of confidence intervals and the implementation of hypothesis tests for $\theta$ under each type of model specification using standard errors constructed from the corresponding variance-covariance matrix. The various forms of the variance-covariance matrix in Theorem (ref) provide a basis for the construction of a range of specification tests, similarly to the information matrix equality test in ML theory (White:1982; CS:1991). Inference using the general asymptotic variance formula in ((ref)) will automatically be robust to all forms of misspecification, and therefore to the presence of heteroskedasticity of unknown form.

For the linear homoskedastic model $Y=X'\beta_{0}+U$ with $E[U\mid X]=0$ and $\textrm{Var}(U\mid X)=\sigma_{0}^{2}$, the MVR and OLS variance-covariance matrices coincide asymptotically, and MVR is efficient. Our numerical simulations in Section (ref) and the Supplementary Material illustrate that there is close to no finite-sample loss in estimating linear homoskedastic models using MVR instead of OLS. If both the conditional mean and variance functions are correctly specified, then GLS with weights $1/\sigma^{2}(x)$ is an efficient estimator for $\beta_{0}$. Letting $\check{Y}:=Y/s(X'\gamma^{*})$, $\check{X}:=X/s(X'\gamma^{*})$ and $\check{s}(X'\gamma):=s(X'\gamma)/s(X'\gamma^{*})$, define the weighted MVR objective \[ Q^{\textrm{WMVR}}(\theta):=E\left[\frac{1}{2}\left\{ \left(\frac{\check{Y}-\check{X}'\beta}{\check{s}(X'\gamma)}\right)^{2}+1\right\} \check{s}(X'\gamma)\right]=E\left[\frac{1}{2}\left\{ e\left(Y,X,\theta\right)^{2}+1\right\} \frac{s(X'\gamma)}{s(X'\gamma^{*})}\right]. \] If $\sigma^{2}(X)=s(X'\gamma_{0})^{2}$, then $\gamma^{*}=\gamma_{0}$ and $Q^{\textrm{WMVR}}(\theta)$ has first-order conditions for $\beta$ \[ \frac{\partial Q^{\textrm{WMVR}}(\theta)}{\partial\beta}=-E\left[\frac{X}{s(X'\gamma_{0})}e\left(Y,X,\theta\right)\right]=0, \] which are satisfied by $\theta=\theta_{0}$ and coincide with the GLS (and ML) first-order conditions for $\beta$ at a solution. In general, the functional form of the conditional variance function is unknown, and the MVR and weighted MVR solutions will differ.

An important implication of Theorem (ref) is that the influence function of the MVR estimator for $\beta$ is proportional to both moment functions $m_{1}$ and $m_{2}$: \[ IF_{\beta}(y,x,\theta)=-(G_{11}-G_{12}G_{22}^{-1}G_{21})^{-1}[m_{1}(y,x,\theta)-G_{12}G_{22}^{-1}m_{2}(y,x,\theta)]. \] The quadratic term $m_{2}$ dominates and an influential observation is defined as having $(y_{i}-x_{i}'\beta)^{2}$ large enough for $e(y_{i},x_{i},\theta)^{2}$ to be large. Observations that are influential for $\beta$ are influential relative to the dispersion of $Y$, accounting for mean misspecification.

When the CMF is well-specified, the variance-covariance matrix takes the form \[ G^{-1}SG^{-1}=\left[

array[array omitted — 138 chars of source]

\right]. \] The influence function of $\beta$ is then proportional to $m_{1}$ only, and the influence function of $\gamma$ is proportional to $m_{2}$ only, since the off-diagonal blocks of $G$ are then $0_{k\times k}$. For the mean parameter $\beta$, an observation $(y_{i},x_{i})$ with large influence will be such that $y_{i}$ is large enough for the standardized residual $e(y_{i},x_{i},\theta)$ to be large. Because $\hat{\beta}$ and $\hat{\gamma}$ are determined simultaneously, the influence of outliers on the mean parameter is limited by the restriction that the sample second moment of $e(y_{i},x_{i},\theta)$ must remain close to one, and be exactly one if the scale is linear. In sharp contrast with OLS, the scale parameter will simultaneously compensate an increase in $Y$ dispersion so as to keep the variance of $e(y_{i},x_{i},\theta)$ constant. The MVR influence function although unbounded for a fixed value of $\gamma$ thus robustifies OLS through the simultaneous reweighting of the residuals, downweighting regions in the covariate space where the information on $Y$ is imprecise, as measured by $s(x'\gamma)$, in the calculation of the regression fit.

In summary, the MVR estimator does not robustify OLS through the bounding of the influence function (Koenker:2005), but by incorporating information about the dispersion of $Y$ across the covariate space in the definition of an influential outlier.

rem(Implementation.) Under our assumptions, the MVR objective is globally convex in $\theta$, and therefore in $\beta$ for any $\gamma\in\Theta_{\gamma,n}$. This implies that for any $\gamma\in\Theta_{\gamma,n}$ there exists a unique corresponding minimizer $\hat{\beta}(\gamma)$. This observation forms the basis of our implementation, and letting $y=(y_{1},\ldots,y_{n})'$, we first obtain $\hat{\gamma}$ by solving \begin{align*} \min_{\gamma\in\mathbb{R}^{k}}\,\frac{1}{n}\sum_{i=1}^{n}\frac{1}{2}\left\{ \left(\frac{y_{i}-x_{i}'\hat{\beta}(\gamma)}{s(x_{i}'\gamma)}\right)^{2}+1\right\} s(x_{i}'\gamma), & \quad\hat{\beta}(\gamma):=[X_{n}'\varOmega_{n}^{-1}(\gamma)X_{n}]^{-1}X_{n}'\varOmega_{n}^{-1}(\gamma)y,\\ s.t.\quad s(x_{i}'\gamma)>0,\quad i=1,\ldots,n, & \quadif s(t)\leq0 for some t\in\mathbb{R}. \end{align*} Concentrating out $\beta$ for each $\gamma$ provides a convenient implementation of the MVR estimator $\hat{\gamma}$, with the final estimate for $\beta$ defined as $\hat{\beta}:=\hat{\beta}(\hat{\gamma})$.\qed

Inference

Given the MVR estimator $\hat{\theta}=(\hat{\beta},\hat{\gamma})'$, inference is performed based on the estimated asymptotic variance-covariance matrix $\hat{V}:=\hat{G}^{-1}\hat{S}\hat{G}^{-1}$, which can be partitioned into 4 blocks \[ \hat{V}=\left[

array[array omitted — 74 chars of source]

\right]. \] The specific form of $\hat{V}$ depends on the specification assumptions made on the conditional mean and variance functions. For $\hat{\beta}_{j}$ and $\hat{\gamma}_{j}$ the $j$th components of $\hat{\beta}$ and $\hat{\gamma}$, respectively, MVR standard errors are obtained as \[ \textrm{s.e.}(\hat{\beta}_{j}):=\left(\frac{1}{n}\left[\hat{V}_{11}\right]_{j,j}\right)^{\frac{1}{2}},\quad\textrm{s.e.}(\hat{\gamma}_{j}):=\left(\frac{1}{n}\left[\hat{V}_{22}\right]_{j,j}\right)^{\frac{1}{2}}, \] with resulting two-sided confidence intervals with nominal level $1-\alpha$, \[ \hat{\beta}_{j}\pm\Phi^{-1}(1-\alpha/2)\times\textrm{s.e.}(\hat{\beta}_{j}),\quad\hat{\gamma}_{j}\pm\Phi^{-1}(1-\alpha/2)\times\textrm{s.e.}(\hat{\gamma}_{j}), \] where $\Phi^{-1}(1-\alpha/2)$ denotes the $1-\alpha/2$ quantile of the Gaussian distribution. A significance test of the null $\beta_{j}=0$ and $\gamma_{j}=0$ can then be performed using the test statistics $\hat{\beta}_{j}/\textrm{s.e.}(\hat{\beta}_{j})$ and $\hat{\gamma}_{j}/\textrm{s.e.}(\hat{\gamma}_{j})$.

Simultaneous significance testing or hypothesis tests on linear combination of multiple parameters can be implemented by a Wald test. For $h\leq2\times k$, letting $R$ be an $h\times(2\times k)$ matrix of constants of full rank $h$ and $r$ be an $h\times1$ vector of constants, define \[ H_{0}:R\theta^{*}-r=0,\quad H_{1}:R\theta^{*}-r\neq0, \] the null and alternative hypotheses for a two-sided tests of linear restrictions on the location-scale model $Y=X'\beta^{*}+s(X'\gamma^{*})e$. It follows from asymptotic normality of $\hat{\theta}$ in ((ref)) that the corresponding MVR Wald statistic $W_{\textrm{MVR}}$ satisfies \[ W_{\textrm{MVR}}:=(R\hat{\theta}-r)'[R(\hat{V}/n)R']^{-1}(R\hat{\theta}-r)\sim\chi_{(h)}^{2}, \] under the null $H_{0}$.

The Wald statistic $W_{\textrm{MVR}}$ can be specialized to formulate a one-step robust MVR-based test for heteroskedasticity. Letting \[ h=k-1,\quad R=\left[

array[array omitted — 37 chars of source]

\right],\quad r=0_{k-1}, \] the statistic $W_{\textrm{MVR}}$ gives a robust test of the null hypothesis $H_{0}:\gamma_{2}^{*}=\ldots=\gamma_{k}^{*}=0$.

remWhen the CMF is linear, robust MVR inference on $\hat{\beta}$ uses the closed-form variance formula \[ \widehat{\textrm{Var}}(\hat{\beta})=n^{-1}(X_{n}'\Omega_{n}^{-1}(\hat{\gamma})X_{n})^{-1}(X_{n}'\hat{\varPsi}_{e}X_{n})(X_{n}'\Omega_{n}^{-1}(\hat{\gamma})X_{n})^{-1}, \] where $\hat{\varPsi}_{e}=\textrm{diag}(\hat{e}_{i}^{2})$.\qed

Numerical Illustrations

All computational procedures can be implemented in the software R (R:2017) using open source software packages for nonlinear optimization such as Nlopt, and its R interface Nloptr (YBE:2018).

Empirical Application: Reversal of Fortune

We apply our methods to the study of the effect of European colonialism on today's relative wealth of former colonies, as in AJR:2002. They show that former colonies that were relatively rich in 1500 are now relatively poor, and provide ample empirical evidence of this reversal of fortune. In particular, they study the relationship between urbanization in 1500 and GDP per capita in 1995 (PPP basis), using OLS regression analysis. The sample size ranges from 17 to 41 former colonies, allowing the illustration of MVR properties in small samples.

We take the outcome $Y$ to be log GDP per capita in 1995 and in the baseline specification $X$ includes an intercept and a measure of urbanization in 1500, a proxy for economic development. We implement MVR with both linear ($\ell$-MVR) and exponential ($e$-MVR) scale functions, and report estimated standard errors robust to mean misspecification according to ((ref)). We also report OLS estimates, with heteroskedasticity-robust standard errors. In the Supplementary Material we also provide results including standard errors with finite-sample adjustments suggested by MKW:1985, and we also report MVR standard errors calculated under correct mean misspecification. Our findings are robust to these modifications.

sidewaystable\begin{tabular}{lcccccccccccc} \toprule & & \multicolumn{11}{c}{Dependent variable is log GDP per capita (PPP) in 1995}\tabularnewline & & & & & & & & & & & & \tabularnewline \cmidrule{3-13} & & OLS & $\ell$-MVR & e-MVR & & OLS & $\ell$-MVR & e-MVR & & OLS & $\ell$-MVR & e-MVR \tabularnewline \cmidrule{3-5} \cmidrule{7-9} \cmidrule{11-13} & & \multicolumn{3}{c}{(1) Base sample} & & \multicolumn{3}{c}{(2) Without North Africa} & & \multicolumn{3}{c}{(3) Without the Americas}\tabularnewline & & \multicolumn{3}{c}{($n=41$)} & & \multicolumn{3}{c}{($n=37$)} & & \multicolumn{3}{c}{($n=17$)}\tabularnewline Urbanization in 1500 & & -0.078 & -0.067 & -0.069 & & -0.101 & -0.099 & -0.099 & & -0.115 & -0.064 & -0.077 \tabularnewline & & (0.023) & (0.028) & (0.026) & & (0.032) & (0.034) & (0.034) & & (0.043) & (0.127) & (0.113)\tabularnewline & & & & & & & & & & & & \tabularnewline \cmidrule{3-5} \cmidrule{7-9} \cmidrule{11-13} & & \multicolumn{3}{c}{(4) Just the Americas} & & \multicolumn{3}{c}{(5) With the continent} & & \multicolumn{3}{c}{(6) Without neo-Europes}\tabularnewline & & \multicolumn{3}{c}{($n=24$)} & & \multicolumn{3}{c}{dummies ($n=41$)} & & \multicolumn{3}{c}{($n=37$)}\tabularnewline Urbanization in 1500 & & -0.053 & -0.045 & -0.044 & & -0.082 & -0.063 & -0.060 & & -0.046 & -0.036 & -0.038 \tabularnewline & & (0.029) & (0.032) & (0.032) & & (0.031) & (0.029) & (0.030) & & (0.021) & (0.023) & (0.023)\tabularnewline & & & & & & & & & & & & \tabularnewline \cmidrule{3-5} \cmidrule{7-9} \cmidrule{11-13} & & \multicolumn{3}{c}{(7) Controlling for Latitude} & & \multicolumn{3}{c}{(8) Controlling for colonial} & & \multicolumn{3}{c}{(9) Controlling for religion}\tabularnewline & & \multicolumn{3}{c}{($n=41$)} & & \multicolumn{3}{c}{origin ($n=41$)} & & \multicolumn{3}{c}{($n=41$)}\tabularnewline Urbanization in 1500 & & -0.072 & -0.069 & -0.070 & & -0.071 & -0.063 & -0.062 & & -0.060 & -0.042 & -0.040 \tabularnewline & & (0.020) & (0.022) & (0.021) & & (0.025) & (0.026) & (0.027) & & (0.027) & (0.029) & (0.029)\tabularnewline & & & & & & & & & & & & \tabularnewline \bottomrule \end{tabular} \caption{Reversal of fortune. Asymptotic heteroskedasticity-robust OLS standard errors and MVR standard errors are in parenthesis.}

Table (ref) reports our results for urbanization in the baseline specification across 5 different sets of countries, and for 4 additional specifications\footnote{We exclude two specifications of Table III in AJR:2002 for which not all types of OLS and MVR standard errors are well-defined.} including continent dummies, and controlling for latitude, colonial origin and religion\footnote{See AJR:2002 for a detailed description of the data.}. A striking feature of the results is the robustness to scale specification of MVR point estimates and standard errors. They are nearly identical across all specifications, except for Panel (3). Compared to OLS, MVR point estimates are all smaller in magnitude, suggesting a negative bias of OLS away from zero while standard errors are of similar magnitude, making it more likely to find a significant relationship with OLS estimates in this empirical application. The urbanization coefficient loses significance with MVR in 4 specifications, mainly as a result of the change in point estimates.

Specifically, we find that MVR provides supporting evidence of a significant statistical relationship between urbanization in 1500 and GDP per capita in 1995 in the whole sample, but also dropping North Africa, including continent dummies, and controlling for latitude and for colonial origin. However, the relationship between urbanization in 1500 and GDP per capita in 1995 is not statistically significant in the four remaining specifications. When the Americas are dropped (Panel (3)), when only former colonies from the Americas are considered (Panel (4)), and when controlling for religion (Panel (9)), the urbanization coefficient is no longer significant with MVR. These results are robust to implementing finite-sample adjustments. Specification (6) drops observations for neo-Europes (United States, Canada, New Zealand, and Australia), and only the $e$-MVR estimate is significant at the 10 percent level when no finite-sample adjustments are implemented, and loses significance otherwise.

MVR results provide renewed empirical support for a subset of the specifications, but overall show that the mean relationship in this empirical application is weaker and less robust than first suggested by the OLS-based analysis\footnote{We also implemented WLS and found that the magnitude of most WLS coefficients is smaller than MVR point estimates. In addition to specifications (3), (4), (6) and (9), specification (5) is also found to be not statistically significant. We report the results in Section 3 of the Supplementary Material.}. These findings illustrate that MVR can substantially alter the conclusions obtained using OLS in practice.

Numerical Simulations

We investigate the properties of MVR in small samples and compare its performance with OLS and WLS by implementing numerical simulations based on the experimental setup in MacKinnon:2013. In the Supplementary Material, we provide additional results for models featuring a nonlinear CMF and report simulation results from an experiment calibrated to a second empirical example. We find that using MVR approximations does not result in a loss in the quality of approximation of nonlinear CMFs compared to OLS, and MVR estimation and inference finite-sample properties compare favorably to both OLS and WLS.

Design of Experiments

The data generating process is

align*[align* omitted — 294 chars of source]

where all regressors are drawn from the standard log-normal distribution, and $z(\alpha)$ is chosen such that the expected variance of $\sigma\varepsilon$ is equal to 1. The log-normal regressors ensure that many samples will include high-leverage points with a few observations taking extreme values. This feature of the design distorts the distribution of test statistics based on heteroskedasticity-robust estimators of OLS standard errors. The parameter coefficient values are set to $\beta_{j}=\gamma_{j}=1$ for $j=0,\ldots,4$.\footnote{This departs slightly from the original MacKinnon:2013 design where $\beta_{4}=\gamma_{4}=0$. We are grateful to James MacKinnon for suggesting this modification that preserves heteroskedasticity in $X_{4}$.} The index $\alpha$ measures the degree of heteroskedasticity in the model, with $\alpha=0$ corresponding to homoskedasticity, and $\alpha=2$ corresponding to high heteroskedasticity. The numerical simulations are implemented for sample sizes $n=20,40,80,160,320,640$ and $1280$.

For each $\alpha$ and sample size, we generate 10000 samples, and implement OLS, WLS and MVR. We implement MVR with both linear ($\ell$-MVR) and exponential ($e$-MVR) scale functions. For WLS we follow the implementations proposed by Romano and Wolf (RW:2017, cf. equation (3.4) and (3.5), p. 4). Denote the OLS estimator by $\hat{\beta}_{\textrm{LS}}$ and let $\tilde{x}_{i}=(x_{1i},x_{2i},x_{3i},x_{4i})'$. We form the OLS residuals $\hat{u}_{i}:=y_{i}-x_{i}'\hat{\beta}_{\textrm{LS}}$, $i=1,\ldots,n$, and perform the OLS regressions \[ \log(\max(\delta^{2},\hat{u}_{i}^{2}))=\nu+\pi'|\tilde{x}_{i}|+\eta_{i}, \] for WLS with linear scale ($\ell$-WLS)\footnote{This regression is performed imposing the $n$ constraints $\nu+\pi|\tilde{x}_{i}|\geq\delta$, $i=1,\ldots,n$, using the $\mathtt{lsei}$ R package (WLH:2017). }, and \[ \log(\max(\delta^{2},\hat{u}_{i}^{2}))=\nu+\pi'\log(|\tilde{x}_{i}|)+\eta_{i}, \] for WLS with exponential scale ($e$-WLS), where $\delta=0.1$ as in the implementation of RW:2017, and with estimates $(\hat{\nu},\hat{\pi})$. The WLS weights are formed as $\hat{w}_{i}^{\ell}:=\hat{\nu}+\hat{\pi}'|\tilde{x}_{i}|$ and $\hat{w}_{i}^{e}:=\exp(\hat{\nu}+\hat{\pi}'\log(|\tilde{x}_{i}|))$, and the WLS estimators are \[ \hat{\beta}_{\textrm{WLS}}^{m}:=[X'_{n}(W_{n}^{m})^{-1}X{}_{n}]^{-1}X'_{n}(W_{n}^{m})^{-1}y,\quad W_{n}^{m}:=\textrm{diag}(\hat{w}_{i}^{m}),\quad m=\ell,e, \] where $y=(y_{1},\ldots,y_{n})'$, $X_{n}$ is the $n\times5$ matrix of explanatory variables values, and $\textrm{diag}(\hat{w}_{i}^{\ell})$ and $\textrm{diag}(\hat{w}_{i}^{e})$ denote the $n\times n$ diagonal matrices with diagonal elements $\hat{w}_{1}^{\ell},\dots,\hat{w}_{n}^{\ell}$ and $\hat{w}_{1}^{e},\dots,\hat{w}_{n}^{e}$, respectively.

In all experiments the results for $\beta_{1}$ to $\beta_{4}$ are similar and we thus only report the results for $\beta_{4}$ for brevity. Also, throughout the relative performance of MVR and WLS is assessed by comparing $\ell$-MVR to $\ell$-WLS, and $e$-MVR to $e$-WLS.

Estimation results

Tables (ref) and (ref) report the ratio of MVR root mean squared errors (RMSE) across simulations over the OLS and WLS RMSEs for the coefficient parameter $\beta_{4}$, each sample size and value of heteroskedasticity parameter $\alpha$, in percentage terms. Denoting an estimator $\tilde{\beta}_{4}^{(s)}$ of $\beta_{4}$ for the $s$th simulation, the RMSE is computed as $\{\frac{1}{S}\sum_{s=1}^{S}(\tilde{\beta}_{4}^{(s)}-\beta_{4})^{2}\}^{1/2}$, for $S=10000$.

Table (ref) shows that the performance of both MVR estimators relative to OLS improves as $n$ and $\alpha$ increase. As expected, for the homoskedastic case $\alpha=0$ the performance of MVR and OLS estimators is very similar, and the ratios converge to 100 from above, reflecting the efficiency of the OLS estimator in that case. The performance of MVR then becomes markedly superior as $n$ and $\alpha$ increase, with ratios that reach $31.7$ for $\ell$-MVR and $22.9$ for e-MVR. The estimator $\ell$-MVR dominates e-MVR slightly for $n=20$. The performance of the estimator e-MVR then becomes superior as the degree of heteroskedasticity and sample size increase, showing higher robustness of the exponential scale specification in more extreme designs in these simulations.

In Table (ref), we find that the relative performance of both MVR estimators relative to WLS also improves as $n$ increases and as $\alpha$ increases from 0.5 to 2. For the homoskedastic case $\alpha=0$, an interesting feature of the simulation results is that the relative performance of MVR and WLS estimators now converges to 100 from below. This reflects the fact that for homoskedastic designs MVR weights are better able to mitigate the cost of reweighting in small samples compared to WLS weights. For other designs with $\alpha>0$, the relative performance of both MVR estimators dominates the performance of WLS with ratios that reach $76.2$ for $\ell$-MVR and $43.5$ for e-MVR. For $\alpha=1$, WLS with linear scale is efficient and dominates slightly $\ell$-MVR as $n$ increases. Compared to OLS and the results of Table (ref), these results show that in this experiment WLS also improves over OLS, that $\ell$-MVR improves over WLS as $n$ increases and $\alpha$ deviates from $1$, with little loss for $\alpha\leq1$, and $e$-MVR yields substantial additional gains over WLS as both $n$ and $\alpha$ increase.

table[table omitted — 2,576 chars of source]

Inference

sloppyIn order to study the finite-sample performance of MVR inference relative to heteroskedasticity-robust OLS and WLS inference, we first compare the rejection probabilities of asymptotic $t$ tests of the null hypothesis $\beta_{4}=1$ based on the standard normal distribution\footnote{We also performed simulations with and tested for $\beta_{4}=0$, and calculated rejection probabilities using a $t_{n-k}$ distribution. The relative performance of the methods remains similar.}. We then compare the lengths of the confidence intervals constructed for the coefficient parameter $\beta_{4}$. MVR standard errors are calculated under mean misspecification according to ((ref)). OLS and WLS standard errors used in the construction of confidence intervals and tests statistics are the asymptotic heteroskedasticity-robust standard errors. For completeness, in the Supplementary Material we also compare rejection probabilities and confidence intervals based on standard errors with finite-sample adjustments suggested by MKW:1985, and we also replicate all experiments using MVR standard errors calculated under correct mean misspecification with the simplifications in ((ref)).

Figure (ref) displays rejection probability curves of asymptotic $t$ tests of the null hypothesis $\beta_{4}=1$ for each sample size and value of the heteroskedasticity parameter $\alpha$. The nominal size of the tests is set to $5\%$. Figures (ref)(a)-(d) show that MVR addresses overrejection of the OLS- and WLS-based tests in the presence of heteroskedasticity ($\alpha>0$). A striking feature of MVR rejection probability curves is that they flatten very quickly across $\alpha$ as $n$ increases, a feature somewhat more pronounced for $\ell$-MVR curves. This is in sharp contrast with OLS and $e$-WLS rejection probability curves that are increasing with the degree of heteroskedasticity $\alpha$. For $\ell$-WLS the rejection curves are distorted around $\alpha=1$, for which it is efficient, and overall the rejection probabilities are much larger than for $\ell$-MVR. The curves for $\ell$-MVR and $\ell$-WLS coincide only for the case where $\ell$-MVR is efficient ($\alpha=1$) at moderate sample size and above ($n\geq320$). The $\ell$-MVR rejection probability curve for $n=20$ (black curve) is not placed above the other curves although it is above the nominal level for all values of $\alpha$. This feature disappears when finite-sample corrections are implemented (Figures 2.1-2.3 in the Supplementary Material).

figure[figure omitted — 803 chars of source]

In order to further investigate the relative performance of MVR-based inference, Tables (ref) and (ref) report the ratio of average MVR confidence interval lengths across simulations over the average OLS and WLS confidence interval lengths for $\beta_{4}$ for each sample size and value of heteroskedasticity index $\alpha$, in percentage terms. We find that in the presence of heteroskedasticity, the length of MVR confidence intervals is shorter for all designs compared to both OLS and WLS confidence intervals for some $n$ large enough. The only exception is relative to $\ell$-WLS for $\alpha=0.5,1$, as expected for $\alpha=1$ from $\ell$-WLS being efficient in that case. The relatively larger average length of the confidence intervals for $\ell$-MVR when $n=20$ in Tables (ref) and (ref) is very much reduced with finite-sample corrections (Tables 1-4 in the Supplementary Material).

table[table omitted — 2,735 chars of source]

Overall these simulation results demonstrate that MVR achieves large improvements in terms of estimation and inference compared to OLS in the presence of heteroskedasticity, and compared to WLS when the conditional variance function is misspecified. Our numerical simulations confirm MVR robustness to the specification of the scale function, and both $\ell$-MVR and $e$-MVR perform very well in finite samples. In the presence of heteroskedasticity rejection probabilities for MVR are much closer to nominal level than those for OLS and for WLS with misspecified weights. MVR achieves these improvements while simultaneously displaying tighter confidence intervals in all designs for sample sizes large enough. They are also shorter than their WLS counterpart with misspecified weights for sample sizes large enough. The precision of MVR estimates measured in RMSE is also largely superior to OLS under heteroskedasticity and to WLS with misspecified weights, with lower losses than WLS relative to OLS under homoskedasticity. These results and the simulations in the Supplementary Material illustrate the higher precision, improved finite-sample inference, and favorable approximation properties of MVR compared to classical least-squares methods.

Conclusion

We introduce a new loss function for the linear estimation and approximation of CMFs. The proposed alternative generalises OLS, resulting in more robust approximations under misspecification and large improvements in finite samples. Given the importance of the least-squares loss in econometrics and statistics, and the common occurence of heteroskedasticity in empirical practice, the range of applications for simultaneous mean-variance regression will be vast. Examples of natural avenues for future research include the method of instrumental variables, GARCH models, and flexible specification of the conditional variance function for efficient estimation. These extensions will be considered in subsequent work.