EconBase
← Back to paper

Centered and non-centered variance inflation factor

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

19,847 characters · 8 sections · 17 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Centered and non-centered variance inflation factor

abstractThis paper analyzes the diagnostic of near multicollinearity in a multiple linear regression from auxiliary centered regressions (with intercept) and non-centered (without intercept). From these auxiliary regression, the centered and non-centered Variance Inflation Factors are calculated, respectively. It is also presented an expression that relate both of them.

{\it Keywords:} Centered model, non centered model, intercept, essential multicollinearity, non-essential multicollinearity.

Introduction

Considering the following multiple linear model with $n$ observations and $k$ regressors :

equation[equation omitted — 153 chars of source]

where $\mathbf{y}$ is a vector with the observations of the dependent variable, $\mathbf{X}$ is a matrix containing the observations of regressors and $\mathbf{u}$ is a vector representing random disturbance (that is supposed to be spherical). When this model presents near multicollinearity, it is to say, when the linear relation between the regressors affects to the numerical and/or statistical analysis of the model, it is usual to transform the data (see, for example, Belsley Belsley1984, Marquardt Marquardt1980 or, more recently, Velilla Velilla2018).

In these cases, the first column of matrix $\mathbf{X}$ is said to be composed by ones to denote that the model contains an intercept. Thus, $\mathbf{X} = [ \mathbf{1} \ \mathbf{X}_{2} \dots \mathbf{X}_{k}]$ where $\mathbf{1}_{n \times 1} = (1 \ 1 \dots 1)^{t}$. This model is considered to be centered. On the other hand, transformed models are considered to be non-centered, since the transformations (centering, typification or standardization) imply the elimination of the intercept in the model. Note that even after transforming the data, it is possible to recover the original model (centered) from the estimations of the transformed model (non-centered model). However, in this paper we refer to centered and non-centered model depending on the inclusion of intercept. Thus, it is considered that the model is centered if $\mathbf{X} = [ \mathbf{1} \ \mathbf{X}_{2} \dots \mathbf{X}_{k}]$ and non-centered if $\mathbf{X} = [ \mathbf{X}_{1} \ \mathbf{X}_{2} \dots \mathbf{X}_{k}]$ being $\mathbf{X}_{j} \not= \mathbf{1}$ with $j=1,\dots,k$.\\

The main contribution of the paper is the analysis of the variance inflation factor, which is widely applied to detect multicollinearity in model ((ref)) by considering that the auxiliary regression used to its calculation is centered or not. For a better understanding of this paper, could be interesting to distinguish between the following two kinds of near multicollinearity that can be found in model ((ref)), (see Marquardt and Snee MarquardtSnee):

description• Near linear relation between the intercept and at least one of the independent variables. • Near linear relation between at least two of the independent variables (excluding the intercept).

The structure of the paper is as follows: Section (ref) presents some preliminaries and introduces the main questions that will be answered through the paper, Section (ref) presents the non-centered auxiliary regressions and, finally, section (ref) summarizes the main conclusions.

Preliminaries

Considering $k=3$ in ((ref)), Belsley Belsley1984 used the non-centered coefficient of determination (without considering the intercept) of the regression of $\mathbf{X}_{3}$ as a function of $\mathbf{X}_{2}$ to calculate the non-centered variance inflation factor (denoted as VIFnc) of the following regression: $$\mathbf{y} = \beta_{1} \cdot \mathbf{1} + \beta_{2} \cdot \mathbf{X}_{2} + \beta_{3} \cdot \mathbf{X}_{3} + \mathbf{u},$$ Thus, Belsley Belsley1984 used the coefficient of determination of the following auxiliary regression: $$\mathbf{X}_{3} = \delta \cdot \mathbf{X}_{2} + \mathbf{v},$$ taking into account the values \footnote{Variables $\mathbf{y}$, $\mathbf{X}_{2}$ and $\mathbf{X}_{3}$ were originally used by Belsley Belsley1984. Variable $\mathbf{X}_{4}$ has been randomly generated to obtain a variable linearly independent to the rest.} displayed in Table (ref). In these data, it is intuited the existence of near non-essential multicollinearity, it is to say, relation between the intercept and at least one of the independent variables of the model.

table[table omitted — 1,403 chars of source]

By using the original variables applied by Belsley, the traditional VIF (from centered model, see Theil Theil1971) provides a value equal to 1 (its minimum possible value), while the VIFnc is equal to 100032.1. However, if an additional variable is included (generated from a normal distribution with mean equal to 4 and variance equal to 16), $\mathbf{X}_{4}$, the following values are obtained for the VIF and the VIFnc of the three variables: $$1.155364, 1.084168, 1.239559, \qquad 100453.8, 100490.6, 1.773768.$$ Thus, the VIF is not detecting the existence of essential near multicollinearity, (see Salmer\'on et al Salmeron2018) while the VIFnc does detect it.

However, since the calculation of VIFnc excludes the constant term, the detected relation refers to the one between $\mathbf{X}_{2}$ and $\mathbf{X}_{3}$, and not to the relation between $\mathbf{X}_{2}$ and/or $\mathbf{X}_{3}$ with the intercept. These results suggest a new definition of non-essential multicollinearity as the relation between at least two variables with little variability. Thus, the particular case when one of these variables is the intercept leads to the definition initially given by Marquardt y Snee MarquardtSnee.

The following values are obtained for the VIF and VIFnc of the second and fourth variables, respectively: $$1.143328, \qquad 1.765676,$$ while, for the third and fourth: $$1.072873, \qquad 1.766323.$$ Thus, it is not detected in any case a relation between $\mathbf{X}_{2}$ or $\mathbf{X}_{3}$ with the intercept.

With these results and by following Salmer\'on et al Salmeron2018, it can be concluded that the VIF only detects the near essential multicollinearity and the VIFnc only detects the non essential near multicollinearity. It will be interesting to analyze if the VIFnc could also detect the essential one. It could be also interesting to determine a threshold for the VIFnc from which the multicollinearity will be considered worrying. \\ On the other hand, given the model ((ref)), the expression obtained for the variance of the estimator is given by:

equation[equation omitted — 125 chars of source]

where $RSS_{j}$ is the residual sum of squares (RSS) of the auxiliary regression of the $-j$ independent variable as a function of the rest of independent variables. Taking into account the decomposition of the squared sums, this expression is equivalent to:

equation[equation omitted — 140 chars of source]

where $SST_{j}$ is the residual sum of squares (TSS) of the previous auxiliary regression. But, this decomposition is only verified if there is intercept in the auxiliary regression, in other case, it is not possible to state that expressions ((ref)) and ((ref)) are equivalent.

In this case, while the participation in the calculation of $var (\widehat{\beta}_{j})$ was not showed, it is not appropriate to consider that the VIFnc is a factor that inflate the variance. Anyway, it is evident that the VIFnc is able to detect near multicollinearity in a linear regression model.

In the following section, we answer the following questions:

enumerate[a)] • Is the VIFnc a factor which inflate the variance? • What kind of multicollinearity is able to detect?

Finally, main results are presented in Section (ref).

Auxiliary non-centered regressions

This section presents the calculation of the VIFnc considering that the auxiliary regression is non-centered, it is to say, it has no intercept. Firstly, it is presented how to calculate the coefficient of determination for non-centered models.

Non-centered coefficient of determination

Given the linear regression ((ref)) with or without intercept, it is verified the following decomposition for the sum of squares:

equation[equation omitted — 167 chars of source]

where $\widehat{\mathbf{y}}$ represents the estimation of the dependent variable of the model fitted by ordinary least squares (OLS) and $\mathbf{e} = \mathbf{y} - \widehat{\mathbf{y}}$ the residuals obtained from that fit. In this case, the coefficient of determination is obtained by the following expression:

equation[equation omitted — 224 chars of source]

Comparing the decomposition of the sums of squares given by ((ref)) with the traditionally applied to calculated the coefficient of determination in models with intercept as in model ((ref)):

equation[equation omitted — 216 chars of source]

it is noted that both coincide if the dependent variable has zero mean. If the mean is different to zero, both models present the same residual sum of squares and different explained and total sum of squares. Thus, these models lead to the same value for the coefficient of determination (and, as consequence, for the VIF) only if the dependent variable presents a mean equal to zero.

Non-centered variance inflator factor

The VIFnc is obtained from expression:

equation[equation omitted — 99 chars of source]

being $R^{2}nc(j)$ the coefficient of determination, calculated by following ((ref)), of the non-centered auxiliary regression:

equation[equation omitted — 106 chars of source]

where $\mathbf{X}_{-j}$ is equal to the matrix $\mathbf{X}$ after eliminating the variable $\mathbf{X}_{j}$, for $j = 2,\dots, k$, and it has not a vector of ones representing the intercept.

In this case:

itemize$\sum \limits_{i=1}^{n} X_{ij}^{2} = \mathbf{X}_{j}^{t} \mathbf{X}_{j}$, and • $\sum \limits_{i=1}^{n} \widehat{X}_{ij}^{2} = \widehat{\mathbf{X}}_{j}^{t} \widehat{\mathbf{X}}_{j} = \mathbf{X}_{j}^{t} \mathbf{X}_{-j} \cdot \left( \mathbf{X}_{-j}^{t} \mathbf{X}_{-j} \right)^{-1} \cdot \mathbf{X}_{-j}^{t} \mathbf{X}_{j}$ due to $\widehat{\mathbf{X}}_{j} = \mathbf{X}_{-j} \cdot \left( \mathbf{X}_{-j}^{t} \mathbf{X}_{-j} \right)^{-1} \cdot \mathbf{X}_{-j}^{t} \mathbf{X}_{j}$.

and then:

eqnarray[eqnarray omitted — 735 chars of source]

Thus, the VIFnc coincides to the expression given by Stewart Stewart1987 for the VIF and denoted as $k_{j}^{2}$, it is to say, $VIFnc(j) = k_{j}^{2}$.

However, recently, Salmer�n et al. Salmeron2019 showed that the index presented by Stewart has been, even by the own Stewart, misleading identified as the VIF, verifying the following relation between both measures:

equation[equation omitted — 141 chars of source]

where $\overline{\mathbf{X}}_{j}$ is the mean of the $-j$ variable of $\mathbf{X}$.

From expression ((ref)) it is shown that the VIFnc and the VIF only coincide if the associated variable has zero mean, analogously to what happens in the decomposition of the sum of squares.

Does the VIFnc inflate the variance?

From expression ((ref)) and considering that expression ((ref)) can be rewritten as: $$VIFnc(j) = \frac{\mathbf{X}_{j}^{t} \mathbf{X}_{j}}{RSS_{j}},$$ it is possible to obtain:

equation[equation omitted — 199 chars of source]

It is required to establish a model as a reference to conclude if the variance has been inflated (see, for example, Cook Cook1984). Thus, if the variables in $\mathbf{X}$ are orthogonal, it is verified that $\mathbf{X}^{t} \mathbf{X} = diag( d_{1},\dots, d_{k})$ where $d_{j} = \mathbf{X}_{j}^{t} \mathbf{X}_{j}$. In this case, $\left( \mathbf{X}^{t} \mathbf{X} \right)^{-1} = diag( 1/d_{1},\dots, 1/d_{k})$ and, consequently:

equation[equation omitted — 167 chars of source]

In this case, $$\frac{var (\widehat{\beta}_{j})}{var (\widehat{\beta}_{j,o})} = VIFnc(j), \quad j=2,\dots,k,$$ and, then, it is possible to state that the VIFnc is a factor that inflate the variance.

What kind of near multicollinearity detects the VIFnc?

Introduction section showed that the VIFnc detects the traditional definition of non-essential near multicollinearity. When the VIF is related to the index of Stewart, see expression ((ref)), it is possible to conclude that the VIFnc is also able to detect the essential near multicollinearity.

However, the introduction also showed that the VIFnc does not detect the relation between the intercept and the rest of the independent variables, which is explained when the intercept is eliminated in the auxiliary regression. This fact is contradictory to the fact that the VIFnc coincides with the index of Stewart since this measure is able to detect the non essential multicollinearity (see Salmer�n et al. Salmeron2019).

Although, the VIFnc could be fooled including the constant as an independent variable in a model without intercept, it is to say: $$\mathbf{y} = \beta_{1} \cdot \mathbf{X}_{1} + \beta_{2} \cdot \mathbf{X}_{2} + \beta_{3} \cdot \mathbf{X}_{3} + \mathbf{u}.$$

The following results are obtained for the VIFnc of auxiliary regressions of this model from expression ((ref)) for $\mathbf{X}_{1}$, $\mathbf{X}_{2}$ and $\mathbf{X}_{3}$ $$400031.4, 199921.7, 200158.3,$$ while, only considering $\mathbf{X}_{1}$ and $\mathbf{X}_{2}$ it is obtained 199921.7 and for $\mathbf{X}_{1}$ and $\mathbf{X}_{3}$ is equal to 200158.3.

Thus, considering the centered model and calculating the coefficient of determination of the auxiliary regressions as if the model was non-centered, it is possible to detect the non-essential multicollinearity.

Conclusions

This work analyzes the detection of near multicollinearity from non-centered auxiliary regressions obtaining that the VIF obtained from them, VIFnc, coincides with the index of Stewart. It is also provided a definition of the non-essential multicollinearity that generalizes the definition given by Marquardt and Snee MarquardtSnee. Finally, it will be interesting as future work to determine thresholds for the VIFnc from which determine that the degree of near multicollinearity detected is worrying.

thebibliography{999} \bibitem{Belsley1984} Belsley, D.A. (1984). Demeaning conditioning diagnostics throught centering. The American Statistician, 38 (2), pag. 73--77. \bibitem{Cook1984} Cook, R.D. (1984). Comment on Demeaning conditioning diagnostics throught centering by Belsley, D.A. The American Statistician, 38 (2), pag. 78--79. \bibitem{MarquardtSnee} Marquardt, D.W. y Snee, R. (1975). Ridge regression in practice. The American Statistician, 29(1), pag. 3--20. \bibitem{Marquardt1980} Marquardt, D.W. (1980). Comment on A Critique of Some Ridge Regression Methods by Smith, G. and Campbell, F.: You Should Standardize the Predictor Variables in Your Regression Models. Journal of the American Statistical Association, 75 (369), pag. 87--91. \bibitem{Salmeron2018} Salmer�n, R., Garc�a, C.B. y Garc�a, J. (2018). Variance inflation factor and condition number in multiple linear regression. Journal of Statistical Computation and Simulation, 88 (12), pag. 2365--2384. \bibitem{Salmeron2019} Salmer�n, R., Rodr�guez, A. y Garc�a, C.B. (2019). \emph{Diagnosis and quantification of the non-essential collinearity}. Computational Statistics (under review). \bibitem{Stewart1987} Stewart, G.W. (1987). \emph{Collinearity and least squares regression}. Statistical Science, 2 (1), pag. 68--100. \bibitem{Theil1971} Theil, H. (1971). Principles of Econometrics. John Wiley & Sons, New York. \bibitem{Velilla2018} Velilla, S. (2018). \emph{A note on collinearity diagnostics and centering}. The American Statistician, 72, pag. 140--146.