Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
40,710 characters · 7 sections · 26 citation commands
Regression based thresholds in principal loading analysis
Principal loading analysis (PLA) is a tool developed by BD21 to reduce dimensions. The method chooses a subset of observed variables by discarding the other variables based on the impact of the eigenvectors on the covariance matrix. BA21 extended the method to correlation-based PLA in order to provide a scale invariant approach.
By nature, some parts of PLA correspond with principal component analysis (PCA) which is a popular dimension reduction technique first formulated by PE01 and HO33. Both are based on the eigendecomposition of the covariance (correlation) matrix and take into account the magnitude of dispersion of each eigenvector. Despite the intersections however, the outcome is different because PCA yields a reduced set of variables, named principal components (PCs), by transforming the original variables.
PCA has been also extended to ordinary least squares (OLS) regression contexts by HO57 and KE57. Instead of regressing on the original variables, in principal component regression the PCs, which are orthogonal by construction, are used as predictors to address multicollinearity. There also exist approaches to reduce the number of PCs based on principal component regression (see HO76 and MW77 among others).
However, even if reducing the number of PCs yields yet another reduced set of all original variables and not a subset of the original variables, it is natural to examine an extension to linear regression contexts also for PLA. We provide such an extension to multivariate linear regression (MLR). Linear regression is then a special case of our results with reduced dimension. Two things have to be taken into account: firstly, MLR can be interpreted as a dimensionality reduction tool itself when predictors are eliminated if the corresponding regression coefficients are not significantly different from zero. Hence, a subset of the original variables is selected by discarding the non-significant variables, which can be done by backwards elimination. Secondly, MLR considers predictors $\bm{X}_\mathcal{K}$ that are regressed on independent variables $\bm{X}_\mathcal{D}$, while PLA considers cross-sectional data. To address this difference, PLA is conducted on $\bm{X} \equiv (\bm{X}_\mathcal{K},\bm{X}_\mathcal{D})$ in this work to take into account both the predictors as well as the independent variables. The idea to perform the method not only on $\bm{X}_\mathcal{D}$ but rather on $\bm{X}$ is not new and has been used for latent root regression, a variation of PC regression, which has been formulated by HA73 and WGM74, however, not in a multivariate regression context.
In this work, we check for intersections of MLR and PLA in a dimensionality reduction-sense. That is, we examine the conditions under which the same variables are selected hence the same variables are discarded by MLR and PLA. We provide necessary as well as necessary and sufficient conditions for those intersections. AN63, KN93 and NW90 provide convergence rates for the sample noise of covariance matrices which are useful for our results because PLA and MLR only share the same selection of variables in the sample case and not in the population case except for the trivial case of no perturbation of the covariance matrix. We interpret sample values as their perturbed population counterparts. Hence, our results are based on matrix perturbation theory which originally comes from algebraic fields such as in IN09 and SS90. The contributed bounds for the cut-off value are based on the Davis-Kahan theorem (DK70 and YW15) as it was done in the work of BD21.
This article is organized as follows: Section (ref) provides notation needed for the remainder of this work. In Section (ref), we recap PLA and introduce the concept of our approach for comparison. An intuition for our conclusions is provided by Lemma (ref) in Section (ref) which is the base for all remaining results. Further, we contribute bounds for the perturbation matrices under which PLA and MLR regression select the same variables and under which they yield different results. Same is done for the threshold value in PLA in Section (ref). We give an example in Section (ref) and take a resume as well as suggest extensions in Section (ref).
We first state some notation and assumptions. Generally, we use index sets to assign submatrices which is extensively needed in this work. For any set $\mathcal{K} = \{k_1,\ldots,k_K \} \subset \mathbb{N}$ with $k_1<\ldots<k_K$, $\bm{A}_{[\{r\},\mathcal{K}]}$ denotes the $r$th row of columns $k_1,\ldots,k_K$ for a matrix $\bm{A}$. For columns, $\bm{A}_{[\mathcal{K},\{c\}]}$ follows the same notation. Further, we denote the elements of any column vector $\bm{v}$ by $\bm{v} = ( v^{(1)} , \ldots , v^{(M)} )$.
We consider $\bm{x} = (\bm{x}_1,\ldots,\bm{x}_M) \equiv (\bm{x}_\mathcal{K} , \bm{x}_\mathcal{D})$ $\in\mathbb{R}^{N\times M}$ to be an independent and identically distributed (IID) sample. $\bm{x}$ contains $n\in\{1,\ldots,N\}$ with $N>M$ observations of a random vector $\bm{X} = (X_1,\ldots,X_M) \equiv (\bm{X}_\mathcal{K}, \bm{X}_\mathcal{D})$ with $\bm{X}_\mathcal{K}^\top \in \mathbb{R}^K$ and $\bm{X}_\mathcal{D}^\top \in \mathbb{R}^D$ such that $K + D = M$. The covariance matrix of $\bm{X}$ is given by $\bm{\Sigma} = (\sigma_{i,j})$ for defined indices $i,j\in\{1,\ldots,M\}$ throughout this work. Further, $\bm{\Sigma}_1$ and $\bm{\Sigma}_2$ are the covariance matrices of $\bm{X}_\mathcal{K}$ and $\bm{X}_\mathcal{D}$ respectively such that
Analogously to $\bm{x}$, we introduce the IID sample $\tilde{\bm{x}} = (\tilde{\bm{x}}_1,\ldots,\tilde{\bm{x}}_M) \equiv (\tilde{\bm{x}}_\mathcal{K} , \tilde{\bm{x}}_\mathcal{D})$ $\in\mathbb{R}^{N\times M}$ drawn from a random vector $\tilde{\bm{X}} = (\tilde{X}_1,\ldots,\tilde{X}_M) \equiv (\tilde{\bm{X}}_\mathcal{K}, \tilde{\bm{X}}_\mathcal{D})$ with covariance matrix $\tilde{\bm{\Sigma}}$. The intuition is that $\bm{X}$ and $\tilde{\bm{X}}$, hence also $\bm{x}$ and $\tilde{\bm{x}}$, behave similarly, however, with the distinction that $\tilde{\bm{\Sigma}} \equiv \bm{\Sigma} + \bm{\mathrm{E}}$ is slightly perturbed by a sparse matrix $\bm{\mathrm{E}} = (\varepsilon_{i,j})$. $\bm{\mathrm{E}}$ is a technical construction that contains small components extracted from $\bm{\Sigma}$ hence $\varepsilon_{i,j} \neq 0 \Rightarrow \sigma_{i,j} = 0$. The sample counterparts of $\bm{\Sigma}$ and $\tilde{\bm{\Sigma}}$ are given by
and
respectively. $\bm{\mathrm{H}}$ and $\tilde{\bm{\mathrm{H}}}{}$ are perturbations in the form of random noise matrices. The noise is due to a finite number of sample observations. $\hat{\bm{\Sigma}}$, $\hat{\tilde{\bm{\Sigma}}}$, and all of their compound matrices follow an analogous block-structure as $\bm{\Sigma}$ given in ((ref)).
Further, we consider the MLR models $\bm{x}_\mathcal{K} = \bm{x}_\mathcal{D} \bm{\beta} + \bm{u}$ and $\tilde{\bm{x}}_\mathcal{K} = \tilde{\bm{x}}_\mathcal{D} \tilde{\bm{\beta}} + \tilde{\bm{u}}$ with $\bm{u},\tilde{\bm{u}}\in\mathbb{R}^{N\times K}$, $\bm{\beta} = (\bm{\beta}_1,\ldots,\bm{\beta}_K) \in \mathbb{R}^{D\times K}$ and $\tilde{\bm{\beta}} = (\tilde{\bm{\beta}}_1,\ldots,\tilde{\bm{\beta}}_K)\in \mathbb{R}^{D\times K}$ where $\bm{\beta}_k,\tilde{\bm{\beta}}_k \in \mathbb{R}^D$ for $k\in\{1,\ldots,K\}$. We denote the estimators for $\bm{\beta}$ and $\tilde{\bm{\beta}}$ by $\hat{\bm{\beta}}$ and $\hat{\tilde{\bm{\beta}}}$ respectively.
We use the $\mathcal{L}_1$ and $\mathcal{L}_2$ vector norms $\Vert\bm{v}\Vert_1 \equiv \sum_i |v^{(i)}|$ and $\Vert\bm{v}\Vert_2 \equiv \left(\sum_i (v^{(i)})^2\right)^{1/2}$, and the Frobenius norm $\Vert\bm{A}\Vert_F$ as a subordinate extension for matrices. Further, the trace of a matrix $\bm{A}$ is given by ${\rm tr}(\bm{A})$. $\mathcal{O}_p(\cdot)$ denotes stochastic boundedness and we use the symbol "$\perp \!\!\! \perp$" as a shortcut that variables have zero correlation and therefore not necessarily for independence. Further, $\bm{0}$ denotes a matrix containing zeros and its dimension its always self-explanatory.
Variables or blocks of variables are linked to eigenvectors in PLA. Hence, for practical purposes we define $\mathcal{D} \subset \{1,\ldots,M\}$ with ${\rm card}(\mathcal{D}) = D$ such that $1\leq D < M$. $\mathcal{D}$ will contain the indices of a variable-block $\{\tilde{X}_d\} \equiv \{\tilde{X}_d\}_{d\in\mathcal{D}}$ we discard according to PLA, and we introduce the quasi-complement $\mathcal{K} \equiv \{1,\ldots,M\} \backslash \mathcal{D}$ with ${\rm card}(\mathcal{K}) = K$ to label the remaining variables that are kept. In an analogue manner, $\mathit{\Delta}$ with ${\rm card}(\mathit{\Delta}) = D$ will be used to index eigenvectors with respective eigenvalues linked to the $\{\tilde{X}_d\}$. $\mathcal{D}\times\mathit{\Delta} \equiv \{(d,\delta) : d\in\mathcal{D}, \; \delta\in\mathit{\Delta}\}$ denotes the Cartesian product and therefore contains all possible combinations of paired indices $d$ linked to $\delta$. For convenience purposes, we consider only a single block $\{\tilde{X}_d\}$ to be discarded throughout this work except for the example in Section (ref). However, our results remain true for any number of blocks.
Given the notation above, the following assumptions are made:
Zero correlation among the random variables is a technical construction that allows to extract the correlations among $(\tilde{X}_1,\ldots,\tilde{X}_{M})$ using the perturbation $\bm{\mathrm{E}}$ of the covariance matrix. Hence, we do not assume that there is zero correlation among the random variables $(\tilde{X}_1,\ldots,\tilde{X}_{M})$.
We assume variables to have mean zero for convenience purposes to simplify the estimators of the covariance matrices to $\hat{\bm{\Sigma}} = (N-1)^{-1}\bm{x}^\top\bm{x}$ and $\hat{\tilde{\bm{\Sigma}}} = (N-1)^{-1}\tilde{\bm{x}}^\top\tilde{\bm{x}}$ respectively.
It is assumed that the regression coefficients can be estimated using OLS. The respective sample linear models are $\bm{x}_\mathcal{K} = \bm{x}_\mathcal{D} \hat{\bm{\beta}} + \bm{u}$ and $\tilde{\bm{x}}_\mathcal{K} = \tilde{\bm{x}}_\mathcal{D} \hat{\tilde{\bm{\beta}}} + \tilde{\bm{u}}$ with $\hat{\bm{\beta}} = \hat{\bm{\Sigma}}_2^{-1} \hat{\bm{\Sigma}}_{12}^\top \equiv (\bm{x}_\mathcal{D}^\top\bm{x}_\mathcal{D})^{-1} \bm{x}_\mathcal{D}^\top \bm{x}_\mathcal{K}$ and $\hat{\tilde{\bm{\beta}}} = \hat{\tilde{\bm{\Sigma}}}_2^{-1} \hat{\tilde{\bm{\Sigma}}}_{12}^\top \equiv (\tilde{\bm{x}}_\mathcal{D}^\top\tilde{\bm{x}}_\mathcal{D})^{-1} \tilde{\bm{x}}_\mathcal{D}^\top \tilde{\bm{x}}_\mathcal{K}$ due to Assumption (ref). If the regression coefficient of a variable is not-significantly different from zero, we say that the respective variable is a non-significant predictor and that the variable can be discarded according to MLR. We implicitly assume that the fourth-order moments $ \mathbb{E}( \bm{X}^\top\bm{X} \otimes \bm{X}^\top\bm{X}) < \infty$ and $ \mathbb{E}( \tilde{\bm{X}}^\top\tilde{\bm{X}} \otimes \tilde{\bm{X}}^\top\tilde{\bm{X}}) < \infty$ are finite because OLS estimation is sensitive to outliers. This further allows us to obtain the rate of convergence for the sample covariance matrices by applying the Lindeberg-L\'evy central limit theorem.
This assumption is in line with the work of BD21 where the asymptotics of eigenvalues were considered. The reason behind this is, that asymptotics for eigenvalues with a multiplicity greater than one are problematic to obtain as pointed out by DPR82. In this work, we are not concerned with the limits of eigenvalues. However, we make usage of the Davis-Kahan Theorem and hence divide by the difference of eigenvalues. Assumption (ref) therefore is rather strict, but by loosing it one has to assume that the difference of the respective eigenvalues in Section (ref) is unequal to zero. Further, we assume that the population covariance matrix $\bm{\Sigma}$ and the sample covariance matrix $\hat{\tilde{\bm{\Sigma}}}$ of $\tilde{\bm{x}}$ are positive definite and hence invertible.
In this section, we recap PLA and state our general approach to derive the conditions under which PLA and MLR share the same selection of variables. The following summary of PLA is based on the work of BD21 and we refer to this for a more elaborate explanation.
PLA is a tool for dimension reduction where a subset of existing variables is selected while the other variables are discarded. The intuition is that blocks of variables are discarded which distort the covariance matrix only slightly. This is done by checking if all $K$ elements $k\in\mathcal{K}$ of $D$ eigenvectors $\{\tilde{v}_{\delta}^{(k)}\}_{\delta\in\mathit{\Delta}}$ of the covariance matrix $\tilde{\bm{\Sigma}}$ are smaller in absolute terms than a certain cut-off value $\tau$. That is to say, we check if $|\tilde{v}_{\delta}^{(k)}| \leq \tau$ for all $(k,\delta)\in\mathcal{K}\times\mathit{\Delta}$. If this is the case, then most distortion of the variables $\{\tilde{X}_d\}$ is geometrically represented by the eigenvectors $\{\tilde{\bm{v}}_{\delta}\}_{\delta\in\mathit{\Delta}}$. The explained variance of the $\{\tilde{X}_d\}$ can then be evaluated either by $N^{-1}\sum_{d\in\mathcal{D}} \tilde{\sigma}_{d,d}$ or approximated by $( \sum_{m\in\mathcal{P}} \tilde{\lambda}_m )^{-1} ( \sum_{\delta\in\mathit{\Delta}} \tilde{\lambda}_{\delta} ) $ where $\tilde{\lambda}_i$ is the eigenvalue corresponding to the eigenvector $\tilde{\bm{v}}_i$. If the explained variance is sufficiently small for the underlying purpose of application, the variables $\{\tilde{X}_d\}$ are discarded. In practice, the population eigenvectors and eigenvalues are replaced by their sample counterparts.
For now, we consider the extreme case that a variable is discarded according to PLA using a threshold $\tau = 0$. According to this, the eigenvectors need to contain zeros which reflects zero correlation among variables. In MLR on the other hand, a predictor that is neither correlated with the remaining predictors nor with the dependent variables has regression coefficients equal to zero. Hence, if the goal is to select a number of predictors, the variable that is not correlated is discarded. Therefore, there is a natural link between PLA and MLR based on the correlation structure and we cover those underlying structures that result in the same selections of variables, both according to PLA and MLR.
We do not take the explained variance of blocks into account despite its importance for PLA. The reason is that it is only relevant if predictors lie in the same block as the independent variable to select them in a PLA sense.
In this section, we derive conditions for the perturbation matrices that have to hold in order to obtain the same selected variables by PLA and MLR. To be more precise: we provide conditions under which discarded variables according to PLA are also considered to be non-significant predictors in MLR. Here, a non-significant predictor means that the regression coefficients of the respective variables are not significantly different from zero.
The following theorem provides an intuition of the relationship between the perturbation matrices:
The intuition is that the perturbation together with the noise of $\tilde{\bm{\Sigma}}$ shall not billow more than pure noise itself. Simply put, this can either happen in the trivial case when $\bm{\mathrm{E}} = \bm{0}$, or when $\bm{\mathrm{E}}$ and $\tilde{\bm{\mathrm{H}}}{}$ fluctuate towards different directions hence pull themselves back below $\bm{\mathrm{H}}$. However, strictly spoken the perturbations are sandwiched hence we have to take multiplication with $(\bm{\Sigma}_2^{-1})_{[\{{d^\ast}\},\mathcal{D}]} $ into account. Nonetheless, Lemma (ref) provides an intuition regarding the perturbations.\\ More detailed results regarding the perturbation and noise are given in Theorem (ref) and in Corollary (ref). The results provide necessary as well as necessary and sufficient conditions under which PLA and MLR share the same selection of variables.
If $\tilde{X}_{d^\ast}$ is discarded according to PLA, the variable was weakly correlated with all $\{\tilde{X}_k\}_{k\in\mathcal{K}}$. Further, the smaller the partial correlation of $\tilde{X}_{d^\ast}$ within $\{\tilde{X}_d\}$, the more likely $\tilde{X}_{d^\ast}$ is also considered to be a non-significant predictor. In practice, $(\bm{\Sigma}_2^{-1})_{[\{{d^\ast}\},\mathcal{D}]}$ can be estimated by $(\hat{\bm{\Sigma}}_2^{-1})_{[\{{d^\ast}\},\mathcal{D}]} $. We can further bound the perturbations of $\tilde{\bm{\Sigma}}$ by noise of $\bm{\Sigma}$ only, which reflects the intuition discussed in Lemma (ref).
Corollary (ref) reflects what we discussed above, that is that the perturbation together with the noise of $\tilde{\bm{\Sigma}}$ shall not billow more than pure noise itself. Further, it is indicated that there are conditions under which PLA and MLR select the same variables in the sample case, however not necessarily in the population case. This is due to the perturbation $\bm{\mathrm{E}}$ as stated in the following result.
BD21 provided convergence rates for the noise present in all sample parts needed for PLA. However, a bound for the perturbation $\bm{\mathrm{E}}$ was not given. Looking from an MLR perspective onto PLA, $\bm{\mathrm{E}}$ has to be bounded by $\mathcal{O}_p(N^{-1/2})$. For the trivial case that $\bm{\mathrm{E}} = \bm{0}$, MLR and PLA share the same selection of variables also for the population case. When perturbation is present however, that is to say when the variables to be discarded according to PLA are correlated with the remaining variables, both methods can only select the same variables in the finite sample case.
Choosing an appropriate threshold $\tau$ is crucial for PLA. BD21 and BA21 provide feasible thresholds for PLA based on simulations. In this section, we translate our results from Section (ref) to the choice of such a cut-off value. We provide bounds for the threshold $\tau$ that ensure discarding according to PLA as well as testing the respective regression coefficient to be not significantly different from zero on a $(1-\alpha)\cdot 100\%$ level hence ensure discarding in a multivariate regression sense. Therefore, the threshold depends on the significance level $\alpha$. Further, we provide bounds for the perturbations $\bm{\mathrm{E}} + \tilde{\bm{\mathrm{H}}}{}$ in this section.
The procedure to test if all regression coefficients are significantly different from zero is well known. Commonly, the tests based on Roy's largest root, Wilks' Lambda, Hotelling-Lawley trace and Pillai-Bartlett trace are used. In a multivariate regression sense, all variables $\{\tilde{X}_d\}$ are discarded if we fail to reject the null of $\bm{\beta} = \bm{0}$.
Choosing $\tau_\alpha \equiv \dfrac{2^{2/3} F_\alpha \Vert \hat{\tilde{\bm{\Sigma}}}_1^{-1} \Vert_F^{-1/2} \Vert \hat{\tilde{\bm{\Sigma}}}_2^{-1} \Vert_F^{-1/2} S D_1 D_2^{-1} }{\min_{\delta\in\mathit{\Delta}} \{\lambda_{\delta-1} - \lambda_\delta, \lambda_\delta - \lambda_{\delta+1}\}} $ is sufficient, not necessary and sufficient. This is because the terms are bound from above in the derivations.
Since $\hat{\tilde{\bm{\Sigma}}} \equiv \bm{\Sigma} + \bm{\mathrm{E}} + \tilde{\bm{\mathrm{H}}}{}$, it is difficult to obtain bounds for $\bm{\mathrm{E}}$ and $\tilde{\bm{\mathrm{H}}}{}$ having only access to $\hat{\tilde{\bm{\Sigma}}}$ in practice. However, if the $\{\tilde{X}_d\}$ are not discarded according to multivariate regression, we can find a lower bound for $\bm{\mathrm{E}}+ \tilde{\bm{\mathrm{H}}}{}$.
For completion, we further give necessary conditions in line with the intuition from Section (ref).
Note that the case when we focus only on a single variable which belongs to a block $\{\tilde{X}_d\}$ in PLA-sense instead of considering the whole block itself, the maximization operator in Theorem (ref) simplifies. The reason is that the link between the variable and the eigenvector with respective eigenvalue is evident in this case. For a block of variables it is not obvious which variable is linked to which eigenvectors and hence all eigengaps $ \{\lambda_{\delta-1} - \lambda_\delta, \lambda_\delta - \lambda_{\delta+1}\}$ for $\delta\in\mathit{\Delta}$ have to be considered.
Same holds also for Corollary (ref) which also simplifies when considering only a single variable.
In practice, the eigenvalues $\lambda_\delta$ can be estimated by $\hat{\tilde{\lambda}}_\delta$. BD21 outlined that estimation is biased by $\Vert \bm{\mathrm{E}} \Vert_F$ which is small, however, due to the sparseness of $\bm{\mathrm{E}}$.
We provide an example how multivariate regression can support the threshold choice in PLA. Firstly, we conduct PLA with a threshold $\tau$ to detect the underlying blocks. Afterwards, the threshold is adjusted to $\tau_\alpha$ according to a MLR-sense. The blocks can then either be discarded by PLA with threshold $\tau_\alpha$ as well, which corresponds to the case that both methods choose the same variables, or are not discarded by PLA with threshold $\tau_\alpha$. The latter one corresponds to the case where variables are discarded by PLA with threshold $\tau$ but not according to MLR. Further, we give an example where a variable is discarded by MLR, however, not by PLA. All values are rounded to four decimal places.
We assume a sample to be drawn from the random vector $(\tilde{X}_1, \tilde{X}_2, \tilde{X}_3, \tilde{X}_4, \tilde{X}_5)$ and we generated a respective sample $\tilde{\bm{x}}$ of $N=100$ observations. The sample covariance matrix $\hat{\tilde{\bm{\Sigma}}}$ can be found in (ref) while the eigenvectors $\hat{\tilde{\bm{v}}}_m$ and eigenvalues $\hat{\tilde{\lambda}}_m$ with $m\in\{1,\ldots,5\}$ are in Table (ref).
We choose $\tau = 0.3$ as recommended by BD21, and $\alpha = 0.05$ for our analysis. PLA detects three Blocks: $\tilde{X}_1$ ($1\times1$), $\tilde{X}_2$ ($1\times1$) and $(\tilde{X}_3,\tilde{X}_4,\tilde{X}_5)$ ($3\times3$) as can be seen in Table (ref). We now want to bear the PLA results from a MLR point of view and check the two $1\times1$ blocks $\tilde{X}_1$ and $\tilde{X}_2$ according to Corollary (ref).
Starting with $\tilde{X}_1$, we consider the MLR model $(\tilde{\bm{x}}_2,\tilde{\bm{x}}_3,\tilde{\bm{x}}_4,\tilde{\bm{x}}_5) = \tilde{\bm{x}}_1 \tilde{\bm{\beta}} + \tilde{\bm{u}}$. Hence, $D_1 \equiv 1 \cdot 4 =4$, $D_2 \equiv 1 \cdot (100 - 1 - 4) = 95$, and $F_\alpha = 2.4675$. The block is reflected by the second eigenvector and therefore we have that $\delta = 2$. Since $\Vert \hat{\tilde{\bm{\Sigma}}}_1^{-1} \Vert_F^{-1/2} = 0.9597$, $\Vert \hat{\tilde{\bm{\Sigma}}}_2^{-1} \Vert_F^{-1/2} = 0.4299$ ((ref)), and $\min\{ \hat{\tilde{\lambda}}_1 - \hat{\tilde{\lambda}}_2, \hat{\tilde{\lambda}}_2 - \hat{\tilde{\lambda}}_3\} = 1.3032$, we can conclude that a sufficient threshold to discard $\tilde{X}_1$ by MLR is given by $\hat{\tau}_{\alpha,1} = 0.2753$. From Table (ref) we can see that PLA detects the block containing $\tilde{X}_1$ also for $\hat{\tau}_{\alpha,1}$ making it a feasible threshold for both, PLA and MLR. Therefore, the result we have gotten from PLA in the first place is also supported by regression analysis.
On the other hand, for the block consisting of $\tilde{X}_2$ we consider the MLR model $(\tilde{\bm{x}}_1,\tilde{\bm{x}}_3,\tilde{\bm{x}}_4,\tilde{\bm{x}}_5) = \tilde{\bm{x}}_2 \tilde{\bm{\beta}} + \tilde{\bm{u}}$ with $\Vert \hat{\tilde{\bm{\Sigma}}}_1^{-1} \Vert_F^{-1/2} = 0.5568$, $\Vert \hat{\tilde{\bm{\Sigma}}}_2^{-1} \Vert_F^{-1/2} = 0.9326$ ((ref)), and $\min\{ \hat{\tilde{\lambda}}_4 - \hat{\tilde{\lambda}}_5, \hat{\tilde{\lambda}}_5 - \hat{\tilde{\lambda}}_6\} = \hat{\tilde{\lambda}}_4 - \hat{\tilde{\lambda}}_5 = 3.7965$. $\delta = 5$ because $\tilde{X}_2$ is reflected by the fifth eigenvector. We can conclude that a sufficient threshold to discard $\tilde{X}_2$ by MLR is given by $\hat{\tau}_{\alpha,2} = 0.0751$. Using this threshold for PLA however, $\tilde{X}_2$ is not detected as a block. Hence, this block is not supported by MLR and we can conclude that $\Vert \ \bm{\mathrm{E}}+ \tilde{\bm{\mathrm{H}}}{} \Vert_F > \Vert \hat{\tilde{\bm{\Sigma}}}_1^{-1} \Vert_F^{-1/2} \Vert \hat{\tilde{\bm{\Sigma}}}_2^{-1} \Vert_F^{-1/2} \dfrac{S D_1}{D_2 \cdot F_\alpha^{-1} +D_1} \approx 0.1795$ according to Corollary (ref)
Considering the regression model $(\tilde{\bm{x}}_1,\tilde{\bm{x}}_2,\tilde{\bm{x}}_4,\tilde{\bm{x}}_5) = \tilde{\bm{x}}_3 \tilde{\bm{\beta}} + \tilde{\bm{u}}$, the estimated regression coefficients $\hat{\tilde{\bm{\beta}}}$ are not significantly different from $\bm{0}$ on a 5% level. In fact, the $p$-values for the approximated $F$-tests based on Roy's largest root, Wilks' Lambda, Hotelling-Lawley trace and Pillai-Bartlett trace all equal 0.4685 with the respective value of the test statistic being $0.8978$ ((ref)). Hence, $\tilde{X}_3$ is discarded in a MLR sense. However, $\tilde{X}_3$ is not considered to be a block in PLA-sense which means that $ \Vert \hat{\tilde{\bm{\Sigma}}}_1^{-1} \Vert_F^{-1/2} \Vert \hat{\tilde{\bm{\Sigma}}}_2^{-1} \Vert_F^{-1/2} \dfrac{S D_1}{D_2 \cdot F_\alpha^{-1} +D_1} \approx 0.2259 > \Vert \ \bm{\mathrm{E}}+ \tilde{\bm{\mathrm{H}}}{} \Vert_F > 2^{-2/3} \tau \cdot \min\{ \hat{\tilde{\lambda}}_{\delta-1} - \hat{\tilde{\lambda}}_\delta, \hat{\tilde{\lambda}}_\delta - \hat{\tilde{\lambda}}_{\delta+1}\}^{-1}$ ((ref)). Since $\tilde{X}_3$ is not detected to be a $1\times 1$ block according to PLA, it is difficult to choose the eigenvector that represents $\tilde{X}_3$ in order to calculate the respective eigengaps. However, it is reasonable to assume the vector to be either $\hat{\tilde{\bm{v}}}_3$ or $\hat{\tilde{\bm{v}}}_4$ because $\tilde{X}_3$ obtains the largest values in them. Since $\min_{\delta=3} \{ \hat{\tilde{\lambda}}_{\delta-1} - \hat{\tilde{\lambda}}_\delta, \hat{\tilde{\lambda}}_\delta - \hat{\tilde{\lambda}}_{\delta+1}\} = \hat{\tilde{\lambda}}_3 - \hat{\tilde{\lambda}}_{4} = 1.1340$ and $\min_{\delta=4} \{ \hat{\tilde{\lambda}}_{\delta-1} - \hat{\tilde{\lambda}}_\delta, \hat{\tilde{\lambda}}_\delta - \hat{\tilde{\lambda}}_{\delta+1}\} = \hat{\tilde{\lambda}}_3 - \hat{\tilde{\lambda}}_{4} = 1.1340$ as well, we can conclude that $0.2259 > \Vert \ \bm{\mathrm{E}}+ \tilde{\bm{\mathrm{H}}}{} \Vert_F > 0.1667$.
Overall, we covered three different outcomes. Choosing the threshold based on an OLS regression point of view validates the block containing the single variable $\tilde{X}_1$. Hence, we obtain the block $\tilde{X}_1$ and the block $(\tilde{X}_2,\tilde{X}_3,\tilde{X}_4,\tilde{X}_5)$ respectively. The block consisting of $\tilde{X}_2$ cannot be detected by PLA with threshold $\tau_{\alpha,2}$. Therefore, $\tilde{X}_2$ is not discarded by MLR, however, it is by PLA with threshold $\tau = 0.3$. Lastly, $\tilde{X}_3$ is considered to be redundant in an regression-sense but not by PLA.
We have provided conditions under which PLA and MLR select the same variables and hence discard the same variables. The intuition of those conditions is that the perturbation of PLA together with the noise shall not billow more than pure noise itself. Further, we have contributed bounds for the threshold $\tau$ in PLA that ensure discarding by both PLA and regression analysis. This helps to adjust $\tau$ according to the underlying data rather than applying fixed cut-off values. We also gave bounds for the perturbation matrices.
The choice of $\tau = \tau_\alpha$ is based on the test for regression coefficients with significance level $\alpha$. With the threshold being a function of $\alpha$, $\tau_\alpha$ increases when $\alpha$ decreases. A small $\alpha$ means that the likelihood to test correctly if the regression coefficient equals zero is large. However, it increases the likelihood of conducting a type II error which is not addressed yet and should be part of future research.
Also, our derivations are based on the Pillai-Bartlett trace test which can be constructed to follow an $F$ distribution approximately. Since such approximations also exist for Roy’s largest root, Wilks’ Lambda, and Hotelling-Lawley trace, derivations based on those tests might be useful for small sample size approximations. Hotelling-Lawley trace is similarly constructed like Pillai-Bartlett trace, however, the other tests consider determinants rather than traces which makes it hard to obtain bounds with respect to matrix norms.