Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Testing for Structural Change under Nonstationarity
Wald OLS test for mildly integrated regressors
We consider separately the limiting distribution of the sup Wald-OLS statistic when the regressor is assumed to be generated via a mildly integrated process. The econometric intuition in this case is that since the regressor is mildly integrated then it is expected to behave asymptotically similar to the IVX instrument. Therefore, we replace $z_t$ with $x_t$ into the corresponding sample moments and obtain the corresponding limiting terms. Moreover, intuitively in the case we have a mildly integrated regressor and since the degree of persistence is controlled by the exponent rate $\gamma \in (0,1)$, then we expect that the limiting distribution of the sup Wald-OLS statistic, when testing for an unknown break-point $\pi \in \Pi$, to weakly converge to the standard NBB.
Consider the univariate predictive regression with multiple regressors
align[align omitted — 128 chars of source]
where $I_{1t} := \mathbf{1} \{ t \leq k \}$ and $I_{2t} := \mathbf{1} \{ t > k \}$ with $k = \floor{ T\pi}$. The set of regressors $x_t$ is generated via the following process
align[align omitted — 111 chars of source]
where $\gamma \in (0,1)$ is the exponent rate of the degree of persistence.
Single regressor
Equivalently, to simplify the asymptotics of the following Proposition we consider the univariate predictive regression with a single mildly integrated regressor (no intercept)
align[align omitted — 145 chars of source]
We use the following asymptotic terms
align[align omitted — 524 chars of source]
propositionUnder conditions A1 and A2 of Assumption 1 of the paper, the standard Wald OLS statistic given by the following expression
\begin{align}
\mathcal{W}_T( \pi )
&= \frac{1}{ \hat{\sigma}_u^2 } \left( \hat{\beta}_1 - \hat{\beta}_2 \right)^{\prime} \left[ \mathcal{R} \left( X^{\prime} X \right)^{-1} \mathcal{R}^{\prime}\right]^{-1} \left( \hat{\beta}_1 - \hat{\beta}_2 \right)
\end{align}
for testing the null hypothesis $\mathbb{H}_0: \beta_1 = \beta_2$, when the regressor is assumed to be generated via the following process
\begin{align}
x_t = \left( 1 - \frac{c}{ T^{\gamma} } \right) x_{t-1} + v_{t}, \ \ \ \ with \ \ x_0 = 0, \ \ and \ \ \gamma \in (0,1).
\end{align}
is found to have the following limiting distribution
\begin{align}
\mathcal{W}^{*}_T( \pi ) = \underset{ \pi \in [ \pi_1 , \pi_2 ] }{ sup } \ \mathcal{W}_T( \pi ) \Rightarrow \underset{ \pi \in [ \pi_1 , \pi_2 ] }{ sup } \ \frac{ \bigg[ W( \pi) - \pi W( 1 ) \bigg]^2 }{\pi (1 - \pi)}
\end{align}
remarkProposition 1 above shows that when the regressor has persistence properties assumed to be fall in the realm of mildly integrated processes, then the limiting distribution when testing for a structural change in a univariate predictive regression with a single regressor and no intercept, follows a standard NBB limit similar to the classical linear regression case as proved by Andrews (1993).
proofWe obtain the standard OLS estimators $\hat{\beta}_1$ and $\hat{\beta}_2$ of the corresponding regression coefficients $\beta_1$ and $\beta_2$ as below
\begin{align}
\hat{\beta}_1 &= \frac{ \sum_{t=1}^T x_{t} I_{1t} y_{t+1} }{ \sum_{t=1}^T x^2_{t} I_{1t} } = \beta^0 + \frac{ \sum_{t=1}^T x_{t} I_{1t} u_{t+1} }{ \sum_{t=1}^T x^2_{t} I_{1t} }
\\
\hat{\beta}_2 &= \frac{ \sum_{t=1}^T x_{t} I_{2t} y_{t+1} }{ \sum_{t=1}^T x^2_{t} I_{2t} } = \beta^0 + \frac{ \sum_{t=1}^T x_{t} I_{2t} u_{t+1} }{ \sum_{t=1}^T x^2_{t} I_{2t} }
\end{align}
Thus, assuming that the structural break is at an unknown break point such as $k = [T \pi]$ for $\pi \in (0,1)$ we consider the limiting results under the null hypothesis, $\mathbb{H}_0: \beta_1 = \beta_2$. Note that the FCLT does not apply in this case (mildly integrated predictors). However, we use the limit theory already derived in PM.
\begin{align}
T^{ \frac{1 + \gamma}{2}} \left( \hat{\beta}_1 - \beta^0 \right) = \frac{ \frac{1}{ T^{ \frac{1 + \gamma}{2}} } \sum_{t=1}^{ \floor{ T \pi} } x_{t} u_{t+1} }{ \frac{1}{T^{ 1 + \gamma } } \sum_{t=1}^{ \floor{ T \pi} } x^2_{t} }
\Rightarrow \frac{ \mathcal{N} \left( 0, \pi \frac{ \sigma_u^2 \omega_v^2 }{2c} \right) }{ \pi \frac{ \omega^2_v}{2c} }
\end{align}
\begin{align}
T^{ \frac{1 + \gamma}{2}} \left( \hat{\beta}_2 - \beta^0 \right) = \frac{ \frac{1}{ T^{ \frac{1 + \gamma}{2}} } \sum_{t=\floor{ T \pi} + 1}^{ T } x_{t} u_{t+1} }{ \frac{1}{T^{ 1 + \gamma } } \sum_{t=\floor{ T \pi} + 1}^{ T } x^2_{t} }
\Rightarrow \frac{ \mathcal{N} \left( 0, (1 - \pi) \frac{ \sigma_u^2 \omega_v^2 }{2c} \right) }{ (1 - \pi) \frac{ \omega^2_v}{2c} }
\end{align}
Thus, using (ref) and (ref) we obtain the following simplified expression
\begin{align}
T^{ \frac{1 + \gamma}{2}} \left( \hat{\beta}_1 - \hat{\beta}_2 \right)
\Rightarrow \frac{ \mathcal{N} \left( 0, \pi \frac{ \sigma_u^2 \omega_v^2 }{2c} \right) }{ \pi \frac{ \omega^2_v}{2c} } - \frac{ \mathcal{N} \left( 0, (1 - \pi) \frac{ \sigma_u^2 \omega_v^2 }{2c} \right) }{ (1 - \pi) \frac{ \omega^2_v}{2c} }
\end{align}
Denoting with $X = [ x_{t} I_{1t} \ \ x_{t} I_{2t} ] \equiv [ X_1 \ X_2 ]$ then the Wald test has an equivalent representation as below
\begin{align}
\mathcal{W}_T( \pi )
&= \frac{1}{ \hat{\sigma}_u^2 } \left( \hat{\beta}_1 - \hat{\beta}_2 \right)^{\prime} \left[ \mathcal{R} \left( X^{\prime} X \right)^{-1} \mathcal{R}^{\prime}\right]^{-1} \left( \hat{\beta}_1 - \hat{\beta}_2 \right)
\end{align}
Firstly, it can be easily proved that the following equivalent expression holds, using the orthogonality property of $X_1$ and $X_2$ and via a standard matrix inversion application.
\begin{align*}
\left[ \mathcal{R} \left( X^{\prime} X \right)^{-1} \mathcal{R}^{\prime}\right] &= \left[ \left( \sum_{t=1}^T x^2_{t} I_{1t} \right)^{-1} + \left( \sum_{t=1}^T x^2_{t} I_{1t} \right)^{-1} \right] \\
&= \frac{ \sum_{t=1}^T x^2_{t} I_{1t} + \sum_{t=1}^T x^2_{t} I_{2t} }{ \left( \sum_{t=1}^T x^2_{t} I_{1t} \right) \left( \sum_{t=1}^T x_{t} I_{2t} \right) }
= \frac{ \sum_{t=1}^T x^2_{t} }{ \left( \sum_{t=1}^T x^2_{t} I_{1t} \right) \left( \sum_{t=1}^T x^2_{t} I_{2t} \right) }
\end{align*}
Therefore, the simplified expression of the Wald statistic is given by the expression below in the case of single predictors
\begin{align*}
\mathcal{W}_T( \pi ) = \frac{ \left( \hat{\beta}_1 - \hat{\beta}_2 \right)^2 }{ \hat{\sigma}^2_u } \left[ \frac{ \sum_{t=1}^T x^2_{t} }{ \left( \sum_{t=1}^T x^2_{t} I_{1t} \right) \left( \sum_{t=1}^T x^2_{t} I_{2t} \right) } \right]^{-1}
&=
\frac{ \left( \hat{\beta}_1 - \hat{\beta}_2 \right)^2 }{ \hat{\sigma}^2_u } \frac{ \left( \sum_{t=1}^T x^2_{t} I_{1t} \right) \left( \sum_{t=1}^T x^2_{t} I_{2t} \right) }{ \sum_{t=1}^T x^2_{t} } \\
&=
T^{ 1 + \gamma } \frac{ \left( \hat{\beta}_1 - \hat{\beta}_2 \right)^2 }{ \hat{\sigma}^2_u } \frac{ \left( \sum_{t=1}^T \frac{ x^2_{t} I_{1t}}{ T^{ 1 + \gamma } } \right) \left( \sum_{t=1}^T \frac{ x^2_{t} I_{2t} }{ T^{ 1 + \gamma } } \right) }{ \sum_{t=1}^T \frac{ x^2_{t} }{ T^{ 1 + \gamma } } }
\end{align*}
Moreover, the following asymptotic convergence result also holds
\begin{align}
\frac{ \left( \sum_{t=1}^T \frac{ x^2_{t} I_{1t}}{ T^{ 1 + \gamma } } \right) \left( \sum_{t=1}^T \frac{ x^2_{t} I_{2t} }{ T^{ 1 + \gamma } } \right) }{ \sum_{t=1}^T \frac{ x^2_{t} }{ T^{ 1 + \gamma } } } \Rightarrow \frac{ \pi \frac{ \omega^2_v}{2c} (1 - \pi) \frac{ \omega^2_v}{2c} }{ \frac{ \omega^2_v}{2c} } = \pi(1 - \pi) \frac{ \omega^2_v}{2c}
\end{align}
Thus, we obtain that
\begin{align}
\mathcal{W}_T( \pi )
&\equiv \frac{1}{ \sigma^2_u } \left\{ \frac{ \mathcal{N} \left( 0, \pi \frac{ \sigma_u^2 \omega_v^2 }{2c} \right) }{ \pi \frac{ \omega^2_v}{2c} } - \frac{ \mathcal{N} \left( 0, (1 - \pi) \frac{ \sigma_u^2 \omega_v^2 }{2c} \right) }{ (1 - \pi) \frac{ \omega^2_v}{2c} } \right\}^2 \pi(1 - \pi) \frac{ \omega^2_v}{2c}
\nonumber
\\
\nonumber
\\
&=
\frac{1}{ \sigma^2_u } \frac{ \left\{ (1 - \pi) \ \mathcal{N} \left( 0, \pi \frac{ \sigma_u^2 \omega_v^2 }{2c} \right) - \pi \ \mathcal{N} \left( 0, (1 - \pi) \frac{ \sigma_u^2 \omega_v^2 }{2c} \right) \right\}^2 }{ \pi (1 - \pi) \frac{ \omega^2_v}{2c} }
\end{align}
Thus, simplifying the terms which do not depend on $\pi$ from above expression and using the supremum functional as well since we consider an unknown break-point then we obtain the following expression
\begin{align*}
\mathcal{W}^{*}_T( \pi )
&\Rightarrow \underset{ \pi \in [ \pi_1 , \pi_2 ] }{ sup } \ \frac{ \left\{ (1 - \pi) \ \mathcal{N} \bigg( 0, \pi \bigg) - \pi \ \mathcal{N} \bigg( 0, (1 - \pi) \bigg) \right\}^2 }{ \pi (1 - \pi) }
\nonumber
\\
\nonumber
\\
\nonumber
&\equiv
\underset{ \pi \in [ \pi_1 , \pi_2 ] }{ sup } \ \frac{ \left\{ \mathcal{N} \bigg( 0, \pi \bigg) - \pi \ \mathcal{N} \bigg( 0, 1 \bigg) \right\}^2 }{ \pi (1 - \pi) }
\end{align*}
which shows indeed the weakly convergence to a NBB in the case of a single mildly integrated predictor in the predictive regression model.
In summary, we show that the sup Wald-OLS statistic for testing for a single structural change predictive regressions with mildly integrated predictors, which is the case that the degree of persistence is controlled via the exponent rate $\gamma \in (0,1)$, then we obtain an asymptotically equivalent limiting distribution as in the standard linear regression.
Furthermore, since $\gamma \in (0,1)$ and assuming that the exponent rate for the degree of persistence of the instrument satisfies $\gamma \in ( 0 , \delta )$, then using Lemma 3.5 of PM the following asymptotic terms hold
align[align omitted — 299 chars of source]
Multiple regressors
Next, we consider the case of multiple mildly integrated regressors. First, we consider the following example, which provides useful insights for the related asymptotic terms in the case of the univariate predictive regression with multiple predictors.
exampleConsider the predictive regression with multiple predictors given below
\begin{align}
y_{t+1} = \beta_0 + \beta_1^{\prime} \ x_t + u_{t+1}
\end{align}
We aim to examine the limiting distribution of the parameter vector $\underline{\beta} = ( \beta_0, \underline{\beta}_1^{\prime} )$ with $\beta_0 \in \mathbb{R}$, $\underline{\beta}_1 \in \mathbb{R}^{p \times 1} $ and $\widetilde{x}_t = \left( \underline{1} , \underline{x}^{\prime}_t \right)^{\prime} \in \mathbb{R}^{ T \times (p+1) }$. Note that since the intercept and the vector of predictors have a different convergence rate, we define the normalization matrix: $\mathcal{D}_T = diag( \sqrt{T}, T^{ \frac{ 1 + \gamma }{2} } \text{I}_{p} )\in \mathbb{R}^{(p+1) \times (p+1)}$.
We have that
\begin{align}
\left( \widehat{\beta} - \beta^0 \right) = \left( \sum_{t=1}^T \widetilde{ x }_{t} \widetilde{ x }_{t}^{\prime} \right)^{-1} \left( \sum_{t=1}^T \widetilde{ \underline{x} }_{t} u_{t+1} \right)
\end{align}
where $\underline{\beta}^0$ the true parameter vector under the null hypothesis, $\mathbb{H}_0: \underline{\beta} = \underline{\beta}^0$.
We obtain
\begin{align}
\mathcal{D}_T^{-1} \left( \sum_{t=1}^T \widetilde{\underline{x}}_{t} \widetilde{\underline{x}}_{t}^{\prime} \right) \mathcal{D}_T^{-1}
=
\begin{bmatrix}
1 & \frac{1}{ T^{ \frac{ 1 + \gamma }{2} } } \sum_{t=1}^T \underline{x}_{t}^{\prime}
\\
\\
\frac{1}{ T^{ \frac{ 1 + \gamma }{2} } } \sum_{t=1}^T \underline{x}_{t} & \frac{1}{ T^{ 1 + \gamma }} \sum_{t=1}^T \underline{x}_{t} \underline{x}_{t}^{\prime}
\end{bmatrix}
\Rightarrow
\begin{bmatrix}
1 & \textcolor{red}{0}
\\
\\
\textcolor{red}{0} & V_{zz}
\end{bmatrix}
\end{align}
Similarly, we also obtain the following expression
\begin{align}
\mathcal{D}_T^{-1} \left( \sum_{t=1}^T \widetilde{ \underline{x} }_{t} u_{t+1} \right)
=
\begin{bmatrix}
\frac{1}{ \sqrt{T} } \sum_{t=1}^T u_{t+1}
\\
\frac{1}{ T^{ \frac{ 1 + \gamma }{2} } } \sum_{t=1}^T \underline{x}^{\prime}_t u_{t+1}
\end{bmatrix}
\Rightarrow
\begin{bmatrix}
B_u (r)
\\
\\
\mathcal{N} \bigg( 0, \sigma_u^2 V_{zz} \bigg)
\end{bmatrix}
\end{align}
Therefore, combining the above results, we obtain
\begin{align}
\mathcal{D}_T \left( \underline{ \widehat{\beta} } - \underline{\beta}^0 \right)
&= \left[ \mathcal{D}_T^{-1} \left( \sum_{t=1}^T \widetilde{ \underline{x} }_{t} \widetilde{ \underline{x} }_{t}^{\prime} \right) \mathcal{D}_T^{-1} \right]^{-1} \mathcal{D}_T^{-1} \left( \sum_{t=1}^T \widetilde{ \underline{x} }_{t}^{\prime} u_{t+1} \right)
\nonumber
\\
&\Rightarrow
\begin{bmatrix}
1 & \textcolor{red}{0}
\\
\\
\textcolor{red}{0} & V_{zz}
\end{bmatrix}^{-1}
\times
\begin{bmatrix}
B_u (r)
\\
\\
\mathcal{N} \bigg( 0, \sigma_u^2 V_{zz} \bigg)
\end{bmatrix}
\end{align}
propositionUnder conditions A1 and A2 of Assumption 1 of the paper, the standard Wald OLS statistic given by the following expression
\begin{align}
\mathcal{W}_T( \pi )
&= \frac{1}{ \hat{\sigma}_u^2 } \left( \widehat{\beta} _1 - \widehat{\beta} _2 \right)^{\prime} \left[ \mathcal{R} \left( X^{\prime} X \right)^{-1} \mathcal{R}^{\prime}\right]^{-1} \left( \widehat{\beta} _1 - \widehat{\beta} _2 \right)
\end{align}
for testing the null hypothesis $\mathbb{H}_0: \underline{\beta}_1 = \underline{\beta}_2$, when the regressor is generated via
\begin{align}
x_t = \left( I_p - \frac{C}{ T^{\gamma} } \right) x_{t-1} + \underline{v}_{t}, \ \ \ \ \text{with} \ \ \underline{x}_0 = 0, \ \ \text{and} \ \ \gamma \in (0,1).
\end{align}
is found to have the following limiting distribution
\begin{align}
\mathcal{W}_T( \pi ) \Rightarrow \chi^2_k + \underset{ \pi \in [ \pi_1 , \pi_2 ] }{ \text{ sup } } \ \ \frac{ \mathcal{B B} (\pi) }{ \pi (1 - \pi) }
\end{align}
proofUnder the null hypothesis of no structural break, $\mathbb{H}_0: \underline{\beta}_1 = \underline{\beta}_2$, we have
\begin{align*}
\mathcal{D}_T \left( \widehat{\beta} _1 - \beta^0 \right)
&= \left[ \mathcal{D}_T^{-1} \left( \sum_{t=1}^T \widetilde{ x }_{t} \widetilde{ x }_{t}^{\prime} I_{1t} \right) \mathcal{D}_T^{-1} \right]^{-1} \mathcal{D}_T^{-1} \left( \sum_{t=1}^T \widetilde{ x }_{t}^{\prime} u_{t+1} I_{1t} \right)
\end{align*}
where $\underline{\beta}^0$, the population value of the model coefficient.
Similarly,
\begin{align*}
\mathcal{D}_T \left( \widehat{\beta} _2 - \underline{\beta}^0 \right)
&= \left[ \mathcal{D}_T^{-1} \left( \sum_{t=1}^T \widetilde{ \underline{x} }_{t} \widetilde{ \underline{x} }_{t}^{\prime} I_{2t} \right) \mathcal{D}_T^{-1} \right]^{-1} \mathcal{D}_T^{-1} \left( \sum_{t=1}^T \widetilde{ \underline{x} }_{t}^{\prime} u_{t+1} I_{2t} \right)
\end{align*}
The weakly convergence result for the estimator of $\beta_1$ follows
\begin{align}
\mathcal{D}_T \left( \underline{ \widehat{\beta} }_1 - \underline{\beta}^0 \right)
&\Rightarrow
\begin{bmatrix}
1 & \textcolor{red}{0}
\\
\\
\textcolor{red}{0} & \pi V_{zz}
\end{bmatrix}^{-1}
\times
\begin{bmatrix}
B_u (\pi)
\\
\\
\mathcal{N} \bigg( 0, \pi \sigma_u^2 V_{zz} \bigg)
\end{bmatrix}
\end{align}
Similarly, for the estimator of $\beta_2$ we have the following weakly convergence result
\begin{align}
\mathcal{D}_T \left( \underline{ \widehat{\beta} }_2 - \underline{\beta}^0 \right)
&\Rightarrow
\begin{bmatrix}
1 & \textcolor{red}{0}
\\
\\
\textcolor{red}{0} & (1 - \pi) V_{zz}
\end{bmatrix}^{-1}
\times
\begin{bmatrix}
B_u (1) - B_u (\pi)
\\
\\
\mathcal{N} \bigg( 0, (1 - \pi) \sigma_u^2 V_{zz} \bigg)
\end{bmatrix}
\end{align}
Therefore, we have that
\begin{align}
\mathcal{D}_T \left( \underline{ \widehat{\beta} }_1 - \underline{\beta}^0 \right) - \mathcal{D}_T \left( \underline{ \widehat{\beta} }_2 - \underline{\beta}^0 \right)
\equiv
\mathcal{D}_T \left( \underline{ \widehat{\beta} }_1 - \underline{ \widehat{\beta} }_2 \right)
\end{align}
Recall that the expression for the Wald statistic is as below
\begin{align}
\mathcal{W}_T( \pi )
&=
\frac{1}{ \hat{\sigma}_u^2 } \left( \underline{ \widehat{\beta} }_1 - \underline{ \widehat{\beta} }_2 \right)^{\prime} \left[ \mathcal{R} \left( X^{\prime} X \right)^{-1} \mathcal{R}^{\prime}\right]^{-1} \left( \underline{ \widehat{\beta} }_1 - \underline{ \widehat{\beta} }_2 \right)
\end{align}
Note that, the following holds
\begin{align}
\left[ \mathcal{R} \left( X^{\prime} X \right)^{-1} \mathcal{R}^{\prime}\right]
&= \left[ \left( X_1^{\prime} X_1 \right)^{-1} + \left( X_2^{\prime} X_2 \right)^{-1} \right]
\nonumber
\\
&= \mathcal{D}_T^{-1} \left[ \bigg( \mathcal{D}_T^{-1} \left( X_1^{\prime} X_1 \right) \mathcal{D}_T^{-1} \bigg)^{-1} + \bigg( \mathcal{D}_T^{-1} \left( X_2^{\prime} X_2 \right) \mathcal{D}_T^{-1} \bigg)^{-1} \right] \mathcal{D}_T^{-1}
\end{align}
Thus, the Wald statistic is expressed as below
\begin{align}
\mathcal{W}_T( \pi )
&=
\frac{1}{ \hat{\sigma}_u^2 } \left( \mathcal{D}_T \left( \underline{ \widehat{\beta} }_1 - \underline{ \widehat{\beta} }_2 \right)\right) ^{\prime} \mathcal{D}_T^{-1} \left[ \bigg( \mathcal{D}_T^{-1} \left( X_1^{\prime} X_1 \right) \mathcal{D}_T^{-1} \bigg)^{-1} + \bigg( \mathcal{D}_T^{-1} \left( X_2^{\prime} X_2 \right) \mathcal{D}_T^{-1} \bigg)^{-1} \right]^{-1} \mathcal{D}_T^{-1}
\nonumber
\\
&\ \ \ \ \ \times \mathcal{D}_T \left( \underline{ \widehat{\beta} }_1 - \underline{ \widehat{\beta} }_2 \right)
\end{align}
Now, we can consider the limiting distribution of the Wald OLS statistic by replacing the asymptotic terms for both the distance measure and the covariance matrix in the expression for the test statistic. We obtain the following
\begin{align}
\mathcal{D}_T \left( \underline{ \widehat{\beta} }_1 - \underline{ \widehat{\beta} }_2 \right)
&=
\begin{bmatrix}
1 & \textcolor{red}{0}
\\
\\
\textcolor{red}{0} & \pi V_{zz}
\end{bmatrix}^{-1}
\times
\begin{bmatrix}
B_u (\pi)
\\
\\
\mathcal{N} \bigg( 0, \pi \sigma_u^2 V_{zz} \bigg)
\end{bmatrix}
\nonumber
\\
&-
\begin{bmatrix}
1 & \textcolor{red}{0}
\\
\\
\textcolor{red}{0} & (1 - \pi) V_{zz}
\end{bmatrix}^{-1}
\times
\begin{bmatrix}
B_u (1) - B_u (\pi)
\\
\\
\mathcal{N} \bigg( 0, (1 - \pi) \sigma_u^2 V_{zz} \bigg)
\end{bmatrix}
\end{align}
Note that the off-diagonal elements converge in probability to zero. Thus, we obtain
\begin{align}
\mathcal{D}_T \left( \underline{ \widehat{\beta} }_1 - \underline{ \widehat{\beta} }_2 \right)
&=
\begin{bmatrix}
1 & 0
\\
\\
0 & \frac{1}{ \pi } V_{zz}^{-1}
\end{bmatrix}
\times
\begin{bmatrix}
B_u (\pi)
\\
\\
\mathcal{N} \bigg( 0, \pi \sigma_u^2 V_{zz} \bigg)
\end{bmatrix}
\nonumber
\\
&-
\begin{bmatrix}
1 & 0
\\
\\
0 & \frac{1}{1 - \pi} V_{zz}^{-1}
\end{bmatrix}
\times
\begin{bmatrix}
B_u (1) - B_u (\pi)
\\
\\
\mathcal{N} \bigg( 0, (1 - \pi) \sigma_u^2 V_{zz} \bigg)
\end{bmatrix}
\nonumber
\\
\nonumber
\\
&=
\begin{bmatrix}
B_u (\pi)
\\
\\
\frac{1}{ \pi } V_{zz}^{-1} \mathcal{N} \bigg( 0, \pi \sigma_u^2 V_{zz} \bigg)
\end{bmatrix}
-
\begin{bmatrix}
B_u (1) - B_u (\pi)
\\
\\
\frac{1}{1 - \pi } V_{zz}^{-1} \mathcal{N} \bigg( 0, (1 - \pi) \sigma_u^2 V_{zz} \bigg)
\end{bmatrix}
\nonumber
\\
\nonumber
\\
&=
\begin{bmatrix}
- B_u (1)
\\
\\
\frac{1}{ \pi } V_{zz}^{-1} \mathcal{N} \bigg( 0, \pi \sigma_u^2 V_{zz} \bigg) - \frac{1}{1 - \pi } V_{zz}^{-1} \mathcal{N} \bigg( 0, (1 - \pi) \sigma_u^2 V_{zz} \bigg)
\end{bmatrix}
\end{align}
which implies that
\begin{align}
\mathcal{D}_T \left( \underline{ \widehat{\beta} }_1 - \underline{ \widehat{\beta} }_2 \right)
&=
\begin{bmatrix}
- B_u (1)
\\
\\
\frac{ V_{zz}^{-1} }{ \pi (1 - \pi) } \left\{ \mathcal{N} \bigg( 0, \pi \sigma_u^2 V_{zz} \bigg) -\pi \ \mathcal{N} \bigg( 0, \sigma_u^2 V_{zz} \bigg) \right\}
\end{bmatrix}
\end{align}
The asymptotic convergence of the covariance matrix is given by \
\begin{align}
\left\{
\begin{bmatrix}
1 & 0
\\
\\
0 & \frac{1}{ \pi } V_{zz}^{-1}
\end{bmatrix}
+
\begin{bmatrix}
1 & 0
\\
\\
0 & \frac{1}{1 - \pi } V_{zz}^{-1}
\end{bmatrix}
\right\}^{-1}
&=
\begin{bmatrix}
1 & 0
\\
\\
0 & \frac{1}{\pi( 1 - \pi) } V_{zz}^{-1}
\end{bmatrix}^{-1}
\nonumber
\\
\nonumber
\\
&=
\begin{bmatrix}
1 & 0
\\
\\
0 & \pi( 1 - \pi) V_{zz}
\end{bmatrix}
\end{align}
Therefore, we obtain
\begin{align}
\mathcal{W}_T( \pi )
&\Rightarrow
\frac{1}{ \sigma_u^2 }
\begin{bmatrix}
- B_u (1)
\\
\\
\frac{ V_{zz}^{-1} }{ \pi (1 - \pi) } \left\{ \mathcal{N} \bigg( 0, \pi \sigma_u^2 V_{zz} \bigg) -\pi \ \mathcal{N} \bigg( 0, \sigma_u^2 V_{zz} \bigg) \right\}
\end{bmatrix}^{\prime}
\begin{bmatrix}
1 & 0
\\
\\
0 & \pi( 1 - \pi) V_{zz}
\end{bmatrix}
\nonumber
\\
\nonumber
\\
&\ \ \ \ \ \ \times
\begin{bmatrix}
- B_u (1)
\\
\\
\frac{ V_{zz}^{-1} }{ \pi (1 - \pi) } \left\{ \mathcal{N} \bigg( 0, \pi \sigma_u^2 V_{zz} \bigg) -\pi \ \mathcal{N} \bigg( 0, \sigma_u^2 V_{zz} \bigg) \right\}
\end{bmatrix}
\nonumber
\\
\nonumber
\\
&=
\begin{bmatrix}
- W_u (1) ,
&
V_{zz}^{-1} \left\{ \mathcal{N} \bigg( 0, \pi V_{zz} \bigg) -\pi \ \mathcal{N} \bigg( 0, V_{zz} \bigg) \right\}
\end{bmatrix}
\begin{bmatrix}
- W_u (1)
\\
\frac{ V_{zz}^{-1} }{ \pi (1 - \pi) } \left\{ \mathcal{N} \bigg( 0, \pi V_{zz} \bigg) -\pi \ \mathcal{N} \bigg( 0, V_{zz} \bigg) \right\}
\end{bmatrix}
\nonumber
\\
\nonumber
\\
&= \bigg( W_u (1) \bigg)^2 + \underset{ \pi \in [ \pi_1 , \pi_2 ] }{ \text{ sup } } \ \frac{ \big[ W(\pi) - \pi W(1) \big]^{\prime} \big[ W(\pi) - \pi W(1) \big] }{ \pi (1 - \pi) }
\nonumber
\\
\nonumber
\\
&:= \chi^2_p + \underset{ \pi \in [ \pi_1 , \pi_2 ] }{ \text{ sup } } \ \ \frac{ \mathcal{B B} (\pi) }{ \pi (1 - \pi) }
\end{align}
Asymptotic Distribution of sup Wald-IVX statistic
Consider the univariate predictive regression with multiple predictors
align[align omitted — 128 chars of source]
and $x_t$ is generated via a LUR process as below
align[align omitted — 100 chars of source]
where $I_{1t} := \mathbf{1} \{ t \leq k \}$ and $I_{2t} := \mathbf{1} \{ t > k \}$ with $k = \floor{ T\pi}$.
Denote with $X_1 \in \mathbb{R}^{T \times (p+1)}$ to represent the corresponding matrix stacking $I_{1t}$, i.e., $x_t I_{1t} \equiv x_{1t}$ and similarly $X_2 \in \mathbb{R}^{T \times (p+1)}$ represents the matrix stacking $I_{2t}$, i.e, $x_t I_{2t} \equiv x_{2t}$ including a column with one's to capture the model intercept in both regimes. We consider the following two normalization matrices, that is, $\mathcal{D}_1 = \text{diag}\left( T^{ \frac{\delta}{2} }, T^{ \frac{ 1 + \delta }{ 2 } } \text{I}_p \right)$ and $\mathcal{D}_2 = \text{diag}\left( T^{ 1 - \frac{\delta}{2} }, T^{ \frac{ 1 + \delta }{ 2 } } \text{I}_p \right)$, where $\text{I}_p$ the $( p \times p)$ identity matrix.
The IVX instrumentation implies that
align[align omitted — 160 chars of source]
Note that all matrices, $\left\{ X_1, X_2, Z_1, Z_2 \right\} \in \mathbb{R}^{ T \times (p+1)}$, include the first column to be a column vector of ones with the remaining columns to represent the corresponding stacked values from the set of $p-$predictors which are included in the model. Furthermore, for simplicity of notation, we consider that $Z_j$ for $j = 1,2$ represents the corresponding IVX instruments, constructed via the IVX instrumentation procedure of PM. Moreover, we operate under the assumption of an unknown break-point $\pi \in \Pi$.
The Wald IVX statistic has the following form
align*[align* omitted — 365 chars of source]
where the covariance matrix $\mathcal{Q}_{ \mathcal{R} }$ is defined as below
align[align omitted — 291 chars of source]
We denote with $\theta_i = ( \alpha_i, \beta_i )$ for $i = 1,2$, the IVX estimator which is expressed as below
align*[align* omitted — 285 chars of source]
Thus,
align*[align* omitted — 212 chars of source]
Similarly,
align*[align* omitted — 280 chars of source]
Thus,
align*[align* omitted — 212 chars of source]
Therefore,
align*[align* omitted — 382 chars of source]
and/or equivalently,
align*[align* omitted — 406 chars of source]
Furthermore, for the covariance matrix we have that
align*[align* omitted — 893 chars of source]
Therefore,
align*[align* omitted — 382 chars of source]
which gives
align[align omitted — 283 chars of source]
\color{black}
Wald IVX test for mildly integrated regressors
Consider the univariate predictive regression with multiple predictors
align*[align* omitted — 129 chars of source]
and $x_t$ is generated via the following process
align*[align* omitted — 112 chars of source]
where $I_{1t} := \mathbf{1} \{ t \leq k \}$ and $I_{2t} := \mathbf{1} \{ t > k \}$ with $k = \floor{ T\pi}$.
Proof. We consider the weakly convergence of the following sample moments
align[align omitted — 1,060 chars of source]
align[align omitted — 484 chars of source]
align[align omitted — 517 chars of source]
Furthermore, for each estimator we have that
align[align omitted — 1,616 chars of source]
\color{blue}
We have the following formula for the inverse of a partitioned matrix
align*[align* omitted — 280 chars of source]
where
align*[align* omitted — 82 chars of source]
\color{black}
Thus, for the inversion of $\mathcal{A}_1$ we have that
align[align omitted — 364 chars of source]
since
align[align omitted — 110 chars of source]
Similarly, for the inversion of $\mathcal{A}_2$ we have that
align[align omitted — 466 chars of source]
Furthermore, for each estimator we have that
align[align omitted — 1,584 chars of source]
landscape\begin{align*}
\mathcal{D}_1 \mathcal{Q}_{ \mathcal{R} } \mathcal{D}_2
&=
\bigg[ \mathcal{D}_2^{-1} \left(Z_1^{\prime} X_1 \right) \mathcal{D}_1^{-1} \bigg]^{-1} \bigg[ \mathcal{D}_2^{-1} \left( Z_1^{\prime} Z_1 \right) \mathcal{D}_1^{-1} \bigg] \bigg[ \mathcal{D}_2^{-1} \left(Z_1^{\prime} X_1 \right) \mathcal{D}_1^{-1} \bigg]^{-1}
\nonumber
\\
&\ +
\bigg[ \mathcal{D}_2^{-1} \left(Z_2^{\prime} X_2 \right) \mathcal{D}_1^{-1} \bigg]^{-1} \bigg[ \mathcal{D}_2^{-1} \left( Z_2^{\prime} Z_2 \right) \mathcal{D}_1^{-1} \bigg] \bigg[ \mathcal{D}_2^{-1} \left(Z_2^{\prime} X_2 \right) \mathcal{D}_1^{-1} \bigg]^{-1}
\\
\\
&=
\begin{bmatrix}
\frac{1}{\pi_0} & \ \ 0
\\
\\
\frac{1}{\pi_0^2} \Omega_{vv}^{-1} J_c( \pi_0 ) & \ \ - \frac{1}{\pi_0} \Omega_{vv}^{-1}
\end{bmatrix}
\times
\begin{bmatrix}
\pi_0 & \ \ 0
\\
\\
J_{c} ( \pi_0 ) & \pi_0 V_{zz}
\end{bmatrix}
\times
\begin{bmatrix}
\frac{1}{\pi_0} & \ \ 0
\\
\\
\frac{1}{\pi_0^2} \Omega_{vv}^{-1} J_c( \pi_0 ) & \ \ - \frac{1}{\pi_0} \Omega_{vv}^{-1}
\end{bmatrix}
\\
\\
&+
\begin{bmatrix}
\frac{1}{1 - \pi_0} & \ \ 0
\\
\\
\frac{1}{(1 - \pi_0)^2} \Omega_{vv}^{-1} \bigg( J_c( 1 ) - J_c( \pi_0 ) \bigg) & \ \ - \frac{1}{1 - \pi_0} \Omega_{vv}^{-1}
\end{bmatrix}
\times
\begin{bmatrix}
1 - \pi_0 & \ \ 0
\\
\\
J_{c} ( 1 ) - \underline{J}_{c} ( \pi_0 ) & (1 - \pi_0) V_{zz}
\end{bmatrix}
\times
\\
&\times
\begin{bmatrix}
\frac{1}{1 - \pi_0} & \ \ 0
\\
\\
\frac{1}{(1 - \pi_0)^2} \Omega_{vv}^{-1} \bigg( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \bigg) & \ \ - \frac{1}{1 - \pi_0} \Omega_{vv}^{-1}
\end{bmatrix}
\end{align*}
\begin{align}
\Phi_1
&:=
\begin{bmatrix}
\frac{1}{\pi_0} & \ \ 0
\\
\\
\frac{1}{\pi_0^2} \Omega_{vv}^{-1} \underline{J}_c( \pi_0 ) & \ \ - \frac{1}{\pi_0} \Omega_{vv}^{-1}
\end{bmatrix}
\times
\begin{bmatrix}
\pi_0 & \ \ 0
\\
\\
\underline{J}_{c} ( \pi_0 ) & \pi_0 V_{zz}
\end{bmatrix}
\times
\begin{bmatrix}
\frac{1}{\pi_0} & \ \ 0
\\
\\
\frac{1}{\pi_0^2} \Omega_{vv}^{-1} \underline{J}_c( \pi_0 ) & \ \ - \frac{1}{\pi_0} \Omega_{vv}^{-1}
\end{bmatrix}
\nonumber
\\
\nonumber
\\
&=
\begin{bmatrix}
1 & \ \ 0
\\
\\
\frac{1}{\pi_0} \Omega_{vv}^{-1} \underline{J}_c( \pi_0 ) - \frac{1}{\pi_0} \Omega_{vv}^{-1} \underline{J}_c( \pi_0 ) & \ \ - \Omega_{vv}^{-1} V_{zz}
\end{bmatrix}
\times
\begin{bmatrix}
\frac{1}{\pi_0} & \ \ 0
\\
\\
\frac{1}{\pi_0^2} \Omega_{vv}^{-1} \underline{J}_c( \pi_0 ) & \ \ - \frac{1}{\pi_0} \Omega_{vv}^{-1}
\end{bmatrix}
\nonumber
\\
\nonumber
\\
&=
\begin{bmatrix}
\frac{1}{\pi_0} & \ \ 0
\\
\\
\left( \Phi_1 \right)_{21} & \ \ \frac{1}{\pi_0} \Omega_{vv}^{-1} V_{zz} \Omega_{vv}^{-1}
\end{bmatrix}
\end{align}
where
\begin{align}
\left( \Phi_1 \right)_{21} = - \frac{1}{\pi_0^2} \Omega_{vv}^{-1}V_{zz} \Omega_{vv}^{-1} \underline{J}_c( \pi_0 )
\end{align}
\begin{align*}
\Phi_2
&=
\begin{bmatrix}
\frac{1}{1 - \pi_0} & \ 0
\\
\\
\frac{1}{(1 - \pi_0)^2} \Omega_{vv}^{-1} \bigg( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \bigg) & - \frac{1}{1 - \pi_0} \Omega_{vv}^{-1}
\end{bmatrix}
\begin{bmatrix}
1 - \pi_0 & \ 0
\\
\\
\underline{J}_{c} ( 1 ) - \underline{J}_{c} ( \pi_0 ) & (1 - \pi_0) V_{zz}
\end{bmatrix}
\begin{bmatrix}
\frac{1}{1 - \pi_0} & \ \ 0
\\
\\
\frac{1}{(1 - \pi_0)^2} \Omega_{vv}^{-1} \bigg( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \bigg) & - \frac{1}{1 - \pi_0} \Omega_{vv}^{-1}
\end{bmatrix}
\\
\nonumber
\\
&=
\begin{bmatrix}
1 & \ 0
\\
\\
\frac{1}{1 - \pi_0} \Omega_{vv}^{-1} \bigg( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \bigg) - \frac{1}{1 - \pi_0} \Omega_{vv}^{-1} \bigg( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \bigg) & - \Omega_{vv}^{-1} V_{zz}
\end{bmatrix}
\times
\begin{bmatrix}
\frac{1}{1 - \pi_0} & \ \ 0
\\
\\
\frac{1}{(1 - \pi_0)^2} \Omega_{vv}^{-1} \bigg( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \bigg) & - \frac{1}{1 - \pi_0} \Omega_{vv}^{-1}
\end{bmatrix}
\\
\nonumber
\\
&=
\begin{bmatrix}
\frac{1}{1 - \pi_0} & \ 0
\\
\\
\left( \Phi_2 \right)_{21} & \frac{1}{1 - \pi_0} \Omega_{vv}^{-1} V_{zz} \Omega_{vv}^{-1}
\end{bmatrix}
\end{align*}
where
\begin{align}
\left( \Phi_2 \right)_{21} = - \frac{1}{(1 - \pi_0)^2} \Omega_{vv}^{-1} V_{zz} \Omega_{vv}^{-1} \bigg( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \bigg)
\end{align}
Thus, we have that
\begin{align}
\mathcal{D}_1 \mathcal{Q}_{ \mathcal{R} } \mathcal{D}_2 = \Phi_1 + \Phi_2
&=
\begin{bmatrix}
\frac{1}{\pi_0} \ \ & \ \ 0
\\
\\
\left( \Phi_1 \right)_{21} \ \ & \ \ \frac{1}{\pi_0} \Omega_{vv}^{-1} V_{zz} \Omega_{vv}^{-1}
\end{bmatrix}
+
\begin{bmatrix}
\frac{1}{1 - \pi_0} \ \ & \ \ 0
\\
\\
\left( \Phi_2 \right)_{21} \ \ & \ \ \frac{1}{1 - \pi_0} \Omega_{vv}^{-1} V_{zz} \Omega_{vv}^{-1}
\end{bmatrix}
\nonumber
\\
\nonumber
\\
&=
\begin{bmatrix}
\frac{1}{\pi_0 (1 - \pi_0)} \ \ & \ \ 0
\\
\\
\left( \Phi_1 \right)_{21} + \left( \Phi_2 \right)_{21} \ \ & \ \ \frac{1}{\pi_0 (1 - \pi_0)} \Omega_{vv}^{-1} V_{zz} \Omega_{vv}^{-1}
\end{bmatrix}
\end{align}
where
\begin{align}
\left( \Phi_1 \right)_{21} + \left( \Phi_2 \right)_{21}
&=
- \frac{1}{\pi_0^2} \Omega_{vv}^{-1}V_{zz} \Omega_{vv}^{-1} \underline{J}_c( \pi_0 )
- \frac{1}{(1 - \pi_0)^2} \Omega_{vv}^{-1} V_{zz} \Omega_{vv}^{-1} \bigg( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \bigg)
\nonumber
\\
\nonumber
\\
&=
- \frac{1}{\pi_0^2} \Omega_{vv}^{-1}V_{zz} \Omega_{vv}^{-1} \underline{J}_c( \pi_0 )
- \frac{1}{(1 - \pi_0)^2} \Omega_{vv}^{-1} V_{zz} \Omega_{vv}^{-1} \underline{J}_c( 1 )
+ \frac{1}{(1 - \pi_0)^2} \Omega_{vv}^{-1} V_{zz} \Omega_{vv}^{-1} \underline{J}_c( \pi_0 )
\nonumber
\\
\nonumber
\\
&=
\frac{1 - 2 \pi_0 }{\pi_0^2(1 - \pi_0)^2 } \Omega_{vv}^{-1}V_{zz} \Omega_{vv}^{-1} \underline{J}_c( \pi_0 )
- \frac{\pi_0^2}{\pi_0^2(1 - \pi_0)^2} \Omega_{vv}^{-1} V_{zz} \Omega_{vv}^{-1} \underline{J}_c( 1 )
\nonumber
\\
\nonumber
\\
&=
\frac{1}{\pi_0^2(1 - \pi_0)^2 } \Omega_{vv}^{-1}V_{zz} \Omega_{vv}^{-1} \bigg\{ \left( 1 - 2 \pi_0 \right) \underline{J}_c( \pi_0 ) - \pi_0^2 \underline{J}_c( 1 ) \bigg\} := \Delta
\end{align}
Furthermore, we have that
\begin{align}
\bigg[ \mathcal{D}_1 \mathcal{Q}_{ \mathcal{R} } \mathcal{D}_2 \bigg]^{-1}
&\equiv
\begin{bmatrix}
\frac{1}{\pi_0 (1 - \pi_0)} \ \ & \ \ 0
\\
\\
\Delta \ \ & \ \ \frac{1}{\pi_0 (1 - \pi_0)} \Omega_{vv}^{-1} V_{zz} \Omega_{vv}^{-1}
\end{bmatrix}^{-1}
\nonumber
\\
\nonumber
\\
&=
\begin{bmatrix}
\pi_0 (1 - \pi_0) \ \ & \ \ 0
\\
\\
K \ \ & \ \ \pi_0 (1 - \pi_0) \Omega_{vv} V_{zz}^{-1} \Omega_{vv}
\end{bmatrix}
\end{align}
where
\begin{align*}
K
&=
- \pi_0^2 (1 - \pi_0)^2 \Omega_{vv} V_{zz}^{-1} \Omega_{vv} \times \Delta
\nonumber
\\
\nonumber
\\
&= - \pi_0^2 (1 - \pi_0)^2 \Omega_{vv} V_{zz}^{-1} \Omega_{vv} \times \frac{1}{\pi_0^2(1 - \pi_0)^2 } \Omega_{vv}^{-1}V_{zz} \Omega_{vv}^{-1} \bigg\{ \left( 1 - 2 \pi_0 \right) \underline{J}_c( \pi_0 ) - \pi_0^2 \underline{J}_c( 1 ) \bigg\}
\nonumber
\\
\nonumber
\\
&=
- \bigg\{ \left( 1 - 2 \pi_0 \right) \underline{J}_c( \pi_0 ) - \pi_0^2 \underline{J}_c( 1 ) \bigg\}
\end{align*}
\color{black}
landscapeAlso, the statistical distance measure is given by
\begin{align}
\mathcal{D}_1 \left( \widehat{\theta}_1 - \widehat{\theta}_2 \right)
=
\begin{bmatrix}
0
\\
\\
- \Omega_{vv}^{-1} \left\{ \frac{1}{\pi_0} B( \pi_0 ) - \frac{1}{1 - \pi_0} \bigg( B( 1 ) - B( \pi_0 ) \bigg) \right\}
\end{bmatrix}
=
\begin{bmatrix}
0
\\
\\
- \frac{1}{\pi_0( 1 - \pi_0)} \Omega_{vv}^{-1} \bigg\{ B( \pi_0) - \pi_0 B( 1 ) \bigg\}
\end{bmatrix}
\end{align}
Thus, the Wald IVX statistic for the case of mildly integrated regressors becomes
\begin{align}
\mathcal{W}_T( \pi)
&\Rightarrow
\frac{1}{\sigma_u^2}
\color{blue}
\begin{bmatrix}
a_1
\\
\\
a_2
\end{bmatrix}^{\prime}
\color{black}
\begin{bmatrix}
\pi_0 (1 - \pi_0) & \ 0
\\
\\
\textcolor{red}{K} & \pi_0 (1 - \pi_0) \Omega_{vv} V_{zz}^{-1} \Omega_{vv}
\end{bmatrix}
\begin{bmatrix}
0
\\
\\
- \frac{1}{\pi_0( 1 - \pi_0)} \Omega_{vv}^{-1} \bigg\{ B( \pi_0) - \pi_0 \underline{B}( 1 ) \bigg\}
\end{bmatrix}
\end{align}
\color{red}
\underline{Note:} See next Section, for the asymptotic convergence of the blue vector above (since it has a different normalization matrix).
\color{black}
Statistical Distance measure with $\mathcal{D}_2$ normalization matrix
We have that
align*[align* omitted — 416 chars of source]
Therefore,
align*[align* omitted — 382 chars of source]
Then, we consider the weakly convergence of the following sample moments
align[align omitted — 1,047 chars of source]
align[align omitted — 470 chars of source]
align[align omitted — 503 chars of source]
Furthermore, for each estimator we have that
align[align omitted — 1,584 chars of source]
\color{blue}
We have the following formula for the inverse of a partitioned matrix
align*[align* omitted — 280 chars of source]
where
align*[align* omitted — 82 chars of source]
\color{black}
Thus, for the inversion of $\mathcal{A}_1$ we have that
align[align omitted — 365 chars of source]
since
align[align omitted — 110 chars of source]
Similarly, for the inversion of $\mathcal{A}_2$ we have that
align[align omitted — 466 chars of source]
Furthermore, for each estimator we have that
align[align omitted — 1,856 chars of source]
landscapeThus, the statistical distance measure is given by
\begin{align*}
\color{blue}
\mathcal{D}_2 \left( \widehat{\theta}_1 - \widehat{\theta}_2 \right)
\color{black}
&=
\begin{bmatrix}
\frac{1}{\pi_0^2} \Omega_{vv}^{-1} J_c( \pi_0 ) B( \pi_0 )
\\
\\
- \frac{1}{\pi_0} \Omega_{vv}^{-1} B( \pi_0 )
\end{bmatrix}
-
\begin{bmatrix}
\frac{1}{(1 - \pi_0)^2} \Omega_{vv}^{-1} \bigg( J_c( 1 ) - J_c( \pi_0 ) \bigg)\bigg( B( 1 ) - \underline{B}( \pi_0 ) \bigg)
\\
\\
- \frac{1}{1 - \pi_0} \Omega_{vv}^{-1} \bigg( \underline{B}( 1 ) - \underline{B}( \pi_0 ) \bigg)
\end{bmatrix}
\\
\\
&=
\begin{bmatrix}
\frac{1}{\pi_0^2} \Omega_{vv}^{-1} \underline{J}_c( \pi_0 ) \underline{B}( \pi_0 ) - \frac{1}{(1 - \pi_0)^2} \Omega_{vv}^{-1} \bigg( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \bigg)\bigg( \underline{B}( 1 ) - \underline{B}( \pi_0 ) \bigg)
\\
\\
- \frac{1}{\pi_0} \Omega_{vv}^{-1} \underline{B}( \pi_0 ) + \frac{1}{1 - \pi_0} \Omega_{vv}^{-1} \bigg( \underline{B}( 1 ) - \underline{B}( \pi_0 ) \bigg)
\end{bmatrix}
\end{align*}
Therefore, the Wald IVX statistic for the case of mildly integrated regressors becomes
\begin{align}
\mathcal{W}_T( \pi)
&\Rightarrow
\underset{ \pi \in [ \pi_1, \pi_2 ] }{ { \text{sup} } } \
\frac{1}{\sigma_u^2}
\color{blue}
\bigg[ \mathcal{D}_2 \left( \widehat{\theta}_1 - \widehat{\theta}_2 \right) \bigg]^{\prime}
\color{black}
\begin{bmatrix}
\pi_0 (1 - \pi_0) & \ 0
\\
\\
\textcolor{red}{K} & \pi_0 (1 - \pi_0) \Omega_{vv} V_{zz}^{-1} \Omega_{vv}
\end{bmatrix}
\begin{bmatrix}
0
\\
\\
- \frac{1}{\pi_0( 1 - \pi_0)} \Omega_{vv}^{-1} \bigg\{ \underline{B}( \pi_0) - \pi_0 \underline{B}( 1 ) \bigg\}
\end{bmatrix}
\nonumber
\\
&=
\underset{ \pi \in [ \pi_1, \pi_2 ] }{ { \text{sup} } } \
\frac{1}{\sigma_u^2}
\begin{bmatrix}
\textcolor{blue}{a_1}
\\
\\
- V_{zz}^{-1} \Omega_{vv} \bigg( \underline{B}( \pi_0 ) - \pi_0 \underline{B}( 1 ) \bigg)
\end{bmatrix}^{\prime}
\begin{bmatrix}
0
\\
\\
- \frac{1}{\pi_0( 1 - \pi_0)} \Omega_{vv}^{-1} \bigg\{ \underline{B}( \pi_0) - \pi_0 \underline{B}( 1 ) \bigg\}
\end{bmatrix}
\nonumber
\\
\nonumber
\\
&=
\underset{ \pi \in [ \pi_1, \pi_2 ] }{ { \text{sup} } } \ \frac{ \bigg[ \underline{W}( \pi) - \pi \underline{W}( 1 ) \bigg]^{\prime} \bigg[ \underline{W}( \pi) - \pi \underline{W}( 1 ) \bigg] }{\pi (1 - \pi)}
\end{align}
where the elements for the first vector are as below:
\begin{align}
\textcolor{blue}{a_1}
&=
\frac{1 - \pi_0}{\pi_0} \Omega_{vv}^{-1} \underline{J}_c( \pi_0 ) \underline{B}( \pi_0 ) - \frac{1}{(1 - \pi_0)} \Omega_{vv}^{-1} \bigg( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \bigg)\bigg( \underline{B}( 1 ) - \underline{B}( \pi_0 ) \bigg)
\nonumber
\\
\nonumber
\\
&+ \left\{ \frac{1}{\pi_0} \Omega_{vv}^{-1} \underline{B}( \pi_0 ) + \frac{1}{1 - \pi_0} \Omega_{vv}^{-1} \bigg( \underline{B}( 1 ) - \underline{B}( \pi_0 ) \bigg) \right\}
\bigg\{ \left( 1 - 2 \pi_0 \right) \underline{J}_c( \pi_0 ) - \pi_0^2 \underline{J}_c( 1 ) \bigg\}
\end{align}
\begin{align}
a_2
&=
\left\{ - \frac{1}{\pi_0} \Omega_{vv}^{-1} \underline{B}( \pi_0 ) + \frac{1}{1 - \pi_0} \Omega_{vv}^{-1} \bigg( \underline{B}( 1 ) - \underline{B}( \pi_0 ) \bigg) \right\} \bigg\{ \pi_0 (1 - \pi_0) \Omega_{vv} V_{zz}^{-1} \Omega_{vv} \bigg\}
\nonumber
\\
\nonumber
\\
&- (1 - \pi_0) \Omega_{vv}^{-1} \Omega_{vv} V_{zz}^{-1} \Omega_{vv} \underline{B}( \pi_0 ) + \pi_0 \Omega_{vv}^{-1} \Omega_{vv} V_{zz}^{-1} \Omega_{vv} \underline{B}( 1 ) - \pi_0 \Omega_{vv}^{-1} \Omega_{vv} V_{zz}^{-1} \Omega_{vv} \underline{B}( \pi_0 )
\nonumber
\\
\nonumber
\\
&= - V_{zz}^{-1} \Omega_{vv} \bigg( \underline{B}( \pi_0 ) - \pi_0 \underline{B}( 1 ) \bigg)
\end{align}
Wald IVX test for LUR regressors
We have the following equivalent expression for the Wald IVX statistic
align*[align* omitted — 271 chars of source]
To determine the limiting distribution of the Wald-IVX statistic we consider the weakly convergence of each of its terms separately. To simplify the algebra we denote with $\displaystyle Q_{xz} := - \left( \int_{ 0 }^1 \underline{J}_c (r) d \underline{J}_v + \Omega_{vv} \right) C_z^{-1}$. Regardless whether the break-point is known or unknown the following holds
align[align omitted — 280 chars of source]
and
align[align omitted — 284 chars of source]
We consider the weakly convergence of the following sample moments
align[align omitted — 1,260 chars of source]
align[align omitted — 484 chars of source]
align[align omitted — 517 chars of source]
Furthermore, for each estimator we have that
align[align omitted — 1,830 chars of source]
Note that, for the asymptotic converges of the terms $\frac{1}{ T^{ \frac{1}{2} + \delta}} \sum_{t=1}^T z_{1t}$ and $\frac{1}{ T^{ \frac{1}{2} + \delta}} \sum_{t=1}^T z_{2t}$, we use the result given by Lemma B1 (i) in the Appendix of KMS. That, is since we have that $\gamma = 1$ in the case of persistent regressors and we assume that the exponent rate of the degree of persistence of the IVX instrument is $\delta \in (0,1)$, then
align[align omitted — 141 chars of source]
Thus, we have that $\frac{1}{ T^{ \frac{1}{2} + \delta}} \sum_{t=1}^{ \floor{T\pi }} \tilde{z}_{t} \Rightarrow \underline{J}_c( \pi_0 )$.
Therefore, we consider the inverse of the partition of the following matrices as below
align*[align* omitted — 631 chars of source]
Useful Notation:
We use the following matrix notation to simplify further the expression for the sup Wald-IVX statistic
align[align omitted — 165 chars of source]
with
equation[equation omitted — 288 chars of source]
\color{blue}
We have the following formula for the inverse of a partitioned matrix
align*[align* omitted — 353 chars of source]
\color{black}
Equivalently, we denote with
align[align omitted — 310 chars of source]
Moreover, we denote with
align[align omitted — 310 chars of source]
Thus, for the inversion of $\mathcal{A}_1$ we have that
align[align omitted — 187 chars of source]
align*[align* omitted — 370 chars of source]
align[align omitted — 460 chars of source]
Similarly, for $\mathcal{A}_2^{-1}$ we have that
align*[align* omitted — 305 chars of source]
Therefore, we obtain an expression for $\mathcal{A}_2^{-1}$ as below
align[align omitted — 544 chars of source]
landscape\begin{align*}
\mathcal{D}_1 \mathcal{Q}_{ \mathcal{R} } \mathcal{D}_2
&=
\bigg[ \mathcal{D}_2^{-1} \left(Z_1^{\prime} X_1 \right) \mathcal{D}_1^{-1} \bigg]^{-1} \bigg[ \mathcal{D}_2^{-1} \left( Z_1^{\prime} Z_1 \right) \mathcal{D}_1^{-1} \bigg] \bigg[ \mathcal{D}_2^{-1} \left(Z_1^{\prime} X_1 \right) \mathcal{D}_1^{-1} \bigg]^{-1}
\nonumber
\\
&\ +
\bigg[ \mathcal{D}_2^{-1} \left(Z_2^{\prime} X_2 \right) \mathcal{D}_1^{-1} \bigg]^{-1} \bigg[ \mathcal{D}_2^{-1} \left( Z_2^{\prime} Z_2 \right) \mathcal{D}_1^{-1} \bigg] \bigg[ \mathcal{D}_2^{-1} \left(Z_2^{\prime} X_2 \right) \mathcal{D}_1^{-1} \bigg]^{-1}
\\
&=
\begin{bmatrix}
\left( \pi_0 - \left( \int_0^{\pi_0} J^{\top} (r) dr \right) \Delta_1^{-1} J_c( \pi_0 ) \right)^{-1} \ \ & \ \ - \frac{1}{\pi_0} \left( \int_0^{\pi_0} J^{\top} (r) dr \right) \mathcal{S}_1^{-1}
\\
\\
- \mathcal{S}_1^{-1} \frac{ J_c( \pi_0 )}{ \pi_0 } \ \ & \ \ \mathcal{S}_1^{-1}
\end{bmatrix}
\times
\begin{bmatrix}
\pi_0 & \ \ 0
\\
\\
J_{c} ( \pi_0 ) & \pi_0 V_{zz}
\end{bmatrix}
\\
& \times
\begin{bmatrix}
\left( \pi_0 - \left( \int_0^{\pi_0} J^{\top} (r) dr \right) \Delta_1^{-1} \underline{J}_c( \pi_0 ) \right)^{-1} \ \ & \ \ - \frac{1}{\pi_0} \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1}
\\
\\
- \mathcal{S}_1^{-1} \frac{ \underline{J}_c( \pi_0 )}{ \pi_0 } \ \ & \ \ \mathcal{S}_1^{-1}
\end{bmatrix}
\\
\\
&+
\begin{bmatrix}
\left( (1 - \pi_0) - \left( \int_{\pi_0}^{1} \underline{J}^{\top} (r) dr \right) \Delta_2^{-1} \big( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \big) \right)^{-1} & - \frac{1}{1 - \pi_0} \left( \int_{\pi_0}^{1} \underline{J}^{\top} (r) dr \right) \mathcal{S}_2^{-1}
\\
\\
- \mathcal{S}_2^{-1} \frac{ \big( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \big)}{\pi_0 (1 - \pi_0)} \ \ & \ \ \mathcal{S}_2^{-1}
\end{bmatrix}
\begin{bmatrix}
1 - \pi_0 & \ \ 0
\\
\\
\underline{J}_{c} ( 1 ) - \underline{J}_{c} ( \pi_0 ) & (1 - \pi_0) V_{zz}
\end{bmatrix}
\times
\\
&\times
\begin{bmatrix}
\left( (1 - \pi_0) - \left( \int_{\pi_0}^{1} \underline{J}^{\top} (r) dr \right) \Delta_2^{-1} \big( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \big) \right)^{-1} & - \frac{1}{1 - \pi_0} \left( \int_{\pi_0}^{1} \underline{J}^{\top} (r) dr \right) \mathcal{S}_2^{-1}
\\
\\
- \mathcal{S}_2^{-1} \frac{ \big( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \big)}{\pi_0 (1 - \pi_0)} \ \ & \ \ \mathcal{S}_2^{-1}
\end{bmatrix}
\end{align*}
\begin{align}
\Phi_1 &:=
\begin{bmatrix}
\left( \pi_0 - \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \Delta_1^{-1} \underline{J}_c( \pi_0 ) \right)^{-1} \ \ & \ \ - \frac{1}{\pi_0} \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1}
\\
\\
- \mathcal{S}_1^{-1} \frac{ \underline{J}_c( \pi_0 )}{ \pi_0 } \ \ & \ \ \mathcal{S}_1^{-1}
\end{bmatrix}_{ \textcolor{red}{ 2 \times 2 } }
\times
\begin{bmatrix}
\pi_0 & \ \ 0
\\
\\
\underline{J}_{c} ( \pi_0 ) & \pi_0 V_{zz}
\end{bmatrix}_{ \textcolor{red}{ 2 \times 2 } }
\\
&\ \times
\begin{bmatrix}
\left( \pi_0 - \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \Delta_1^{-1} \underline{J}_c( \pi_0 ) \right)^{-1} \ \ & \ \ - \frac{1}{\pi_0} \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1}
\\
\\
- \mathcal{S}_1^{-1} \frac{ \underline{J}_c( \pi_0 )}{ \pi_0 } \ \ & \ \ \mathcal{S}_1^{-1}
\end{bmatrix}_{ \textcolor{red}{ 2 \times 2 } }
\nonumber
\\
\nonumber
\\
&=
\begin{bmatrix}
\pi_0 \left( \pi_0 - \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \Delta_1^{-1} \underline{J}_c( \pi_0 ) \right)^{-1} - \frac{1}{\pi_0} \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1} \underline{J}_{c} ( \pi_0 ) \ \ & \ \ - \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1} V_{zz}
\\
\\
- \mathcal{S}_1^{-1} \underline{J}_c( \pi_0 ) + \mathcal{S}_1^{-1} \underline{J}_c( \pi_0 )\ \textcolor{red}{= 0} \ \ & \ \ \pi_0 \mathcal{S}_1^{-1} V_{zz}
\end{bmatrix}_{ \textcolor{red}{ 2 \times 2 } }
\nonumber
\\
\nonumber
\\
& \times
\begin{bmatrix}
\left( \pi_0 - \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \Delta_1^{-1} \underline{J}_c( \pi_0 ) \right)^{-1} \ \ & \ \ - \frac{1}{\pi_0} \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1}
\\
\\
- \mathcal{S}_1^{-1} \frac{ \underline{J}_c( \pi_0 )}{ \pi_0 } \ \ & \ \ \mathcal{S}_1^{-1}
\end{bmatrix}_{ \textcolor{red}{ 2 \times 2 } }
\nonumber
\\
\nonumber
\\
&=
\begin{bmatrix}
\alpha_1 & \ \ \alpha_2
\\
\\
\alpha_3 & \ \ \alpha_4
\end{bmatrix}_{ \textcolor{red}{ 2 \times 2 } }
\end{align}
where
\begin{align*}
\alpha_1
&=
\left\{ \pi_0 \left( \pi_0 - \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \Delta_1^{-1} \underline{J}_c( \pi_0 ) \right)^{-1} - \frac{1}{\pi_0} \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1} \underline{J}_{c} ( \pi_0 ) \right\} \left( \pi_0 - \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \Delta_1^{-1} \underline{J}_c( \pi_0 ) \right)^{-1}
\\
&+ \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1} V_{zz} \mathcal{S}_1^{-1} \frac{ \underline{J}_c( \pi_0 )}{ \pi_0 }
\\
&=
\pi_0 \left( \pi_0 - \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \Delta_1^{-1} \underline{J}_c( \pi_0 ) \right)^{-2}
-
\frac{1}{\pi_0} \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1} \underline{J}_{c} ( \pi_0 ) \left( \pi_0 - \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \Delta_1^{-1} \underline{J}_c( \pi_0 ) \right)^{-1}
\\
&+ \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1} V_{zz} \mathcal{S}_1^{-1} \frac{ \underline{J}_c( \pi_0 )}{ \pi_0 }
\end{align*}
\begin{align*}
\alpha_2
&=
\left\{ \pi_0 \left( \pi_0 - \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \Delta_1^{-1} \underline{J}_c( \pi_0 ) \right)^{-1} - \frac{1}{\pi_0} \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1} \underline{J}_{c} ( \pi_0 ) \right\} \left\{ - \frac{1}{\pi_0} \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}^{-1} \right\}
\\
&- \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1} V_{zz} \mathcal{S}_1^{-1}
\\
&=
- \left( \pi_0 - \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \Delta_1^{-1} \underline{J}_c( \pi_0 ) \right)^{-1} \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1}
+ \frac{1}{\pi_0^2} \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1} \underline{J}_{c} ( \pi_0 ) \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}^{-1}
\\
&- \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_1^{-1} V_{zz} \mathcal{S}_1^{-1}
\end{align*}
\begin{align*}
\alpha_3 &= - S_1^{-1} V_{zz} S_1^{-1} \underline{J}_c ( \pi_0 )
\\
\\
\alpha_4 &= \pi_0 S_1^{-1} V_{zz} S_1^{-1}
\end{align*}
\begin{align*}
\Phi_2
&:=
\begin{bmatrix}
\left( (1 - \pi_0) - \left( \int_{\pi_0}^{1} \underline{J}^{\top} (r) dr \right) \Delta_2^{-1} \big( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \big) \right)^{-1} & - \frac{1}{1 - \pi_0} \left( \int_{\pi_0}^{1} \underline{J}^{\top} (r) dr \right) \mathcal{S}_2^{-1}
\\
\\
- \mathcal{S}_2^{-1} \frac{ \big( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \big)}{\pi_0 (1 - \pi_0)} \ \ & \ \ \mathcal{S}_2^{-1}
\end{bmatrix}
\begin{bmatrix}
1 - \pi_0 & \ \ 0
\\
\\
\underline{J}_{c} ( 1 ) - \underline{J}_{c} ( \pi_0 ) & (1 - \pi_0) V_{zz}
\end{bmatrix}
\\
&\ \times
\begin{bmatrix}
\left( (1 - \pi_0) - \left( \int_{\pi_0}^{1} \underline{J}^{\top} (r) dr \right) \Delta_2^{-1} \big( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \big) \right)^{-1} & - \frac{1}{1 - \pi_0} \left( \int_{\pi_0}^{1} \underline{J}^{\top} (r) dr \right) \mathcal{S}_2^{-1}
\\
\\
- \mathcal{S}_2^{-1} \frac{ \big( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \big)}{\pi_0 (1 - \pi_0)} \ \ & \ \ \mathcal{S}_2^{-1}
\end{bmatrix}
\\
&=
\begin{bmatrix}
\left\{ (1 - \pi_0) \left( (1 - \pi_0) - \left( \int_{\pi_0}^{1} \underline{J}^{\top} (r) dr \right) \Delta_2^{-1} \big( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \big) \right)^{-1} - \frac{1}{1- \pi_0} \left( \int_{\pi_0}^{1} \underline{J}^{\top} (r) dr \right) \mathcal{S}_2^{-1} \big( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \big) \right\} & - \left( \int_0^{\pi_0} \underline{J}^{\top} (r) dr \right) \mathcal{S}_2^{-1} V_{zz}
\\
\\
- \mathcal{S}_2^{-1} \big( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \big) + \mathcal{S}_2^{-1} \big( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \big) \ \textcolor{red}{= 0} & (1 - \pi_0 ) \mathcal{S}_2^{-1} V_{zz}
\end{bmatrix}
\nonumber
\\
\nonumber
\\
& \times
\begin{bmatrix}
\left( (1 - \pi_0) - \left( \int_{\pi_0}^{1} \underline{J}^{\top} (r) dr \right) \Delta_2^{-1} \big( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \big) \right)^{-1} & - \frac{1}{1 - \pi_0} \left( \int_{\pi_0}^{1} \underline{J}^{\top} (r) dr \right) \mathcal{S}_2^{-1}
\\
\\
- \mathcal{S}_2^{-1} \frac{ \big( \underline{J}_c( 1 ) - \underline{J}_c( \pi_0 ) \big)}{\pi_0 (1 - \pi_0)} \ \ & \ \ \mathcal{S}_2^{-1}
\end{bmatrix}
\nonumber
\\
\nonumber
\\
&=
\begin{bmatrix}
\beta_1 & \ \ \beta_2
\\
\\
\beta_3 & \ \ \beta_4
\end{bmatrix}_{ \textcolor{red}{ 2 \times 2 } }
\end{align*}