EconBase
← Back to paper

Quasi-maximum likelihood estimation of break point in high-dimensional factor models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

82,123 characters · 9 sections · 54 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\fontsize{12}{14pt plus.8pt minus .6pt}\selectfont

center[center omitted — 109 chars of source]

\centerline{ Jiangtao Duan\textsuperscript{1}, Jushan Bai\textsuperscript{2}, Xu Han\textsuperscript{3} } \centerline{\it \textsuperscript{1}Northeast Normal University, \textsuperscript{2}Columbia University and \textsuperscript{3}City University of Hong Kong} \fontsize{9}{11.5pt plus.8pt minus .6pt}\selectfont

quotation{\it Abstract:} This paper estimates the break point for large-dimensional factor models with a single structural break in factor loadings at a common unknown date. We propose a quasi-maximum likelihood (QML) estimator of the change point based on the second moments of factors, which are estimated by a single principal component analysis. We show that the QML estimator is consistent for the true break point when the covariance matrix of the pre- or post-break factor loading (or both) is singular. Consistency here means that the deviation of the estimated break date from the actual break date $k_0$ converges to zero as the sample size grows. This is a much stronger result than the break fraction $\hat k/T$ being $T$-consistent (super-consistent) for $k_0/T$. Also, singularity occurs for most types of structural changes, except for a rotational change. Even for a notational change, the QML estimator is still $T$-consistent in terms of the break fraction. Simulation results confirm the theoretical properties of this estimator, and in fact QML significantly outperforms existing estimators for change points in factor models. Finally we apply the method to estimate the break points in a U.S. macroeconomic dataset and a stock return dataset. {\it Key words and phrases:} Structural break, High-dimensional factor models, Factor loadings

Introduction

Large factor models assume that a few factors can capture the common driving forces of a large number of economic variables. Although factor models are useful, practitioners have to be cautious about the potential structural changes. For example, either the number of factors or the factor loadings may change over time. This concern is empirically relevant because parameter instability is pervasive in large-scale panel data.

So far, many methods have been developed to test structural breaks in factor models (e.g., Stock2008, Breitung2011, and Chen2014). The rejection of the null hypothesis of no structural change leads to the subsequent issues of how to estimate the change point, determine the numbers of pre- and post-break factors, and estimate the factor space. Chen2015 considers a least-squares estimator of the break point and proves the consistency of the estimated break fraction (i.e., the break date $k$ divided by the full time series $T$, $\frac{k}{T}$). Cheng2016 propose a shrinkage method to obtain a consistent estimator of the break fraction. Baltagi2017 develop a least-squares estimator of the change point based on the second moments of the estimated pseudo-factors and show that the estimation error of the proposed estimator is $O_p(1)$, which indicates the consistency of the estimated break fraction. A few recent studies also explore a consistent estimation of break points, which is technically more challenging. Ma_Su2018 develop an adaptive fused group Lasso method to consistently estimate all break points under a multibreak setup. Barigozzi2018 propose a method based on wavelet transformations to consistently estimate the number and locations of break points in the common and idiosyncratic components. Bai2017 establish the consistency of the least-squares estimator of the break point in large factor models when factor loadings are subjected to a structural break and the size of the break is shrinking as the sample size increases. Although the estimators proposed in these studies are consistent under certain assumptions, the simulation results show that they perform poorly when (1) the number of factors changes after the break or (2) the loading matrix undergoes a rotational type of change.

According to the factor model literature, a factor model with a break in factor loadings is observationally equivalent to that with constant loadings and possibly more pseudo-factors (e.g., Han2015 and Bai2016). Thus, the estimation of the change point of factor loadings can be converted into that of the change point of the second moment of the pseudo-factors. We propose a quasi-maximum likelihood (QML) method to estimate the break point based on the second moment of the estimated pseudo-factors; therefore, the number of original factors is not required to be known for computing our estimator. First, we estimate the number of pseudo-factors (defined as the factors in the equivalent representation that ignores the break), and then estimate the pre- and post-break second moment matrices of the estimated pseudo-factors for all possible sample splits. The structural break date is estimated by minimizing the QML function among all possible split points.

This paper makes the following contributions to the literature. First, we establish the consistency of the QML break point estimator if the break leads to more pseudo-factors than the original pre- or post-break factors. This occurs when the break augments the factor space or in the presence of disappearing or emerging factors. Under these circumstances, the covariance matrix of loadings on the pre- or post-break pseudo-factors is singular, which is the key condition to establish the consistency of our QML estimator. To the best of our knowledge, this is the first study that links the consistency of the break point estimator to the singularity of covariance matrices of loadings on pre- and post-break pseudo-factors. In addition, we prove that the difference between the estimated and true change points is stochastically bounded when both pre- and post-break loadings on the pseudo-factors have nonsingular covariance matrices. In this case, the loading matrix only undergoes a rotational change, and both the numbers of pre- and post-break original factors are equal to the number of pseudo-factors.

The aforementioned singularity leads to a technical challenge of analyzing the asymptotic property. The singular population covariance matrix of the pre(post)-break loadings has a zero determinant, whose logarithm is undefined. To resolve this issue, we show that the estimated covariance matrices have nonzero determinants and a well-defined inverse for any given sample size, by obtaining the convergence rate of the lower bound of their smallest eigenvalues. This ensures that the objective function based on the estimated covariance is appropriately defined in any finite sample.

Our second major contribution is that the QML method allows a change in the number of factors. Namely, it allows for disappearing or emerging factors after the break. This is an advantage over the methods developed by Ma_Su2018 and Bai2017, who assume that the number of factors remains constant after the break. Our simulation result indicates that the estimator proposed by Bai2017 is inconsistent when some factors disappear and the remaining factors have time-invariant loadings. Baltagi2017 allow a change in the number of factors; however, their estimation error was only stochastically bounded. In contrast, our QML estimator remains consistent under a varying number of factors.

Finally, the QML method has a substantial computational advantage over the estimators that iteratively implement high-dimensional principal component analysis (PCA). For example, the estimator proposed by Bai2017 runs PCA for pre- and post-split sample covariance matrices for all possible split points. In comparison, our QML runs PCA for the entire sample only once, and thus, is computationally more efficient, especially in large samples.

The rest of this paper is organized as follows. Section 2 introduces the factor model with a single break on the factor loading matrix and describes the QML estimator for the break date. Section 3 presents the assumptions made for this model. Section 4 presents the consistency and asymptotic distribution of the QLM estimator for the break date. Section 5 investigates the finite-sample properties of the QML estimator through simulations. Section 6 implements the proposed method to estimate the break points in a monthly macroeconomic dataset of the United States and a dataset of weekly stock returns of Nasdaq 100 components. Section 7 concludes the study.

The following notations will be used throughout the paper. Let $\rho_i(\mathbb{B})$ denote the $i$-th eigenvalue of an $n\times n$ symmetric matrix $\mathbb{B}$, and $\rho_1(\mathbb{B})\geq \rho_2(\mathbb{B})\geq \cdots \geq \rho_n(\mathbb{B})$. For an $m\times n$ real matrix $\mathbb{A}$, we denote its Frobenius norm as $\|\mathbb{A}\|= [tr(\mathbb{A}\mathbb{A}^{'})]^{1/2}$, its MP inverse as $\mathbb{A}^{-}$, its $i$-th singular value as $\sigma_{i}(\mathbb{A})$, and its adjoint matrix as $\mathbb{A}^{\#}$ when $m=n$. Let $\mathrm{Proj}(\mathbb{A}|\mathbb{Z})$ denote the projection of matrix $\mathbb{A}$ onto the columns of matrix $\mathbb{Z}$. For a real number $x$, $[x]$ represents the integer part of $x$.

Model and estimator

Let us consider the following factor model with a common break at $k_0$ in the factor loadings for $i=1,\cdots,N$:

eqnarray[eqnarray omitted — 178 chars of source]

where $f_t$ is an $r-$dimensional vector of unobserved common factors; $r$ is the number of pseudo-factors; $k_0(T)$ is the unknown break date; $\lambda_{i1}$ and $\lambda_{i2}$ are the pre- and post-break factor loadings, respectively; and $e_{it}$ is the error term allowed to have serial and cross-sectional dependence as well as heteroskedasticity. $\tau_0\in (0,1)$ is a fixed constant and $[x]$ represents the integer part of $x$. For notational simplicity, hereinafter, we suppress the dependence of $k_0$ on $T$. Note that the dimension of $f_t$ is the same as that of the pseudo-factors (to be defined soon) instead of the original underlying factors. This formulation simplifies the representation of various types of breaks in a unified framework, which will be clarified in the examples below.

In vector form, model ((ref)) can be expressed as

eqnarray[eqnarray omitted — 170 chars of source]

where $x_t=[ x_{1t},\cdots,x_{Nt} ]^{'}$, $e_t=[ e_{1t},\cdots,e_{Nt} ]^{'}$, $\Lambda_1=[ \lambda_{11},\cdots,\lambda_{N1} ]^{'}$, and $\Lambda_2=[ \lambda_{12},\cdots,\lambda_{N2} ]^{'}$.

For any $k=1,\cdots,T-1$, we define $$X_{k}^{(1)}=[x_1,\cdots,x_k]^{'},X_{k}^{(2)}=[x_{k+1},\cdots,x_T]^{'},$$ $$F_{k}^{(1)}=[f_1,\cdots,f_k]^{'},F_k^{(2)}=[f_{k+1},\cdots,f_T],$$ $$\mbox{\boldmath $e$}_{k}^{(1)}=[e_1,\cdots,e_k]^{'},\mbox{\boldmath $e$}_k^{(2)}=[e_{k+1},\cdots,e_T],$$ where the subscript $k$ denotes the date at which the sample is to be split, and the superscripts $(1)$ and $(2)$ denote the pre- and post-$k$ data, respectively. We rewrite ((ref)) using the following matrix representation:

eqnarray[eqnarray omitted — 725 chars of source]

where $F_{k_0}^{(1)}$ and $F_{k_0}^{(2)}$ have dimensions $k_0\times r$ and $(T-k_0)\times r$, respectively, and $\Lambda$ is an $N\times r$ matrix with full column rank. The pre- and post-break loadings are modeled as $\Lambda_1=\Lambda B$ and $\Lambda_2=\Lambda C$, respectively, where $B$ and $C$ are some $r\times r$ matrices. Both $\Lambda_1$ and $\Lambda_2$ have dimension $N\times r$. In this model, $r_1=rank(B)\leq r$ and $r_2=rank(C)\leq r$ denote the numbers of original factors before and after the break, respectively. We refer to $G$ in ((ref)) as the pseudo-factors because the last line of ((ref)) provides an observationally equivalent representation without a change in the loadings matrix $\Lambda$. In other words, if the break is ignored in the estimation process, then the factors being estimated by a full-sample PCA are actually the pseudo-factors $G$ in ((ref)). It is well known that the break can augment the factor space; thus, $r_1 \leq r$ and $r_2 \leq r$, with $rank(G)=r$. Our representation in ((ref)) allows for changes in the factor loadings and the number of factors. Below, several examples are provided to illustrate that the pseudo-factor representation in ((ref)) is general enough to cover three types of breaks.

Type 1. Both $B$ and $C$ are singular. In this case, the number of original factors is strictly less than that of the pseudo-factors both before and after the break (i.e., $r_{1}<r$ and $r_{2}<r$). This means that the structural break in the factor loadings augments the dimension of the factor space. Let us consider the following example.

Example (1): Let $\mathbb{F}_{k_{0}}^{(1)}$$(k_{0}\times r_{1})$ and $\mathbb{F}_{k_{0}}^{(2)}$$((T-k_{0})\times r_{2})$ denote the original factors before and after the break, respectively, and $\Theta_{1}$ and $\Theta_{2}$ denote the pre- and post-break loadings on these factors. Thus, this model can be represented and transformed as

eqnarray[eqnarray omitted — 737 chars of source]

where $\Lambda=[\Theta_{1},\Theta_{2}]$, $B=diag(I_{r_{1}},0_{r_{2}\times r_{2}})$, $C=diag(0_{r_{1}\times r_{1}},I_{r_{2}})$, $F_{k_{0}}^{(1)}=[\mathbb{F}_{k_{0}}^{(1)}\;\vdots\;*]$, $F_{k_{0}}^{(2)}=[*\;\vdots\;\mathbb{F}_{k_{0}}^{(2)}]$, and the asterisk denotes some unidentified numbers such that all rows in $F_{k_{0}}^{(1)}$ and $F_{k_{0}}^{(2)}$ have the same variance (to satisfy Assumption (ref) in Section 3). (Note that the asterisk entries are cancelled due to multiplication by zero in $B$ and $C$.) In the special case of $r_{1}=r_{2}$, $\Lambda$ is of full rank $2r_{1}$ (i.e., the dimension of the pseudo-factor space is twice that of the original factor space) if the shift in the loading matrix $\Theta_{2}-\Theta_{1}$ is linearly independent of $\Theta_{1}$. We refer to this special case as the shift type of change, because the augmentation of the factor space is induced by a linearly independent shift in the loading matrix. Hence, Type 1 covers the shift type of change.

Type 2. Only $B$ or $C$ is singular. In this case, emerging or disappearing factors are present in the model. Let us consider the following example of disappearing factors.

Example (2): Without loss of generality, let us assume that $r_{2}<r_{1}$ and $\Theta_{2}$ is equal to the first $r_{2}$ columns of $\Theta_{1}$; thus, the last $r_{1}-r_{2}$ factors disappear after the break. Therefore, we can obtain the pseudo-factors by using the following transformation from the original factors $\mathbb{F}$:

eqnarray[eqnarray omitted — 646 chars of source]

where $F_{k_{0}}^{(1)}=\mathbb{F}_{k_{0}}^{(1)}$, $F_{k_{0}}^{(2)}=[\mathbb{F}_{k_{0}}^{(2)}\;\vdots\;*]$, $C=diag(I_{r_{2}},0_{(r_{1}-r_{2})\times(r_{1}-r_{2})})$, $\Lambda=\Theta_{1}$, and the asterisk is defined in a similar manner to that in (ref). In this example, $B=I_{r_{1}}$, $r=r_{1}$, and $r_{2}=\mathrm{rank}(C)<r$. Symmetrically, if $B$ is singular and $C=I_{r_{2}}$, then $r_{2}=r$ and $r_{1}=\mathrm{rank}(B)<r$, which means that certain factors emerge after the break point. Type 2 changes are important in empirical analysis. Please refer to Mcalinn2018 for empirical evidence regarding the varying number of factors in the U.S. macroeconomic dataset. For Types 1 and 2, we obtain a significant result that $P(\hat{k}\ifmmode\begingroup\defbold{bold} \text{\ifx\math@versionbold\bfseries\fi\textminus}\endgroup\else\textminus\fik_{0}=0)\to1$ as $N,T\to\infty$. \footnote{Technically, Types 1 and 2 can be combined into one type that involves singularity, which renders our QML estimator consistent. We consider Type 2 separately to emphasize the case of emerging and disappearing factors. }

Type 3. Both $B$ and $C$ are nonsingular. In this case, the loadings on the original factors undergo a rotational change, and the dimension of the original factors is the same as that of the pseudo-factors.

Example (3): Let us assume that $r_{2}=r_{1}$ and $\Theta_{2}=\Theta_{1}C$ for a nonsingular matrix $C$. The model with the original factors $\mathbb{F}$ can be transformed into the following pseudo-factor representation:

eqnarray[eqnarray omitted — 531 chars of source]

where $F_{k_{0}}^{(1)}=\mathbb{F}_{k_{0}}^{(1)}$, $F_{k_{0}}^{(2)}=\mathbb{F}_{k_{0}}^{(2)}$, and $\Lambda=\Theta_{1}$. In this example, $B=I_{r_{1}}$ and $r=r_{1}=r_{2}$, and the factor dimension remains constant. In the observationally equivalent pseudo-factor representation, the loading is time-invariant and the original post-break factors $\mathbb{F}_{k_{0}}^{(2)}$ are rotated by $C$. We refer to this as the rotation type of change.

The above examples show that a factor model with any of these three types of change can be unified and reformulated by the representation in ((ref)) with pseudo-factors. This representation controls the break type by varying the settings for $B$ and $C$, and thus, is convenient for our theoretical analysis.

Bai2017 rule out the rotation type of change because the break date is not identifiable by minimizing the sum of squared residuals. Baltagi2017 allow changes in the number of factors and rotation type of change; however, the difference between their estimator and the true break point is only stochastically bounded (i.e., their estimator is not consistent). Ma and Su's (2018) setup requires $r_1 = r_2$; thus, Type 2 is ruled out under their assumptions. Our simulation result shows that Ma and Su's estimator does not perform well under rotational changes (Type 3), whereas our QML method can handle changes in all three types discussed above. We obtain a significant result that $\hat{k}-k_0=O_p(1)$ if both $B$ and $C$ are of full rank (i.e., Type 3) and $\hat{k}-k_0=o_p(1)$ if $B$ or $C$, or both, is singular (i.e., Type 1 and Type 2).

In this paper, we consider the QML estimator of the break date for model ((ref)):

eqnarray[eqnarray omitted — 97 chars of source]

where $[\tau_1 T]$ and $[\tau_2 T]$ denote the prior lower and upper bounds for the real break point $k_0$ with $\tau_1,\tau_2\in (0,1)$ and $\tau_1 \leq \tau_0 \leq \tau_2$. The QML objective function $U_{NT}(k)$ is equal to

eqnarray[eqnarray omitted — 111 chars of source]

where $\hat{\Sigma}_1$ and $\hat{\Sigma}_{2}$ are defined as

eqnarray[eqnarray omitted — 190 chars of source]

and $\hat{g}_t$ is the PCA estimator of $g_t$ (i.e., the transpose of the $t$-th row of $G$). We define $\Sigma_{G,1}=E(g_t g_t^{'})$ for $t\leq k_0$, $\Sigma_{G,2}=E(g_t g_t^{'})$ for $t>k_0$, and $\Sigma_G=\tau_0 \Sigma_{G,1}+(1-\tau_0)\Sigma_{G,2}$. We define $\Sigma_\Lambda$ as the covariance matrix of $\Lambda$. The PCA estimator $\hat{g}_t$ is asymptotically close to $H^{'}g_t$ for a rotation matrix $H$, and $H \xrightarrow{p} H_0=\Sigma_{\Lambda}^{1/2}\Phi V^{-1/2}$ as $(N,T)\rightarrow \infty$, where $V$ and $\Phi$ are the eigenvalue and eigenvector matrices of $\Sigma_{\Lambda}^{1/2}\Sigma_G \Sigma_{\Lambda}^{1/2}$, respectively. Evidently, the second moment of $H_0 g_t$ shares the same change point as that of $g_t$. Therefore, we proceed to estimate the pre- and post-break second moments of $g_t$ by using the estimated factors $\hat{g}_t$, and then use ((ref)) to obtain the QML break point estimator $\hat{k}_{QML}$. Similar QML objective functions have been used for multivariate time series with observed data (e.g., Bai2000).

Assumptions

In this section, we state the assumptions made for establishing the consistency and asymptotic distribution of the QML estimator.

assum(i) $E\left\|f_t \right\|^4<M<\infty$, $E(f_tf_t^{'})=\Sigma_F$, where $\Sigma_F$ is positive definite, and $\frac{1}{k_0}\sum_{t=1}^{k_0}f_tf_t^{'}\xrightarrow{p}\Sigma_F,\frac{1}{T-k_0}\sum_{t=k_0+1}^{T}f_tf_t^{'}\xrightarrow{p}\Sigma_F$; (ii) There exists $d>0$ such that $\left\|\Delta\right\|\geq d>0$, where $\Delta =B\Sigma_FB^{'}- C\Sigma_FC^{'}$ and $B,C$ are $r\times r$ matrices.
assum$\left\| \lambda_{\ell i} \right\|\leq \bar{\lambda}<\infty$ for $\ell=1,2$, $i=1,\cdots,N$, $\left\| \frac{1}{N}\Lambda^{'}\Lambda-\Sigma_{\Lambda} \right\|\rightarrow 0$ for some $r\times r$ positive definite matrix $\Sigma_{\Lambda}$.
assumThere exists a positive constant $M<\infty$ such that \begin{itemize} • $E(e_{it})=0$ and $E|e_{it}|^8\leq M$ for all $i=1,\cdots,N$ and $t=1,\cdots,T$; • $E(\frac{e_s^{'}e_t}{N})=E(N^{-1}\sum_{i=1}^Ne_{is}e_{it})=\gamma_N(s,t)$ and $\sum_{s=1}^{T}|\gamma_N(s,t)|\leq M$ for every $t\leq T$; • $E(e_{it}e_{jt})=\tau_{ij,t}$ with $|\tau_{ij,t} |<\tau_{ij}$ for some $\tau_{ij}$ and for all $t=1,\cdots,T$ and $\sum_{j=1}^{N}|\tau_{ij}|\leq M$ for every $i\leq N$; • $E(e_{it}e_{js})=\tau_{ij,ts}$, \begin{equation*} \frac{1}{NT}\sum\limits_{i,j,t,s=1}|\tau_{ij,ts}|\leq M; \end{equation*} • For every $(s,t)$, $E\left| N^{-1/2}\sum_{i=1}^{N}(e_{is}e_{it}-E[e_{is}e_{it}]) \right|^4\leq M$. \end{itemize}
assumThere exists a positive constant $M<\infty$ such that \begin{eqnarray*} E(\frac{1}{N}\sum\limits_{i=1}^{N}\left\| \frac{1}{\sqrt{k_0}}\sum\limits_{t=1}^{k_0}f_te_{it} \right\|^2)&\leq& M,\\ E(\frac{1}{N}\sum\limits_{i=1}^{N}\left\| \frac{1}{\sqrt{T-k_0}}\sum\limits_{t=k_0+1}^{T}f_te_{it}\right\|^2)&\leq& M. \end{eqnarray*}
assumThe eigenvalues of $\Sigma_G\Sigma_\Lambda$ are distinct.
assumLet us define $\epsilon_t=f_tf_t^{'}-\Sigma_F$. According to the data-generating process (DGP) of factors, the H\'{a}jek-R\'{e}nyi inequality applies to the processes $\{\epsilon_t,t=1,\cdots,k_0\}$, $\{\epsilon_t,t=k_0,\cdots,1\}$, $\{\epsilon_t,t=k_0+1,\cdots,T\}$, and $\{\epsilon_t,t=T,\cdots,k_0+1\}$.
remarkUsing the H\'{a}jek-R\'{e}nyi equality on $\epsilon_t$, we can ensure that $\max\limits_{k_0<k\leq [\tau_2 T]}\| \frac{1}{T-k} \sum\limits_{t=k+1}^{T} f_tf_t^{'}-\Sigma_F\|=O_p(\frac{1}{\sqrt{T}})$ in Lemma (ref) and $\max\limits_{[\tau_1 T]\leq k<k_0}\| \frac{1}{k_0-k} \sum\limits_{t=k+1}^{k_0} g_tg_t^{'} \|=O_p(1), \max\limits_{k_0<k\leq [\tau_2 T]}\| \frac{1}{k-k_0} \sum\limits_{t=k_0+1}^{k} g_tg_t^{'} \|=O_p(1)$ in Lemmas (ref) and (ref).
assumThere exists an $M<\infty$ such that (i) For each $s=1,\cdots,T$, \begin{eqnarray*} E(\max_{k<k_0}\frac{1}{k_0-k}\sum_{t=k+1}^{k_0}|\frac{1}{\sqrt{N}}\sum_{i=1}^{N}[e_{is}e_{it}-E(e_{is}e_{it})]|^2)&\leq& M,\\ E(\max_{k> k_0}\frac{1}{k-k_0}\sum_{t=k_0+1}^{k}|\frac{1}{\sqrt{N}}\sum_{i=1}^{N}[e_{is}e_{it}-E(e_{is}e_{it})]|^2)&\leq& M; \end{eqnarray*} (ii) \begin{eqnarray*} E(\max_{k<k_0}\frac{1}{k_0-k}\sum_{t=k+1}^{k_0}\left\|\frac{1}{\sqrt{N}}\sum_{i=1}^{N} \lambda_i e_{it} \right\|^2)&\leq &M,\\ E(\max_{k>k_0}\frac{1}{k_0-k}\sum_{t=k_0+1}^{k}\left\|\frac{1}{\sqrt{N}}\sum_{i=1}^{N} \lambda_i e_{it} \right\|^2)&\leq &M. \end{eqnarray*}
assumThere exists an $M<\infty$ such that for all values of $N$ and $T$, (i) for each $t$, \begin{eqnarray*} E\left(\max_{1 \leq k<k_0}\frac{1}{k_0-k}\sum_{t=k+1}^{k_0} \left\| \frac{1}{\sqrt{NT}}\sum_{s=1}^{T}\sum_{i=1}^Nf_s[e_{is}e_{it}-E(e_{is}e_{it})] \right\|^2\right)&\leq& M,\\ E\left(\max_{k_0<k\leq T}\frac{1}{k-k_0}\sum_{t=k_0+1}^{k} \left\| \frac{1}{\sqrt{NT}}\sum_{s=1}^{T}\sum_{i=1}^Nf_s[e_{is}e_{it}-E(e_{is}e_{it})] \right\|^2\right)&\leq& M; \end{eqnarray*} (ii) the $r\times r$ matrix satisfies \begin{eqnarray*} E\left\| \frac{1}{\sqrt{NT}}\sum_{t=1}^{T}\sum_{i=1}^Nf_t\lambda_i^{'}e_{it} \right\|^2\leq M. \end{eqnarray*}

Asymptotic properties of the QML estimator

In this section, we derive the asymptotic properties of the QML estimator for various breaks. In the literature of structural breaks for a fixed-dimensional time series, conventional break point estimators, such as the least-squares (LS) estimator of Bai1997 or the QML estimator of Qu_Perron2007, are usually inconsistent. The estimation error of these conventional estimators is $O_p(1)$ when the break size is fixed. To reach consistency, the cross-sectional dimension of the time series must be large (e.g., Bai2010 and Kim2011).

Recall that the observationally equivalent representation in ((ref)) has time-invariant loadings and varying pseudo-factors. Hence, our problem converges to estimating the break point in the $r$-dimensional time series $g_t$, where $r$ is fixed. Theorems (ref) and (ref) below show that, for rotational breaks (Type 3), the convergence rate and limiting distribution are similar to those available in the literature. However, for Type 1 and 2 breaks, Theorem (ref) derives a much more significant result than that available in the literature, according to which our QML estimator is consistent even if our $g_t$ has only a fixed cross-sectional dimension $r$.

theoremUnder Assumptions (ref)--(ref), when both $B$ and $C$ are of full rank, $\hat{k}-k_0=O_p(1)$.

This theorem implies that the difference between the QML estimator and the true change point is stochastically bounded in model ((ref)). Although the estimation errors of both Baltagi2017 and our QML methods are bounded, the QML estimator has much better finite sample properties. To confirm this theoretical result, we conduct a simulation where the factor loadings have a rotational change (see DGP 1.B in Section 5). Table (ref) presents the MAEs and RMSEs of different estimators. The simulation result shows that the QML estimators have much smaller MAEs and RMSEs than other methods. In addition, $\hat{k}$ does not collapse to $k_0$, leading to a nondegenerate distribution. We will state the limiting distribution in Theorem (ref). Nevertheless, this theorem shows that the break point can be appropriately estimated because $\hat{\tau}=\hat{k}/T$ is still consistent for $\tau_0$.

To make an inference regarding the change point when both $B$ and $C$ are of full rank, we derive the limiting distribution of $\hat{k}$. Let us define

eqnarray*[eqnarray* omitted — 166 chars of source]

where $\Sigma_{1}=H_{0}^{'}\Sigma_{G,1}H_{0}$ and $\Sigma_{2}=H_{0}^{'}\Sigma_{G,2}H_{0}$ are the pre- and post-breaks of $H_{0}^{'}E(g_{t}g_{t}^{'})H_{0}$. The limiting distribution of $\hat{k}$ is given by the following theorem:

theoremUnder Assumptions (ref)--(ref), when both $B$ and $C$ are of full rank, \begin{eqnarray*} \hat{k}-k_0\xrightarrow{d} \arg\min\limits_\ell W(\ell), \end{eqnarray*} where \begin{eqnarray*} &&W(\ell)=\sum\limits_{t=k_0+\ell}^{k_0-1}tr((\Sigma_2^{-1}-\Sigma_1^{-1})\xi_t)-\left( tr(\Sigma_1\Sigma_2^{-1})-r-\log|\Sigma_1\Sigma_2^{-1}| \right)\ell\\ &&for \ell=-1,-2,\cdots,\\ &&W(\ell)=0 for \ell=0,\\ &&W(\ell)=\sum\limits_{t=k_0+1}^{k_0+\ell}tr((\Sigma_1^{-1}-\Sigma_2^{-1})\xi_t)+\left( tr(\Sigma_1^{-1}\Sigma_2)-r-\log|\Sigma_1^{-1}\Sigma_2| \right)\ell\\ &&for \ell=1,2,\cdots. \end{eqnarray*}

This result shows that the limiting distribution depends on $\xi_t$. If $\xi_t$ is independent over time, then $W(\ell)$ is a two-sided random walk. If $f_t$ is stationary, then $\xi_t$ is stationary in each regime. Here, the limiting distribution of the estimated break date is dependent on the generation processes of the unobserved factors, and thus, cannot be directly used to construct a confidence interval for a true break point. Bai2017 propose a bootstrap method to construct a confidence interval for $k_0$ when the change in the factor loading matrix shrinks as $N\to \infty$. However, their bootstrap procedure lacks robustness in the cross-sectional correlation in the error terms. In the current setup, the break magnitude $\left\| \Sigma_2-\Sigma_1 \right\|$ is fixed and we leave the case of shrinking break magnitude as a future topic.

Next, we establish a much stronger result than that available in the literature, which states that the QML estimator remains consistent when $B$ or $C$, or both, is singular. We make the following additional assumptions.

assumWith probability approaching one (w.p.a.1), the following inequalities hold: \begin{align*} &0<c\le \min_{[\tau_{1}T]\le k\le k_{0}}\rho_{j}\left(\frac{1}{Nk}\sum_{t=1}^{k}\Lambda^{\prime}e_{t}e_{t}^{\prime}\Lambda\right),\\ &0<c\le \min_{k_{0}\le k\le[\tau_{2}T]}\rho_{j}\left(\frac{1}{N(T-k)}\sum_{t=k+1}^{T}\Lambda^{\prime}e_{t}e_{t}^{\prime}\Lambda\right),\,for j=1,\cdots,r;\\ &\rho_{1}\left(\frac{1}{NT}\sum_{t=1}^{T}\Lambda^{\prime}e_{t}e_{t}^{\prime}\Lambda\right)\le \overline{c}<+\infty, \end{align*} as $N,T\to\infty$, where $\underline{c}$ and $\overline{c}$ are some constants.
assum\begin{align*} \max_{[\tau_{1}T]\le k\le k_{0}}\left\Vert \frac{1}{\sqrt{Nk}}\sum_{t=1}^{k}\sum_{i=1}^{N}f_{t}e_{it}\lambda_{i}^{\prime}\right\Vert & =O_{p}(1),\\ \max_{k_{0}\le k\le[\tau_{2}T]}\left\Vert \frac{1}{\sqrt{N(T-k)}}\sum_{t=k+1}^{T}\sum_{i=1}^{N}f_{t}e_{it}\lambda_{i}^{\prime}\right\Vert & =O_{p}(1).\end{align*}

Assumption (ref) is useful to derive the lower bound of the smallest eigenvalue of $\hat{\Sigma}_1$ (or $\hat{\Sigma}_2$) if $B$ (or $C$) is a singular matrix. Assumption (ref) strengthens Assumption (ref)(ii), which is similar to Assumption F2 of Bai2003. Note that the summation $\sum\limits_{t=1}^k$ in Assumptions (ref)-(ref) involves a positive fraction of observations over time since the lower bound of $k$ is $\tau_1 T$, with $\tau_1\in (0,1)$.

Also, as the log of matrix determinant is involved in the QML function, a natural problem is that the log determinant of a singular population covariance matrix is undefined when $B$ or $C$, or both, is singular. Fortunately, the determinants of $\hat{\Sigma}_1=\frac{1}{k}\sum\limits_{t=1}^k \hat{g}_t \hat{g}_t^{'}$ and $\hat{\Sigma}_2=\frac{1}{T-k}\sum\limits_{t=k+1}^T \hat{g}_t \hat{g}_t^{'}$ are small but not equal to zero in finite samples, when $\Sigma_1$ and $\Sigma_2$ are singular matrices. The following proposition develops a lower bound for the smallest eigenvalues of $\hat{\Sigma}_1$ and $\hat{\Sigma}_2$.

propositionUnder Assumptions (ref)--(ref), for $k\ge k_{0}$ and $k\le[\tau_{2}T]$, if $C$ is singular and $\sqrt{N}/T\to0$ as $N,T\to\infty$, then there exist constants $c_{U}\ge c_{L}>0$ such that \begin{align*} P\left(\min_{k\in[k_{0},[\tau_{2}T]]}\rho_{j}(\hat{\Sigma}_{2})\ge\frac{c_{L}}{N}\right) & \to1,\\ P\left(\max_{k\in[k_{0},[\tau_{2}T]]}\rho_{j}(\hat{\Sigma}_{2})\le\frac{c_{U}}{N}\right) & \to1,\end{align*} for $j=r_{2}+1,...,r$.

In proposition (ref), the lower bound of the smallest eigenvalue of the estimated sample covariance matrix $\hat{\Sigma}_2$ is $c_L/N$ for a constant $c_L>0$ w.p.a.1. A similar lower bound for the smallest eigenvalue of $\hat{\Sigma}_1$ can be obtained when $B$ is singular under the same assumptions. This ensures a lower bound for the determinants of the estimated sample covariance matrices. Proposition (ref) provides a useful tool to establish the consistency of our QML estimator. Although this technical result is a byproduct in our analysis, we believe that it is of independent interest and useful in other contexts.

assum(i) $[B,C]$ is of full row rank. (ii) $C^{\#}Bf_{k_0}\neq 0$ when $r-1=r_2>0$; and $B^{\#}Cf_{k_0+1}\neq 0$ when $r-1=r_1>0$, where $\mathbb{A}^{\#}$ denotes the adjoint matrix for a singular matrix $\mathbb{A}$. (iii) $\|Bf_{k_{0}}-\mathrm{Proj}(Bf_{k_{0}}|C)\|\ge d>0$ when $r-r_2\geq 2$ or $r_{2}=0$; and $\|Cf_{k_{0}+1}-\mathrm{Proj}(Cf_{k_{0}+1}|B)\|\ge d>0$ when $r-r_1\geq 2$ or $r_{1}=0$, where $\mathrm{Proj}(\mathbb{A}|\mathbb{Z})$ denotes the projection of $\mathbb{A}$ onto the columns of $\mathbb{Z}$, and $d$ is a constant.

Assumption (ref)(i) implies that $\Sigma_G$ is positive definite.$\footnote{Since

eqnarray*[eqnarray* omitted — 473 chars of source]

and $1<\tau_0<1$, Assumption (ref) implies that $\Sigma_G$ is a positive definite matrix. } $ Assumptions \ref{B_C_full_rank_project}(ii) implies that $B^{\#}C\neq 0$ when $r-1=r_1>0$, and $C^{\#}B\neq 0$ when $r-1=r_2>0$. It also excludes the possibility that $f_{k_0}$ and $f_{k_0+1}$ are in the null space of $C^{\#}B$ and $B^{\#}C$, respectively. Similarly, Assumption \ref{B_C_full_rank_project}(iii) rules out the cases that $Bf_{k_0}$ lies in the column space of $C$ when $r-r_2\geq 2$ or $r_2=0$ and that $Cf_{k_0+1}$ lies in the column space of $B$ when $r-r_1\geq 2$ or $r_1=0$.\footnote{Note that $r_2=0$ means $C=0$, so $B$ has to be nonsingular by Assumption \ref{B_C_full_rank_project}(i). Thus, Assumption \ref{B_C_full_rank_project}(iii) implies that $f_{k_0}\neq 0$ when $C=0$.} Assumption \ref{B_C_full_rank_project} is used to establish Lemma \ref{differ2}, which is useful for validating the consistency result that $Prob(\hat{k}-k = 0) \to 1$ in the proof of Theorem \ref{consistency}. It ensures that the value of the objective function becomes larger even if $\hat{k}$ slightly deviates from the true break point in large samples. Assumption \ref{B_C_full_rank_project} is flexible enough to allow various data generating processes for $f_t$. For example, if $f_{k_{0}}$ and $f_{k_{0}+1}$ have continuous probability distribution functions, then Assumptions \ref{B_C_full_rank_project}(ii)-(iii) just exclude a zero probability event since $C^{\#}B$ and $B^{\#}C$ are not equal to zero.

It is remarkable that existing estimators such as Baltagi2017 and Bai2017 are not consistent even if Assumptions (ref) -- (ref) hold. In contrast, our QML estimator is shown to be consistent under these additional assumptions. The following theorem summarizes the result.

theoremUnder Assumptions (ref)--(ref) and $\frac{N}{T}\to\kappa$, as $N,T\to\infty$ for $0<\kappa<\infty$, when $B$ or $C$, or both, is singular, $Prob(\hat{k}-k = 0) \to 1.$

Theorem (ref) shows that the estimated change point converges to the true change point w.p.a.1 when $B$ or $C$, or both, is singular (Types 1 and 2 in Section 2). This result is much more significant than that obtained by Baltagi2017, who show that the distance between the estimated and true break dates is bounded for Types 1--3. Note that the case in which only $B$ (or $C$) is singular corresponds to Type 2 with emerging (or disappearing) factors. Our QML estimator is consistent under this type of change, whereas Bai2017 and Ma_Su2018 rule out this type by assumption. In empirical applications, the conditions of theorem (ref) are rather flexible and likely to hold and the consistency of the break date estimator is expected in most economic data for the factor analysis.

remarkAn important contribution of Theorem (ref) is to link the consistency of the QML estimator with the singularity of the covariance matrices of the pre- or post-break factor loadings. The singularity is generated by the special structure of the pseudo-factors $g_t$ shown in ((ref)) and ((ref)) in the presence of a structural change. The PCA estimator $\hat{g}_t$ is consistent for $g_t$ (up to some rotation) for large $N$ and $T$, so the singularity structure is maintained in $\hat{\Sigma}_1$ and $\hat{\Sigma}_2$ and hence contributes to the consistency of our QML estimator. The result in Theorem (ref) is in contrast to conventional break point estimators, which only have $O_p(1)$ estimation errors in multivariate time series with a small cross-sectional dimension (e.g., Bai1997; Qu_Perron2007). Although our $\hat{g}_t$ has a fixed dimension, the divergence rate of the objective function depends on $N$. $\footnote{This is because the convergence rate of the smallest eigenvalue of $\hat{\Sigma}_2$ is $N^{-1}$ for $k = k_0$ when $C$ is singular. See Proposition 1.}$ In other words, our QML estimator still implicitly utilizes the information in the large cross-sectional dimension, which is the source of our consistency.
remarkThe conditions that $B$ or $C$, or both, is singular and $\frac{N}{T}\rightarrow \kappa\in (0,\infty)$ are likely to hold in many economic datasets for factor analysis. If both $B$ and $C$ are singular, the break occurs such that the number of pseudo-factors in the entire factor model is larger than that of the factors in the pre- and post-break subsamples. This can happen when the factor loadings undergo a shift type of change, as discussed in Example (1) for Type 1 changes. If $B$ is of full rank and $C$ is singular, some factors become irrelevant, and thus, the loading coefficients attached to these disappearing factors become zero. For example, in the momentum portfolio, some risks are not part of the firm's long-run structure as only sorting based on recent returns works; the reward is high but disappears within less than a year. If $B$ is singular and $C$ is of full rank, some factors emerge after the break date, increasing the dimension of the post-break factor space. For example, changes in the technology or policy may produce certain new factors.
remarkTheorem (ref) indicates that $U_{NT}(k)$ can be minimized to consistently estimate $k_0$. The intuition for this is that $U_{NT}(k)-U_{NT}(k_0)$ is always larger than zero, even if $k$ deviates only slightly from the true break point $k_0$, so that $\hat{k}$ must be equal to $k_0$ to minimize $U_{NT}(k)-U_{NT}(k_0)$. For example, in Type 1, when both $B$ and $C$ are singular for $k<k_0$, we can decompose $\hat{\Sigma}_2$ as $\hat{\Sigma}_2=\frac{1}{T-k}\sum\limits_{t=k+1}^{k_0} \hat{g}_t \hat{g}_t^{'}+\frac{1}{T-k}\sum\limits_{t=k_0+1}^T \hat{g}_t \hat{g}_t^{'}$, and the term $\frac{1}{T-k}\sum\limits_{t=k+1}^{k_0} \hat{g}_t \hat{g}_t^{'}$ results in a larger determinant of $\hat{\Sigma}_2$ than that of $\hat{\Sigma}_2^{0}=\frac{1}{T-k_0}\sum\limits_{t=k_0+1}^T \hat{g}_t \hat{g}_t^{'}$. By symmetry, we obtain a similar result for $k>k_0$. (See Lemmas (ref) and (ref) for more technical details.) Thus, $U_{NT}(k)-U_{NT}(k_0)>0$ w.p.a.1 as $N,T\to\infty$ if $k\neq k_0$.
remarkWith the QML estimator, we do not need to know the numbers of original factors $r_1$ and $r_2$ before and after the break point, but only the number of pseudo-factors in the entire sample. Bai2017 and Ma_Su2018 require knowledge of the number of original factors, which is much more difficult to estimate due to the augmented factor space resulting from the break. In practice, the number of pseudo-factors is much easier to estimate by using one of a number of estimators, such as the information criteria developed by Bai2002.

Simulation

In this section, we consider DGPs corresponding to Types 1--3 to evaluate the finite sample performance of the QML estimator. We compare the QML estimator with three other estimators. As shown below, $\hat{k}_{BKW}$ is the estimator proposed by Baltagi, Kao, and Wang (2017, BKW hereafter); $\hat{k}_{BHS}$ is the estimator proposed by Bai, Han, and Shi (2020, BHS hereafter); $\hat{k}_{MS}$ is the estimator proposed by Ma and Su (2018, MS hereafter); and $\hat{k}_{QML}$ is the QML estimator. Barigozzi2018 develops a change point estimator using wavelet transformation, which exhibits similar performance to that of the estimator proposed by Ma_Su2018. Hence, the comparison with the estimator proposed by Barigozzi2018 is not reported here, but the result is available upon request. The DGP roughly follows BKW, which can be used to examine various elements that may affect the finite sample performance of the estimators, and we use this DGP for model ((ref)). We calculate the root mean square error (RMSE) and mean absolute error (MAE) of these change point estimators $\hat{k}_{BKW}$, $\hat{k}_{BHS}$, and $\hat{k}_{QML}$, and each experiment is repeated 1000 times, where RMSE$=\sqrt{\frac{1}{1000}\sum\limits_{s=1}^{1000}(\hat{k}_s-k_0)^2}$ and MAE$=\frac{1}{1000}\sum\limits_{s=1}^{1000}|\hat{k}_s-k_0|$. When $T$ is small, there is a possibility that Ma and Su's (2018) method detects no break or multiple breaks; thus, the definition of the estimation error for a single break point in such cases is not straightforward. For a comparison, we compute the RMSE and MAE of the MS estimator by only using the results obtained by the MS estimator when it successfully detects a single break. As the computation of $\hat{k}_{BHS}$ and $\hat{k}_{MS}$ requires the number of original factors and that of $\hat{k}_{BKW}$ and $\hat{k}_{QML}$ requires the number of pseudo-factors, we set $\hat{r}=r_0$ for $\hat{k}_{BHS}$ and $\hat{k}_{MS}$ and $\hat{r}=r$ for $\hat{k}_{QML}$ and $\hat{k}_{BKW}$, where $r_0$ is the number of original factors and $r$ is the number of pseudo-factors.

We generate factors and idiosyncratic errors using a DGP similar to that of BKW. Each factor is generated by the following AR(1) process:

eqnarray*[eqnarray* omitted — 96 chars of source]

where $u_t=(u_{t,1},\cdots,u_{t,r_0})^{'}$ is i.i.d. $N(0,I_{r_0})$ for $t=2,\cdots,T$ and $f_1=(f_{1,1},\cdots,f_{1,r_0})^{'}$ is i.i.d. $N(0,\frac{1}{1-\rho^2}I_{r_0})$. The scalar $\rho$ captures the serial correlation of factors, and the idiosyncratic errors are generated by

eqnarray*[eqnarray* omitted — 96 chars of source]

where $v_t=(v_{1,t},\cdots,v_{N,t})^{'}$ is i.i.d. $N(0,\Omega)$ for $t=2,\cdots,T$ and $e_1=(e_{1,1},\cdots,e_{N,1})^{'}$ is $N(0,\frac{1}{(1-\alpha^2) \Omega})$. The scalar $\alpha$ captures the serial correlation of the idiosyncratic errors, and $\Omega$ is generated as $\Omega_{ij}=\beta^{|i-j|}$ so that $\beta$ captures the degree of cross-sectional dependence of the idiosyncratic errors. In addition, $u_t$ and $v_t$ are mutually independent for all values of $t$. We set $r_0=3$ and $k_0=T/2$. We consider the following DGPs for factor loadings and investigate the performance of the QML estimator for the three types of breaks discussed in Section 2.

DGP 1.A We first consider the case in which $C$ is singular, and set $C=[1,0,0;0,1,0;0,0,0]$. This setup aims to model ((ref)). In the pre-break regime, all elements of $\lambda_{i,1}$ are i.i.d. $N(0,\frac{1}{r_0^2}I_{r_0})$ across $i$. In the post-break regime, $\Lambda_2=(\lambda_{1,2},\cdots,\lambda_{N,2})^{'}=\Lambda_1C$. This case corresponds to a Type 2 change with a disappearing factor. The number of pseudo-factors is the same as $r_0$, so $r=3$, and the numbers of pre- and post-break factors are 3 and rank$(C)=2$, respectively. Table (ref) lists the RMSEs and MAEs of three estimators for different values of $(\rho, \alpha, \beta)$. In all cases, $\hat{k}_{QML}$ has much smaller MAEs and RMSEs than $\hat{k}_{BKW}$ and $\hat{k}_{BHS}$. Moreover, the MAEs and RMSEs of $\hat{k}_{QML}$ tend to decrease as $N$ and $T$ increase. This confirms the consistency of $\hat{k}_{QML}$ established in Theorem (ref). In addition, the RMSEs and MAEs of $\hat{k}_{BKW}$ do not converge to zero as $N$ and $T$ increase, which confirms that $\hat{k}_{BKW}$ has a stochastically bounded estimation error. $\hat{k}_{BHS}$ does not appear to be consistent when a factor disappears after the break. Moreover, a larger AR(1) coefficient $\rho$ tends to deteriorate the performance of $\hat{k}_{BKW}$, but does not have much impact on our QML estimator.

DGP 1.B We next consider the case in which $C$ is of full rank. We set $C$ as a lower triangular matrix. The diagonal elements are equal to $0.5$, $1.5$, and $2.5$, and the elements below these diagonal elements are i.i.d. and drawn from a standard normal distribution. Under this DGP, we have $r = r_0$. Table (ref) reports the performance of three estimators for different values of $(\rho, \alpha, \beta)$. In all cases, $\hat{k}_{BKW}$ and $\hat{k}_{QML}$ appear to have stochastically bounded estimation errors, which confirms Theorem 1 of BKW and Theorem (ref) of this paper. Both $\hat{k}_{QML}$ and $\hat{k}_{BKW}$ are inconsistent under this DGP; however, under all settings, our QML estimator tends to have much smaller RMSEs and MAEs than the estimator of BKW. The MAEs and RMSEs of $\hat{k}_{BHS}$ appear to increase with the sample size; thus, the BHS method cannot handle this case.

DGP 1.C In this case, we set $C=[1,0,0;2,1,0;3,2,m]$ and $m\in \{1,0.8,0.5,0.1,0\}$. As $m$ decreases to zero, the matrix $C$ changes from full rank to singular. We still consider serial correlation in factors and serial correlation and cross-sectional dependence in idiosyncratic errors simultaneously with $N=100, T=100$. Table (ref) shows that the MAEs and RMSEs of $\hat{k}_{QML}$ monotonically decrease with $m$, which confirms our findings in Theorems (ref) and (ref). In addition, the RMSEs and MAEs of $\hat{k}_{BKW}$ and $\hat{k}_{BHS}$ are much larger than those of $\hat{k}_{QML}$, and do not tend toward zero as $m$ decreases. For each value of $m$, the experiment is repeated 10000 times to more accurately estimate and compare the RMSEs (MAEs) of our QML estimator across different values of $m$.

DGP 1.D This DGP considers a Type 1 break. In the first regime, the last elements of $\lambda_{i,1}$ are zeros for all $i$, and the first two elements of $\lambda_{i,1}$ are both i.i.d. $N(0,\frac{1}{2}I_{r_0})$. In the second regime, $\lambda_{i,2}$ is i.i.d. $N(0,\frac{1}{3}I_{r_0})$ across $i$. As $\lambda_{i,1}$ and $\lambda_{i,2}$ are independent, the numbers of factors in the two regimes are $r_1=2$ and $r_2=3$, respectively, and the number of pseudo-factors is $r=5$. Because the numbers of pre- or post-break factors are smaller than that of the pseudo-factors, both $\Sigma_1$ and $\Sigma_2$ are singular matrices. Table (ref) reports the MAEs and RMSEs of $\hat{k}_{QML}$, $\hat{k}_{BHS}$, and $\hat{k}_{BKW}$ under this DGP. Table (ref) shows the performances of both $\hat{k}_{BHS}$ and our $\hat{k}_{QML}$. Their MAEs (RMSEs) are less than 0.05 (0.25) for all combinations of $N$, $T$, $\rho$, $\alpha$, and $\beta$. Although $\hat{k}_{BHS}$ is consistent under this DGP, our QML estimator still has smaller RMSEs than $\hat{k}_{BHS}$ in most cases reported in Table (ref). In addition, $\hat{k}_{BKW}$ performs better under this DGP than DGPs 1.A--1.C. However, its estimation error is much larger than that of our QML estimator. This is not surprising because $\hat{k}_{BKW}$ is not consistent. Finally, a larger AR(1) coefficient $\rho$ tends to yield a larger bias for $\hat{k}_{BKW}$, but does not have much effect on the performances of $\hat{k}_{BHS}$ and $\hat{k}_{QML}$.

In summary, Tables (ref) and (ref) show that the QML estimator performs much better than $\hat{k}_{BHS}$ under Type 2 and 3 breaks, which are ruled out under the assumptions of Bai2017. Table (ref) shows that the QML estimator often slightly outperforms $\hat{k}_{BHS}$, even though the latter is known to be consistent and has excellent finite-sample performance under Type 1 breaks. Note that the strength of the BHS method is the consistent estimation of break point for Type 1 break, especially when the size of the break is shrinking as the sample size increases. Under settings with a shrinking break size, the QML method will lose its power because the dimension of $G$ (determined by the IC criterion in Bai2002) will not be augmented, which means that the singularity does not show up in the covariance if breaks are small enough.

table[table omitted — 2,753 chars of source]
table[table omitted — 2,819 chars of source]
table[table omitted — 2,554 chars of source]
table[table omitted — 2,383 chars of source]

Tables (ref)--(ref) present the probabilities of the correct estimation of the break date. The results are consistent with those displayed in Tables (ref)--(ref): the QML estimator $\hat{k}_{QML}$ can detect the true break date with higher probabilities than others regardless of the values of $(\rho,\alpha,\beta)$. The MS method sometimes detects more than one or no break; hence, we only compute its probability of correctly estimating $k_0$ under the condition that it detects a single break. The probabilities of a correct estimation of the QML method increase with the sample sizes $N$ and $T$ in Tables (ref), (ref), and (ref).

Table (ref) shows that the probabilities of correct estimation of the QML estimators increase as $m$ decreases. A smaller $m$ means that $C$ is closer to a singular matrix. Table (ref) is consistent with Table (ref), and confirms Theorems (ref) and (ref). To explore in more detail the effect of changes in $m$ on the QML estimator, we vary the value of $m$ using finer grids and find a similar pattern to that shown in Table (ref). The results are reported in the supplementary appendix.

Figures (ref) and (ref) show the frequency of the estimated change points under DGP 1.A for $N=100,T=100$ and $N=500,T=500$ for 1000 replications. According to these figures, the QML estimators exhibit the highest frequency around the true break under different settings. When we increase the $(N,T)$ value from $100$ to $500$, the frequency at the true break point increases and the simulated distribution becomes tighter. This indicates that the QML estimators are highly likely to identify the true break point. This is consistent with our theory. However, the other three methods are found to have much larger variation and substantially lower probabilities to correctly estimate the break point. Thus, the QML estimators are advantageous in this case. Moreover, the simulation result indicates that for a sample size exceeding $N = 5000, T = 1000$, the probabilities of correctly estimating the QML estimator exceed $90\%$.

Recall that BKW and QML only have $O_p(1)$ estimation errors under DGP 1.B. However, Table (ref) shows that in all cases, the probabilities of correct estimation by the QML estimator are much higher than those of correct estimation by the BKW estimator Apparently, the BHS and MS methods cannot accurately estimate the true break point in this case. Figures (ref) and (ref) show the distributions of the estimated change points under (1.B) for $N=100,T=100$ and $N=500,T=500$, indicating that BHS and MS cannot handle rotational changes. Although the estimation errors of BKW and QML are bounded under all settings, the QML estimators have a much tighter distribution around the true break point.

table[table omitted — 2,228 chars of source]
table[table omitted — 2,239 chars of source]
table[table omitted — 2,212 chars of source]
table[table omitted — 1,980 chars of source]
figure[figure omitted — 583 chars of source]
figure[figure omitted — 583 chars of source]
figure[figure omitted — 623 chars of source]
figure[figure omitted — 623 chars of source]

Empirical Application

Macroeconomic data

In the first empirical application, we apply our proposed method to a U.S. macroeconomic dataset (Stock2012) to detect the possible structural breaks in the underlying factor model. We use the dataset adopted by Cheng2016, which comprises monthly observations of 102 U.S. macroeconomic variables. The sample begins after the Great Moderation and ranges from 1985:01 to 2013:01 $(T = 337)$. Following Bai2017, we focus on the subsample period between 2001:12 and 2013:01 $(T=134,N=102)$ because the complete data may have multiple breaks.

Cheng2016 find that 2007:12 is a single-break date, and that the pre-break and post-break subsamples have one factor and two or three factors, respectively. Following Cheng2016, Bai2017 also set the number of factors equal to one and two for the pre- and post-break subsamples, respectively. Then, they implement the LS estimation and obtain the estimated break point $\hat{k}=2008:12$. To implement our QML method, we first use Bai and Ng's information criterion IC1 and determine three pseudo-factors in the complete sample. Based on this result, we compute our QML estimator and obtain 2007:07 as the estimated break point, using which we split the sample into pre- and post-break subsamples. IC1 of Bai2002 detects two pre-break and three post-break factors. Based on the numbers of pre- and post-break factors and that of pseudo-factors, we can conclude that a new factor emerges after the break, so the QML estimator is consistent based on Theorem (ref).

Stock data

The second empirical application uses the weekly rate of return for Nasdaq 100 Index from April 18, 2019, to October 1, 2020. As all companies have data starting from April 18, 2019, we choose that as the start date. Traditionally, the index is limited to 100 common-stock issues, with only one issue allowed per issuer. Now, the index is limited to 100 issuers, some of which may have multiple issues as index components. The current index has 103 components, representing 100 issuers, four of which are from China: Baidu, JD.com, Ctrip, and NetEase. Thus, the sample size is $T=76$ and $N=103$. As IC1 and IC2 of Bai2002, the methods proposed by Onatski2010, Ahn_Horenstein2013, and Fan2020 yield different numbers of pseudo-factors for the sample, we use different number of factors $r=2,3,4,5,6,7$ to estimate the break date by using the QML method, and find that the estimated break date always falls in the week of February 20, 2020. This result agrees with that obtained using the method developed by Baltagi2017. In fact, the stock market began to fall sharply in the week of February 20, 2020, and two weeks later, the circuit breaker was triggered and U.S. stock market trading halted for a couple of times. Thus, the factor loading matrix appears to have changed in the early days of the epidemic.

Conclusions

We study the QML method for estimating the break point in high-dimensional factor models with a single structural change. We consider three types of changes and develop an asymptotic theory for the QML estimator. We show that the QML estimator is consistent when the covariance matrices of the pre- or post-break factor loadings, or both, are singular. In addition, the estimation error of the QML estimator is $O_p(1)$ when there is a rotation type of change in the factor loading matrix. We also derive the limiting distribution of the estimated break point in this case. Moreover, our QML estimator is computationally easy and fast because the eigendecomposition is conducted only once. The simulation results validate the suitable performance of the QML estimator. We use the proposed method to estimate the break point for U.S. macroeconomic data and stocks data.