EconBase
← Back to paper

Change Point Estimation in Panel Data with Time-Varying Individual Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

88,249 characters · 9 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Change Point Estimation in Panel Data with Time-Varying Individual Effects

abstractThis paper proposes a method for estimating multiple change points in panel data models with unobserved individual effects via ordinary least-squares (OLS). Typically, in this setting, the OLS slope estimators are inconsistent due to the unobserved individual effects bias. As a consequence, existing methods remove the individual effects before change point estimation through data transformations such as first-differencing. We prove that under reasonable assumptions, the unobserved individual effects bias has no impact on the consistent estimation of change points. Our simulations show that since our method does not remove any variation in the dataset before change point estimation, it performs better in small samples compared to first-differencing methods. We focus on short panels because they are commonly used in practice, and allow for the unobserved individual effects to vary over time. Our method is illustrated via two applications: the environmental Kuznets curve and the U.S. house price expectations after the financial crisis.

Introduction

In many panel datasets important variables of interest are missing, either because they are not available or because they are inherently unobservable. In a regression model of CO$_2$ emissions on energy consumption, variables such as a country's natural resources, political developments or the influence of environmental groups on decision making are typically not observed. While some of these unobserved variables or unobserved individual effects, such as initial natural resources, may be time-constant, many others, like political developments or the influence of environmental groups, typically vary over time. Unobserved individual effects are common in panels such as cross-country data, survey data, medical studies, and when not properly dealt with, they cause slope estimates to be inconsistent since they introduce an omitted variable bias. To ensure consistency, most panel data methods assume that the individual effects are time-constant and remove them before estimation, through time-demeaning or first-differencing the initial model. In this paper, we show that the individual effects need not be removed for estimating the number and location of multiple change points. Despite the asymptotic bias in the slope estimators introduced by the unobserved individual effects, our method estimates consistently the number and location of change points under reasonable assumptions.

Besides the long line of work on change points in time series models (see interalia Cs\"{o}rg\"{o} and Horv\'{a}th 1997; Bai and Perron 1998; Qu and Perron 2007; Harchaoui and L\`{e}vy-Leduc 2010; Aue and Horv\'{a}th 2013; Chan, Yau, and Zhang 2014; Perron and Yamamoto 2015; Qian and Su 2015) there is a growing body of work on change points in panel data. In panel data models, coefficients may exhibit common changes across customers, firms or countries due to for example policy changes, financial crises, housing bubbles or technological breakthroughs. Besides economics, panel change point methods have also been proposed to study changes in financial networks, in event counts for telecommunication networks, in encryptions of human speech and sound, and in genomic profiles of patients - see interalia Vert and Bleakley (2010); Cho and Fryzlewicz (2015); Bardwell et al (2018).

Regarding testing for change points, several test statistics have been proposed by Emerson and Kao (2001, 2002), de Wachter and Tzavalis (2012), Horv\'{a}th and Hu\v{s}kov\'{a} (2012), among others. Regarding estimation of multiple change points, one part of the literature is concerned with heterogeneous panels where the slope parameters are allowed to change across individuals or other panel units (see, e.g. Bai 2010; Kim 2011; Horv\'{a}th and Hu\v{s}kov\'{a} 2012; Chan, Horv\'{a}th and Hu\v{s}kov\'{a} 2013; Torgovitski 2015; Cho and Fryzlewicz 2015; Cho 2016; Baltagi, Feng and Kao 2016; Okuy and Wang 2018). The other part, including this paper, focuses on homogenous panels, where the slope parameters are constant across individuals. In this setting, Vert and Bleakley (2010) propose estimating the change points via a group least-angle approach; Qian and Su (2016) use an adaptive fused group Lasso (AGFL) method on the first-differenced data; Li, Qian and Su (2016) propose a principal component modified version of the AGFL method for dealing with a particular form of unobserved effects called interactive fixed effects; Baltagi, Kao and Liu (2017) employ an OLS method on the initial as well as the first-differenced data for stationary and nonstationary regressors; Bardwell et al (2018) use a minimum description length criterion.

The majority of theoretical developments on change point estimation, including the ones listed above, focus on long panels (where either only the number of periods $T$ goes to infinity, or both $T$ and the number of individuals $N$ tend to infinity). However, many panels such as surveys, cross-country data and local administrative data remain short, either because the data has been discontinued, just started, or because older data is unreliable. Therefore, many applications using panel data are forced to rely on less than twenty time periods - see interalia B\"orsch-Supan et al (2013) on the well-known European health survey SHARE, Baier and Bergstrand (2007) on a cross-country study of the impact of free trade agreements, Blanco and Ruiz (2013) on survey data analysis of the impact of crime on institutions and democracy, Kyle and Williams (2016) on a cross country analysis of health care and prescription drugs, Niu and van Soest (2014) on the American Life Panel (ALP) survey and Armona, Fuster and Zafar (2018) on another self-collected survey of house price expectations. Only a few papers develop methods for change point estimation in short panels ($N \rightarrow \infty$, $T$ fixed) while at the same time considering unobserved individual effects. Bai (2010) and Torgovitski (2015) do so, but treat the unobserved individual effects as individual means with potential change points, and estimate these change points in the mean without including other regressors in the model; Bai (2010) proposes an OLS method and Torgovitski (2015) focuses on non-parametric estimation. Qian and Su (2016) consider long and short panels, and estimate a panel regression model with fixed effects, where they rely on first-differenced data for consistent estimation of multiple change points.

Our main contribution is to provide an OLS method that consistently estimates the number of change points in short panel regression models without transforming the initial data. In our setting, the unobserved individual specific effects can either be treated as parameters or as random variables, and they may vary over time. Whether they are parameters or random variables, we omit them in the OLS estimation, and therefore a change point is defined as a change in the slope parameters, in the asymptotic bias of the OLS slope estimators, or in both. After all these (pseudo) change points are identified, we consistently estimate the slope estimators via OLS estimation of the demeaned model in each corresponding stable sub-sample. This latter step allows us to test whether each of the identified change points can be attributed to changes in the slope parameters, often the quantities of interest to applied researchers.\footnote{The only changes that cannot be labeled as changes in slope parameters are those that occur exactly one period after another change, because in this case, with no further assumptions, any transformation would just remove the period in-between two adjacent changes and therefore any information about the corresponding slope parameters would be lost. This case is further discussed in Section 2.}

The literature often assumes that the unobserved individual effects are of a particular functional form, such as time-invariant individual effects (fixed effects), additive effects (fixed effects plus cross-section invariant time effects) or interactive fixed effects (fixed effects times a common shock that is cross-section invariant but changes over time, see e.g. Pesaran 2006; Bai 2009; Bai and Li 2014; Moon and Weidner 2015). Since in many applications, these assumptions can be perceived as too restrictive, we adopt a more general specification where individual specific effects can vary over time both in a smooth and abrupt way.

The majority of papers (for long and short panels) that estimate changes in slope coefficients, like Qian and Su (2016), start from the premise that since OLS slope estimators are inconsistent due to the presence of unobserved individual effects, these effects need to be removed before change point estimation by means of some data transformation such as demeaning or first-differencing. In short panels, most of the variation is across individuals, and therefore such transformations, which typically remove a lot of cross-section variation, are problematic because they remove valuable information prior to change point estimation. Additionally, if the individual effects are not constant over the entire sample, first-differencing or any other available transformation does not fully remove them. In contrast to what most literature currently suggests, we prove that it is not necessary to transform the data for the purpose of change point estimation. Additionally, our simulation results show that in terms of correctly estimating the number of change points in small samples, our method performs better than the method in Qian and Su (2016), which relies on first-differencing.

Another contribution of this paper is to derive the asymptotic properties of two slope estimators while allowing for general time dependence and weak cross-section dependence in the level data. The first one is the conventional fixed effects estimator, obtained by OLS estimation of the initial model demeaned over each stable sub-sample, between two change points. For this estimator, we make the additional assumption that the unobserved individual effects change at the same time as the change in the slope parameters or the individual effects bias, which is still more general than assuming fixed, additive or interactive fixed effects. The second estimator is based on full-sample demeaning in the presence of fixed effects. We show that in the presence of fixed effects and change points, full-sample demeaning can lead to more efficient slope estimators as it uses the additional information that the individual effects do not change over time.

Related to this paper, for time-series models with regressors that are correlated with the errors, Perron and Yamamoto (2015) show under which conditions an OLS estimator for (pseudo) change points is consistent. They propose estimation of change points via sequential testing while our method consistently estimates the total number of (pseudo) change points in one step via an information criterion. Additionally, we show that if one imposes more change points than the truth (which may be desirable due to potential finite sample bias of post-selection methods), the set of estimated change points contains all the true ones with probability one in the limit.

The rest of the paper is organized as follows. Section 2 proves that our method consistently estimates the number and location of (pseudo) change points. Section 3 derives the asymptotic properties of the two proposed slope estimators. The finite-sample properties of the change point and slope estimators are studied through simulations in Section 4, and compared to the estimators in Qian and Su (2016). The practical use of our method is illustrated in Section 5 with two applications: the environmental Kuznets curve and the U.S. house price expectations in the aftermath of the financial crisis. In our first application, we show that in the implementation of the Kyoto protocol, major reductions in emission patterns occurred, which were unfortunately to a large extent undone after its implementation. In our second application, we show that determinants of house valuations changed from being largely subjective to being largely objective after the economy recovered from the recent financial crisis. Section 6 concludes. All the proofs are relegated to the Appendix.

Notation: Matrices and vectors are denoted with bold symbols, and scalars are not. Define for a scalar $S$, the generalized vec operator $\mbox{\bf vec}_{1:S}(\bm{A}_s) = (\bm A_1', \ldots, \bm A_S')'$, stacking in order the matrices $\bm A_s, (s=1,\ldots, S)$, which have the same number of columns. Let $\mbox{\bf diag}_{s=1:S}(\bm{A}_s) \equiv \mbox{\bf diag}_{1:S}(\bm{A}_s) = \mbox{\bf diag}(\bm A_1, \ldots, \bm A_{S})$ be the matrix that puts the submatrices $\bm A_1, \ldots, \bm A_S$ on the diagonal. If $S$ is the number of change points, $T_1,\ldots, T_S$ are the ordered candidate change points and $T$ the number of time series observations, let $\lambda_0=0$, $\lambda_{S+1}=1$, and let $\bm{\lambda}_S=(\lambda_0,\mbox{\bf vec}_{1:S}(\lambda_s)',\lambda_{S+1})'$ be a sample partition of the time interval $[1,T]$ divided by $T$, such that $\lambda_0=0$, $\lambda_{S+1}=1$, and $\lambda_s = T_s/T$ for $s=0,\ldots S+1$, with $T_0=0$ and $T_{S+1}=T$. Define constant regimes as $ I_s = [T_{s-1}+1, T_s]$ for $s=1, \ldots, S+1$. Let $\bm X = \mbox{\bf vec}_{1:T}(\bm X_t)$ be the $NT \times p$ matrix that stacks $\bm X_t=\mbox{\bf vec}_{i=1:N}(\bm x_{it}')$ in order. Call $\widetilde{\bm{X}} = \mbox{\bf diag}(\mbox{\bf vec}_{1:T_1}(\bm X_t), \mbox{\bf vec}_{T_1+1:T_2}(\bm X_t), \ldots, \mbox{\bf vec}_{T_{S}+1:T_{S+1}}(\bm X_t))$ the diagonal partition of $\bm X$ at $\bm \lambda_S$, with $\bm{X}_1, \ldots, \bm{X}_T$ on the diagonal and the rest of the elements zero. A superscript of $0$ on any quantity refers to the true quantity. For any random vector or matrix $\bm Z$, denote by $||Z||$ the Euclidean norm for vectors, or the square root of the maximum eigenvalue of $\bm Z'\bm Z$ for matrices. Also denote $||\bm Z||_{q}$ the $\mathcal L_q$ norm, i.e. $||\bm Z||_{q}= E(||\bm Z ||^q)^{1/q}$. For convenience, we denote by $0$ either a scalar, a vector or a matrix of zeros, and we only specify its dimension when it is unclear.

Change point estimation

Assume that the true model is piecewise-linear with $m^0$ change points:

align[align omitted — 134 chars of source]

In (ref), $i=1,\ldots,N$ are individuals, $t=1,\ldots,T$ are time periods, with $N$ large and $T$ fixed, $y_{it}$ are scalar continuous outcomes, $\bm x_{it}$ is a $p \times 1$ vector including the intercept and observed covariates, some of which may be constant over time; $m^0$ is the true unknown number of change points, with $1 \leq m^0 \leq T-1$. Also, $T_j^0,(j=1, \ldots, m^0)$ are the true unknown change points belonging to the sample partition $\bm{T}_{m^0}^0$. {\color{black}The true number and location of change points are properly defined in Assumption A(ref), and in model (1) they should be interpreted as possible changes in the unknown $p \times 1$ slope parameters $\bm{\beta}_j^0$.} Furthermore, $\epsilon_{it}$ are unobserved mean-zero idiosyncratic errors, uncorrelated with $\bm{x}_{it}$, and $c_{it}$ are the time-varying individual specific effects, which are either parameters or unobserved random variables that are uncorrelated with the idiosyncratic effects $\epsilon_{it}$ but possibly correlated with the observed covariates $x_{it}$. For example, in our second application, the subjective house price expectations equation contains unobservables related to individual optimism which may be correlated with covariates such as a home owner's view of his/her economic situation. {\color{black}For the purpose of change point estimation, the time-variation allowed in $c_{it}$ is quite general and further discussed after Assumption A(ref).}

Assume first that the number of change points $m^0$ is known. To describe the least-squares change point estimators $\hat{\mathbf T}_{m^0} = \hat{\bm{\lambda}}_{m^0} T$, let $u_{it} = c_{it} +\epsilon_{it}$, $\bm{u}= \mbox{\bf vec}_{t=1:T}(\mbox{\bf vec}_{i=1:N} (u_{it}))$, $\bm{\beta}^0=\mbox{\bf vec}_{j=1:m^0+1}(\bm{\beta}_j^0)$, $\bm{y}=\mbox{\bf vec}_{t=1:T}(\mbox{\bf vec}_{i=1:N} (y_{it}))$, and $\bm{\widetilde X}^0$ the diagonal partition of $\bm{X}$ at the true partition $\bm{\lambda}_{m^{0}}^0$. Then (ref) becomes:

equation[equation omitted — 79 chars of source]

We propose estimating (ref) by minimizing the sum of squared residuals over all possible sample partitions $\bm{\lambda}_{m^0}$, which is equivalent to regressing $\bm{y}$ on $\bm{\widetilde X}$ {\color{black}(where the latter was defined in the notation section above),}

equation[equation omitted — 322 chars of source]

and where $\hat \bm{\beta}_{\bm{\lambda}_{m^0}} = (\bm{\widetilde X}' \bm{\widetilde X})^{-1} \bm{\widetilde X}' \bm{y}$ is the OLS estimator using $\bm{\lambda}_{m^0}$ as the candidate partition. The minimizer of the above problem is denoted $\hat {\bm{\lambda}}_{m^0}$ or $\hat {\mathbf T}_{m^0} =\hat {\bm{\lambda}}_{m^0} T$ , and we refer to $\hat {\mathbf T}_{m^0}$ as the OLS change point estimators. If the minimizer is not unique, we break the tie by picking the smallest change point estimators. The OLS estimator of $\bm{\beta}^0$ at the estimated partition is denoted by $ \hat{\bm{\beta}} = \hat{\bm{\beta}}_{\hat \bm{\lambda}_{m^0}}= \mbox{\bf vec}_{1:m^0+1}(\hat{\bm{\beta}}_{j,\hat \bm{\lambda}_{m^0}}) $.

In general, $m^0$ is unknown and needs to be estimated. We propose estimating the number of change points by minimizing the following information criterion over $m=0,\ldots, T-1$, similar to BIC and HQIC: $$ IC(m) = \log S_{NT}(\hat \bm{\beta}_{\hat \bm{\lambda}_{m}}, \hat \bm{\lambda}_{m}) +p^*_m\ell_{NT}, $$ where $\ell_{NT}>0$, $\ell_{NT}\rightarrow 0$, $N \ell_{NT}\rightarrow \infty$, and $p_m^*$ is the number of parameters for a model with $m$ change points. Both Nimomiya (2005) - for a mean-shift model - and Hall and Sakkas (2013) - for a general regression model - show that the penalty for one change point should be three, not one, therefore we recommend to use $p_m^* = 3m+(m+1)p$. In the simulation section we show that the HQIC penalty, $\ell_{NT} = \log[\log (NT)]/NT$, is preferred to the BIC penalty, $\ell_{NT}= \log (NT)/NT$. The resulting estimator for the number of change points is $\hat m = \arg \min IC(m)$. Note that the information criteria is defined at the OLS change point estimators for a given number of change points, so we estimate the number and location of changes in one step.

For proving that our method consistently estimates the number and location of change points in $\bm{\gamma}_j^0$, the pseudo-true parameters defined below, we impose the following assumptions.

assnAs $N\rightarrow \infty$: (i) $(NT)^{-1/2}\sum_{i=1}^N \sum_{t=1}^T \bm{x}_{it} \epsilon_{it} \stackrel{d}{\rightarrow} \mathcal N(0, V)$, where $V$ is a positive definite (pd) matrix of constants; (ii) $N^{-1} \sum_{i=1}^N \epsilon_{it} c_{it} \stackrel{p}{\rightarrow} 0$ ;(iii) $ N^{-1} \sum_{i=1}^N \bm{x}_{it} c_{it}\stackrel{p}{\rightarrow} \bm{a}_j^0 $ for $t \in I_j^0$, $j=1,\ldots, m^0+1$; (iv) $ N^{-1} \sum_{i=1}^N \bm{x}_{it}\bm{x}_{it}'\rightarrow \bm{Q}_j^0 $ for $t\in I_j^0$, $j=1,\ldots, m^0+1$, where $ \bm{Q}_j^0$ are pd matrices of constants; (v) let $\bm{\gamma}_j^0 = \bm{\beta}_j^0+ (\bm{Q}_j^0)^{-1} \bm{a}_j^0$; then $\bm{\gamma}_j^0 \neq \bm{\gamma}_{j+1}^0$ for all $j=1,\ldots,m^0$; (vi)$ (NT)^{-1} \sum_{i=1}^N \sum_{t=1}^T \epsilon_{it}^2 \stackrel{p}{\rightarrow} \sigma^2_{\epsilon,T}$ and $ (NT)^{-1} \sum_{i=1}^N \sum_{t=1}^T c_{it}^2 \stackrel{p}{\rightarrow} \sigma^2_{c,T}$.

A(ref)(i) imposes a central limit theorem for {\color{black}sums} of $\bm{x}_{it} \epsilon_{it}$, allowing for general time-series dependence and for weak cross-section dependence. A(ref)(ii) assumes that if the time-varying individual specific effects $c_{it}$ are random variables, they are uncorrelated with the idiosyncratic errors $\epsilon_{it}$, a common assumption in panels with individual specific effects. {\color{black}If they are parameters, then they are also allowed to vary over time, but they are omitted in the estimation.}

A(ref)(iii)-(v) are key assumptions for consistent estimation of the number of change points. Note that they are not very restrictive in the sense that $\bm{\gamma}_j^0$ can change at each point in time, and it may change because of $\bm{\beta}_j^0$ or not. {\color{black}The allowed time variation in $c_{it}$ is implicitly defined by A(ref)(v), allowing $c_{it}$ to exhibit change points, smooth time-variation and/or jumps. The specification for $c_{it}$ includes fixed effects ($c_{it}=c_i$), interactive fixed effects ($c_{it}=c_i f_t$), but also (other) forms of stationary or non-stationary time variation. Since we define the change points as changes in $\bm{\gamma}_j^0$, as long as $\bm{\gamma}_j^0$ does not change, any time-variation in $c_{it}$ will not result in a change point. If $\bm{\gamma}_j^0$ changes because of change points in $c_{it}$, then these change points are identified by our method.}

A(ref)(vi) is a weak law of large numbers for sums of the second moments of $\epsilon_{it}$ and $c_{it}$, ensuring they do not increase with $N$. In Assumption A(ref), we consider a common set of primitive assumptions used for panel data (such as survey data), when the data is independent over $i$. Lemma (ref) below shows that A(ref) satisfies A(ref)(i)-(iv) and A(ref)(vi).

assn(i) $\bm{x}_{it}$ and $\epsilon_{it}$ are independent over $i$ with $E(\epsilon_{it})=0$ and $E(\epsilon_{it} \bm{x}_{it}) =0$; (ii) $\sup_{it} ||\bm{x}_{it}||_{4+\delta} <\infty$ and $\sup_{it} ||\epsilon_{it}||_{4+\delta} <\infty$ for some $\delta>0$. (iii) $E(\bm{x}_{it} \bm{x}_{is}'\epsilon_{it} \epsilon_{is}) =\bm V_{ts}^0$; (iv) $E(\bm{x}_{it} \bm{x}_{it}') =\bm{Q}_j^0$ for $t\in I_j^0$; (v) $c_{it}$ are independent over $i$, with $E(c_{it} \epsilon_{it})=0$, $E(\bm{x}_{it} c_{it}) = \bm{a}_j^0$ for $t \in I_j^0$ and $\sup_{it} ||c_{it}||_{4+\delta}<\infty$ ; (vi) $E(\epsilon_{it}^2)=\sigma_t^2$; (vii) $E(c_{it}^2)=\sigma_{ct}^2.$
lemmaIf A(ref) holds, then A(ref)(i)-(iv) and (vi) holds.
theoremLet A(ref) hold. As $N\rightarrow \infty$,\\ (i) If $m=m^0$, $\lim P(\hat T_j = T_j^0)=1$ and $\hat \bm{\beta}_{j, \hat \lambda_{m^0}} \stackrel{p}{\rightarrow} \bm{\gamma}_j^0$, for $j=1,\ldots, m^0+1$ \\ (ii) $\lim P(\hat m=m^0)=1$.\\ (iii) If $m>m^0$, then there are indices $j_1,\ldots j_{m^0} \in \{1,\ldots, m\}$ such that $\lim P[ \hat T_{j_s} = T_s^0] =1$ for all $s\in \{1,\ldots, m^0\}$.\\

Part (i) states that if we knew the true number of change points, their locations would be consistently estimated, and the corresponding parameter estimates would be consistent for their pseudo-true values $\bm{\gamma}_j^0$. Part (ii) states that the true number of change points is {\color{black}consistently estimated by the information criterion we propose.} Part (iii), a by-product of our proof, shows that if the number of changes imposed is larger than the true number of the changes, then our method estimates all the true change points with probability one in the limit (and some additional spurious change points).\footnote{This implies that an AIC variant of our information criterion, with penalty $\ell_{NT} =2/(NT)$, can also be used if the researcher is worried that the number of change points is underestimated.}

The intuition for the result in Theorem (ref) is similar to Perron and Yamamoto (2015) who proposed using OLS for estimating change points in time series models with regressors that are correlated with errors. While the parameter estimates are in general not consistent for their true values because of the omitted variable bias, they are consistent for the pseudo-true values $\bm{\gamma}_j^0$, therefore we can consistently estimate the number and location of change points in $\bm{\gamma}_j^0$. Note that unlike Perron and Yamamoto (2015), we propose an information criterion that consistently estimates the number of change points in one step. In contrast to Perron and Yamamoto (2015), we can also allow for a change point in each period.

The advantage of our method over Qian and Su (2016) and Li, Qian and Su (2016) is that we allow for time-varying individual effects without {\color{black}specifying a functional form for the time variation.} Moreover, if some covariates are time-invariant as typical in panel data (gender, race), then the method in Qian and Su (2016), relying on first-differencing the data prior to change point estimation, only benefits from one period to estimate a change in the coefficients on these covariates, while our method uses all the periods available. Further advantages are highlighted in the simulation section, where we show that our method is more precise in estimating the number of change points {\color{black}in finite samples}, because it does not remove important variation in the data by first-differencing.

We now discuss the results in Theorem (ref) in connection to {\color{black}assumption} A(ref) and typical panel data assumptions. In its full generality outlined above, the method in this section does not yet indicate which changes occur only in the slope parameters $\bm{\beta}_j^0$. However, because this may be of main interest to the applied researcher, the next section discusses under what conditions $\bm{\beta}_j^0$ can be consistently estimated by two demeaning procedures. We then suggest using these estimators to test $H_0: \bm{\beta}_j^0=\bm{\beta}_{j+1}^0$, $j=1,\ldots, m^0$, therefore identifying which changes pertain to $\bm{\beta}_j^0$ only.

In special cases, the changes in $\bm{\beta}_j^0$ can be directly identified by the methods of this section. The first case is when $c_{it}$ are random effects; in that case, they are uncorrelated with $\bm{x}_{it}$, in which case a weak law of large numbers can be employed to show that $\bm{a}_j^0 = 0$, therefore that $\bm{\gamma}_j^0=\bm{\beta}_j^0$. The second case is if $\bm{Q}_j^0=\bm{Q}^0$ and $\bm{a}_j^0=\bm{a}^0$. In this case, the correlation between the individual specific effects and the regressors does not change over time, and therefore all the changes in $\bm{\gamma}_j^0$ come from changes in $\bm{\beta}_j^0$, and no further testing is needed.

Our method can also be used as a diagnostic tool for modeling either time-varying parameters or time-varying individual effects. If $\hat m$ is large (close or equal to $T-1$), then (ref) should be revisited for better modelling of the time-variation in $\bm{\gamma}_j^0$. If $\hat m$ is small, then a model with interactive fixed effects, i.e. $c_{it} = c_i \times f_t$ might not be inappropriate unless it can be assumed that despite this specification, $\bm{a}_j^0$ does not change often. If the researcher is further willing to assume fixed effects (i.e. $c_{ij}=c_i$), a common assumptions in panel data, then the next section provides a more efficient estimator of $\bm{\beta}_j^0$ than currently available.

If the covariates include a lagged dependent variable $y_{it-1}$, then $E(y_{it-1} c_{it})$ would in general change in each period, leading to $m^0=T-1$ by definition. Therefore, in our analysis, we do no include lagged dependent variables, but allow for time-series dynamics in the error term. {\color{black}Employing time dummies for each time period, a common approach in short panels, is equivalent to imposing a change point in the intercept at each time period, which is not parsimonious nor necessarily justified by the data. We suggest avoiding this approach as our method can directly estimate the number and location of the changes in the intercept without having to assume they change at each period.}

Slope Estimators and Their Asymptotic Properties

In this section, we proceed as if the true number and location of change points in $\bm{\gamma}_j^0$ was known. For implementing the estimators described below, the true number and location of change points should be replaced by their corresponding estimates from Section 2.

To get consistent estimators of the slope parameters, we follow the common approach to remove the individual effects from model (ref) through a transformation of the data. {\color{black}However, for proper removal of these effects, we assume throughout this section that $c_{it} = c_{ij}$ for $t\in I_j^0$, meaning that the individual effects are allowed to change but only at the change points already detected.\footnote{This assumption can be generalized to allow for further time variation in $c_{it}$. However, in this case, any transformations such as demeaning would only remove the mean of $c_{it}$, leaving its variance in the error term. This variance would increase the variance of the slope estimators, but in general it would not be consistently estimable for fixed $T$ without further assumptions.}}

We propose two methods: (1) sub-sample demeaning, that is, demeaning over segments $I_j^0$, corresponding to the usual fixed effects estimator in the absence of change points; and (2) full sample demeaning, which is new. The former is appropriate whether $\bm{a}_j^0$ is changing over $j$ or not, and the latter only when we have time-invariant fixed effects, i.e. $c_{it}=c_i$.

Note that only parameters that are constant for more than one period can be estimated via sub-sample demeaning, because the others are automatically removed. In contrast, when $c_{it}=c_i$, the full-sample demeaning estimator identifies all the parameters, including the ones for which only one time period is available. To our best knowledge, the full-sample demeaning estimator is new, and it is imposing the additional information that the fixed effects are not changing over time, which in principal should lead to more efficient estimators. In Theorem (ref), we give sufficient conditions for this second estimator to be strictly more efficient than the first. Below, we describe these estimators.

For any vector $\bm{z}_{it}$, let $ \overline \bm{z}_i = T^{-1} \sum_{t=1}^T \bm{z}_{it} $ be the full sample average, and $\overline \bm{z}_{ij} = (T_j^0-T_{j-1}^0)^{-1} \sum_{t\in I_j^0} \bm{z}_{it}$ be the sub-sample averages. Then the FE estimator in $I_j^0$ is the OLS estimator in the demeaned sub-sample $I_j^0$ of model (ref): $$\hat{\bm{\beta}}_{FE,j} = \left(\sum_{i=1}^N \sum_{t \in I_j^0} \ (\bm{x}_{it} - \overline \bm{x}_{ij})(\bm{x}_{it} - \overline \bm{x}_{ij})' \right)^{-1} \sum_{i=1}^N \sum_{t \in I_j^0}\ (\bm{x}_{it} - \overline \bm{x}_{ij})(y_{it}-\overline y_{ij}).$$

Let $\bm{w}_{ij} =T^{-1}\sum_{t\in I_j^0} \bm{x}_{it}$. If we demean model (ref) over the full sample, then

align*[align* omitted — 188 chars of source]

Let $ y_{it}^* = y_{it} - \overline y_i $, $ \epsilon_{it}^* = \epsilon_{it} - \overline \epsilon_i $, and $ \widetilde \bm{x}_{it} =(-\bm{w}_1', -\bm{w}_2',\ldots, -\bm{w}_{j-1}',\bm{x}_{it}' - \bm{w}_j', - \bm{w}_{j+1}',\ldots, -\bm{w}_{m^0+1}')'.$ for $t\in I_j^0$. This model can be written more compactly as: $ y_{it}^* = \widetilde \bm{x}_{it}\bm{\beta}^0 + \epsilon_{it}^*$, and the OLS estimator in this equation is the full sample demeaning estimator, which we name as the FFE (full sample fixed effects) estimator: $$\hat{\bm{\beta}}_{FFE,j} = \left(\sum_{i=1}^N \sum_{t=1}^T \widetilde \bm{x}_{it} \widetilde \bm{x}_{it}'\right)^{-1} \left(\sum_{i=1}^N \sum_{t=1}^T \widetilde \bm{x}_{it} y_{it}^*\right). $$ Let $\mathcal S$ be the subset of the number of regimes with at least two observations, and denote a quantity defined over $\mathcal S$ by a subscript $\mathcal S$; for example, $\bm{\beta}_{\mathcal S}^0 = \mbox{\bf vec}_{j\in \mathcal S}(\bm{\beta}_j^0)$.

assnAs $N \rightarrow \infty$, (i) $ N^{-1} \sum_{i=1}^N \bm{x}_{it}\bm{x}_{is}' \stackrel{p}{\rightarrow} \bm{\Omega}_{ts}^0 $, a matrix of constants, for all $ t,s $; (ii) $ N^{-1/2} \sum_{i=1}^N \mbox{\bf vec}_{j \in \mathcal S}\left(\sum_{t\in I_j^0} (\bm{x}_{it}-\overline \bm{x}_{i,j})\epsilon_{it}\right) \stackrel{d}{\rightarrow} \mathcal N(0,\bm{W}_{1,\mathcal S}) $; (iii) $ N^{-1/2} \sum_{i=1}^N \left(\sum_{t=1}^T \widetilde \bm{x}_{it} \epsilon_{it}^* \right) \stackrel{d}{\rightarrow} \mathcal N(0, \bm{W}_{2})$.\footnote{The quantities $\bm{W}_1$ and $\bm{W}_2$ may depend on $T$, but for simplicity we do not explicitly express this in the notation.}

Assumption A(ref) facilitates the presentation of the asymptotic distributions for general time series dependence (including unit root dependence over $t$ in $\epsilon_{it}$) and weak cross-section dependence. Assumptions (ref) gives a primitive assumption that satisfies Assumption A(ref). Let $\bm{X}_i=\mbox{\bf vec}_{t=1:T}(\bm{x}_{it})$. Then:

assnA(ref)(i)-(iv), (vi) holds and: (i) $E(\epsilon_{it}\bm{X}_i) = 0$ for all $i,t$; (ii) $E(\bm{x}_{it} \bm{x}_{is}') =\bm{\Omega}_{ts}^0$; (iii) $E(\bm{x}_{it}\bm{x}_{is}' \epsilon_{it} \epsilon_{is}) = \bm N_{ts}^0$.
lemmaIf A(ref) is satisfied, then A(ref) are satisfied.

Let $\Delta T_j^0 = T_j^0-T_{j-1}^0$, $\bm{\Omega}_{1,\mathcal S} =\mbox{\bf diag}_{j\in \mathcal S}\left( \Delta T_j^0 \bm{Q}_j^0 - (\Delta T_j^0)^{-1} \bm{Q}_{jj}^0 \right)$, $\bm{Q}_{jj}^0=\sum_{t\in I_j^0}\sum_{s\in I_k^0} \bm{\Omega}_{ts}^0$, and let $\bm{\Omega}_2= \mbox{\bf diag}_{j=1:m^0+1}(\Delta T_j^0 \bm{Q}_j^0) - (2T^{-1} -T^{-2}) \widetilde \bm{Q}^0,$ where $\widetilde \bm{Q}^0$ is the $p(m^0+1) \times p(m^0+1)$ matrix with the $(j,k)$ sub-matrix of size $p\times p$ equal to $\bm{Q}_{jk}^0$.

theoremUnder A(ref), A(ref) and $N\rightarrow \infty$,\\ (i) $ \sqrt{N} \left( \hat{\bm{\beta}}_{FE,\mathcal S} - \bm{\beta}_{\mathcal S}^0\right) \stackrel{d}{\rightarrow} \mathcal{N}(0,\bm V_{FE,\mathcal S}), $ where $\bm V_{FE,\mathcal S} = \bm{\Omega}_{1,\mathcal S}^{-1} \bm{W}_{1,\mathcal S}\bm{\Omega}_{1,\mathcal S}^{-1} $; (ii) further assuming that $c_{it}=c_i$ for all $t=1,\ldots, T$, $ \sqrt{N} (\hat{\bm{\beta}}_{FFE} -\bm{\beta}^0) \stackrel{d}{\rightarrow} \mathcal N\left (0, \bm V_{FFE}\right), $ where $\bm V_{FFE} = \bm{\Omega}_2^{-1}\bm{W}_{2} \bm{\Omega}_2^{-1}$.

{\color{black}It is interesting to note that even though the demeaning removes time-invariant regressors such as gender and race for both estimators, the magnitude of change in the slopes of these regressors can be consistently estimated.\footnote{The elements of $\bm {\hat \beta}_{FE,\mathcal S}$ and $\bm \hat \beta_{FFE}$ in Theorem (ref) referring to time-invariant regressors should be replaced by magnitudes of changes, but we did not do this to simplify notation.}}

With no further assumptions on the time series dependence, it is unclear which estimator is more efficient; therefore, when it can be assumed that $c_{it}=c_i$, we suggest stacking the moment conditions implied by these two estimators, resulting in a generalized method of moments estimator which is more efficient than each of the two if the optimal weighting matrix is used.

Theorem (ref) shows that when the data is uncorrelated over time (as in typical static panels), the FFE estimator is strictly more efficient than the FE estimator, so if it can safely be assumed that $c_{it}=c_i$, then the FFE estimator is preferable. Since this result pertains only to panel data models with at least one change point and with two parameters that can be identified by FFE, we impose $T\geq 4$, $m^0\geq 1$ and $\Delta T_j^0 \geq 2$ for at least two regimes.

theoremLet A(ref) hold, $m^0>1$, $T\geq 4$ and $\Delta T_j^0 \geq 2$ for at least two regimes $j \in\{1,\ldots, m^0+1\}$. Assume that $E(\epsilon_{it}^2|\bm{X}_i) = \sigma^2$ for all $t$, and $E(\epsilon_{it}\epsilon_{is}|\bm{X}_i)=0$ and $E(\bm{x}_{it} \bm{x}_{is})=\bm{\Omega}_{ts}^0=0$ for all $t\neq s$. Then $$\bm V_{FE} = \mbox{\bf diag}_{1:m^0+1}\left[\sigma^2 (\bm{Q}_j^0)^{-1} \frac{1}{\Delta T_j^0-1}\right], \bm V_{FFE} = \mbox{\bf diag}_{1:m^0+1}\left[\sigma^2 (\bm{Q}_j^0)^{-1}\frac{(T^2-3T+1) T^2}{(T-1)^4} \frac{1}{\Delta T_j^0}\right],$$ and $\bm V_{FFE}- \bm V_{FE}$ is negative definite.

Theorem (ref) shows that the relative efficiency of the FFE estimator can be explicitly quantified.

Overall modelling strategy. Since this section provides consistent estimators in each period with more than one observation, one can test which parameters are actually changing by testing $H_0: \bm{\beta}_j^0 = \bm{\beta}_{j+1}^0$ for $j \in \mathcal S$, for example, via a Wald test at the level $\alpha$. Using a simple Bonferroni correction to correct for multiple testing, the overall size of the testing procedure is no larger than $\alpha \hat m$. Also note that $H_0: \bm{Q}_j^0 = \bm{Q}_{j+1}^0$ for $j=1,\ldots, m^0+1$ is testable via a Wald test. So if no changes in $\bm{\beta}_j^0, \bm{Q}_j^0$ are detected for two adjacent regimes, it means that all changes are coming from the individual effects, offering evidence for time-varying individual effects. Similarly, one can identify which change points pertain to the intercept alone by testing only the first restriction in $H_0: \bm{\beta}_j^0 = \bm{\beta}_{j+1}^0$, informing the researcher which periods actually need time dummies. {\color{black}If one is worried about post-model selection issues after the estimation of the number of change points, then one should impose more change points than found by the HQIC.}

Simulation Study

This section looks at the finite sample performance of $\hat{\mathbf T}_{\hat m}$, $ \hat{\bm{\beta}}_{OLS} \equiv \hat \bm{\beta}$, $ \hat{\bm{\beta}}_{FE} $ and $ \hat{\bm{\beta}}_{FFE} $. The data generating process (DGP) is based on model ((ref)) {\color{black}with fixed effects:} $c_{it} = c_i$. The idiosyncratic errors $\epsilon_{it}$ and $c_i$ are independently drawn from $N(0,1/4)$. A single regressor is generated as $x_{it}=\sqrt{2}c_i+z_{it}$, with $z_{it}\sim iid~N(0,1/2)$. The vector of slope parameters $\bm{\beta}^0 $ has elements alternating between $-0.1$ and $0.1$ for the different regimes between change points. For the case of one change point, we consider $T^0_1=2$ and $T^0_1=[T/3]$, where $[\cdot]$ is the least integer function; {\color{black}for two change points, we let $T_1^0=[T/3]$ and $T_2^0=[2T/3]$. We let $N=50,100,500$ and $T=20,30,50$. All results are reported for $1000$ replications.}

figure[figure omitted — 892 chars of source]
figure[figure omitted — 229 chars of source]

{\color{black}Figures (ref)-(ref) report histograms of the estimated change point locations, assuming that $m^0$ is known. Similarly, Tables (ref)-(ref) report slope estimators based on the true sample partition of the respective DGP. Since for $T^0_{1}=[T/3]$ the change point location increases proportionally with $T\in\{20,30,50\}$, the number of time periods before and after the change point are more balanced compared to the case with $T^0_{1}=2$. As a consequence, in {\color{black}the left panels of Figure (ref)}, the distribution of estimates is centered at the true change point $[T/3]$, while {\color{black}the right panels show a distribution that is skewed to the right for small $N$}. As $N$ increases, the distribution collapses over the true change point for both choices of $T^0_{1}$. For two change points, Figure (ref) shows that the estimated locations are also increasingly accurate as $N$ grows, as is expected from the consistency result of Theorem (ref).} (ref).}

\FloatBarrier {\color{black}For the same DGPs, Tables (ref) and (ref) report bias, standard error and mean squared error (MSE) of slope estimates based on OLS, FE and FFE estimates averaged over 1000 repetitions. Standard errors are calculated based on Theorem (ref).\footnote{{\color{black}Note that standard errors of the OLS estimators cannot be computed due to not observing the individual effects.}} Due to the dependence of $x_{it}$ on $c_i$, OLS estimators exhibit, as expected, a strong bias. Overall, FE and FFE estimators perform well even in small samples ($N=50$; $T=20$) with average bias close to zero and small standard errors. For both choices of change point locations, FFE estimators have smaller standard errors compared to FE, as expected from Theorem (ref).}heo3}.}

table[table omitted — 5,909 chars of source]
table[table omitted — 6,257 chars of source]

\FloatBarrier

figure[figure omitted — 1,273 chars of source]

Figure (ref) shows the estimated number of change points for the HQIC and BIC criteria defined in Section 2, and DGPs with $0,1$ or $2$ change points. {\color{black} Both HQIC and BIC perform well in large samples, but in smaller samples, the BIC strongly underestimates the number of change points}. For this reason, we use HQIC in both applications of Section (ref).

figure[figure omitted — 1,806 chars of source]
table[table omitted — 1,886 chars of source]

{\color{black}In Figure (ref), we compare the finite sample performance of the AGFL estimator of the number of change points in Qian and Su (2016) with our method for the DGP described above with two change points. To show how the finite sample performance of AGFL and our method depends on the degree of cross-sectional variation, we change the way the regressor is generated to $x_{it} = \sqrt{2}c_i + e_{it}$, with $e_{it} = w g_i + (1-w)\epsilon_{it}$, where $g_i \sim iid~ N(0,1/2)$ and $\epsilon_{it} \sim iid~ N(0,1/2)$. Note that, as before $c_i$ introduces endogeneity in $x_{it}$, while the new variable $g_i$ adds additional exogenous cross-sectional variation to $x_{it}$. The case $w=0$ reflects the usual DGP used throughout this section, which corresponds to $\approx 5 \%$ cross-sectional variation in $e_{it}$, while $w=0.1$ to $\approx 6.5 \%$ and $w=0.3$ to $\approx 20 \%$.\footnote{The cross-sectional variation is the $R$-squared from a regression of $e_{it}$ on the full set of individual dummies.} Figure (ref) shows that even in large samples ($N=500$), the AGFL is very sensitive to this moderate increase in cross-sectional variation, and tends to severely underestimate the number of change points, while our method remains mostly unaffected. Our method is therefore a useful alternative to AGFL for estimating multiple change points in short panels, where most of the variation in the data is in the cross-section dimension.}

{\color{black}Table (ref) contrasts the corresponding slope estimates for the two methods, and only for the cases that the two change points at $[T/3]$ and $[2T/3]$ are estimated correctly. The post AGFL slope estimators have higher bias and higher variance, although this could be due to less cases available for simulation averaging. Among the two estimators we propose, we see that as shown in Theorem (ref), the FFE estimator is more efficient. }

Since our two applications in the next section have sample sizes $\{T,N\}=\{19,106\}$ and $\{T,N\}=\{18,216\}$ and the second application has 15 regressors, we ran simulations of panels with similar properties. The results in Figures (ref) and (ref) report the estimated number of changes and the estimated location of the change point (when the number of changes is correctly estimated) for a change {\color{black}in each parameter} that varies from 0.01 to 0.02. Our method is able to detect a single change point of moderate size (above $0.015$) at the correct location ($T_1^0=14$) for most of the simulations.

figure[figure omitted — 600 chars of source]

Two Applications

Environmental Kuznets Curve

The environmental Kuznets curve (EKC) is often used to capture the relationship between income of a country and its emissions of chemicals such as carbon dioxide (CO$_2$). To detect changes in this relationship and in the emissions due to climate accords, we use yearly panel data on 106 countries and 19 years.\footnote{We use part of the dataset from Li, Qian and Su (2016) based on the World Bank Development Indicators (\href{https://data.worldbank.org/products/wdi}{https://data.worldbank.org/products/wdi}).} Countries with population less than five million and countries with missing observations are not used in our analysis. We start in 1992, because that is the year of the UN Framework Convention on Climate Change (UNFCCC), the first large international step to acknowledge climate change and to attempt to reduce emissions.\footnote{Before this period, several former communist countries would have to be excluded, leading to severe mismeasurement of emissions.}

equation[equation omitted — 172 chars of source]

Here, Emissions$_{it}$ is the logarithm of per capita CO$_2$ emissions in metric tones for country $i$ in year $t$, GDP$_{it}$ represents the logarithm of real gross domestic product in 2000 USD and Energy$_{it}$ is the logarithm of per capita consumption of energy measured in kilogram of oil equivalent. Energy consumption is included in several applications of the EKC with panel data (see Apergis and Payne 2009, Lean and Smyth 2010, Arouri et al 2012 and Farhani et al 2014). The term $c_{ij}$ reflects unobserved country-specific characteristics affecting CO$_2$ emissions such as its geography, resources, political developments, influence of environmental groups and industry composition, which are all likely to be correlated with income and/or energy consumption.

table[table omitted — 3,148 chars of source]

Our method finds three change points in 1997, 2004 and 2007.\footnote{Li et al. (2016) study the environmental Kuznets Curve using a similar dataset with an interactive fixed effects specification. Since our estimator only finds evidence for three change points between 1992 and 2010, an interactive fixed effects specification, i.e. $c_{it}=c_i f_t$, may not be desirable, unless it can be argued that despite this specification, there are only three changes in the pseudo-true parameters $\bm{\gamma}_j^0$.} All these changes can be traced back to steps in the Kyoto Protocol. At the beginning of our sample, the UNFCCC was adopted by 154 countries with the long-term aim of reducing global greenhouse gas emissions. The first major step in this convention was the Kyoto Protocol, adopted by consensus with more than 150 signatories on December 11, 1997. The Protocol included legally binding emissions targets for developed country parties for the six major greenhouse gases (including carbon dioxide). In 2004, Russia and Canada are the last to ratify the Kyoto Protocol, bringing the treaty into effect. In January 2008, the joint implementation mechanism starts.

In Table (ref) columns 1-4 report the corresponding FE estimates of slope coefficients and columns 5-7 their changes from one segment to the next.\footnote{Table (ref) also reports the Wald test of the $H_0$ hypothesis $\beta_j = \beta_{j-1}$. Moreover, as explained in Section (ref), the alternative slope estimator (FFE) relies on the assumption $c_{ij}=c_i$. Since several unobserved country characteristics such as the industry composition are likely to vary over segments $j$, and common shocks may hit countries at the same time in an unobserved way, as supported by the interactive fixed effects specification in Li, Qian and Su (2016), the FE estimator is the preferred choice over FFE in this application.} Columns 5-7 of Table (ref) reveal that the three changes are largely driven by changes in the coefficient of energy consumption, which decreases over the second and third segment, then increases back in the last. This could indicate that the Kyoto Protocol was initially successful in decreasing the elasticity of CO$_2$ emissions with respect to energy consumption. Based on estimates in the first sample segment (1992-1997), a 1$\%$ increase in per capita energy consumed leads to a 1.271$\%$ increase in CO$_2$ emissions per capita. The elasticity had decreased to 0.525$\%$ by the third segment (2005-2007). This decrease was followed by a significantly large increase of the elasticity to 0.981$\%$ in the last segment (2008-2010), reversing the decrease over the past 15 years to a large extent.

In summary, although the Kyoto Protocol is known to have no noticeable impact on global levels of carbon emissions (see, e.g. Helm, 2012), we find that in the course of its implementation, the elasticity of CO$_2$ emissions per capita with respect to energy consumption per capita underwent significant changes.

House Price Expectations in the U.S. after the Financial Crisis

We use data from a quarterly survey on 216 U.S. households to study house price expectations in the aftermath of the subprime mortgage crisis (2009-2013).\footnote{Our dataset is taken from the RAND American Life Panel (ALP), the Office of Federal Housing Enterprise Oversight (\href{http://www.fhfa.gov}{http://www.fhfa.gov}) and the Bureau of Labor Statistics (\href{http://www.bls.gov}{http://www.bls.gov}). We thank G. Niu and A. van Soest who kindly shared the data used in Niu and van Soest (2014), from which we extracted a balanced panel.} Every three months home owners stated their beliefs about the percentage chance that the value of their home will increase by the next year (0-100). We regress these expectations on household characteristics and state-level economic indicators.\footnote{A detailed description of the 15 covariates used can be found in Appendix B.}

The influence of unobserved characteristics of home owner $i$ such as optimism are captured by $c_{ij}$ in model (ref). These unobserved characteristics are likely to be correlated with some of the regressors. For instance, a home owner with an optimistic personality will have a more positive outlook on the price of his house but also on his subjective economic and financial well-being, which is one of the regressors of interest ('economic sentiment').

Our method finds a single change point in 2012Q2. In columns 2-7 of Table (ref) we report estimates of slope coefficients for both the FE and the FFE\ estimators. However, as in the first application, the two coefficient estimates exhibit large differences, leading us to conclude that the assumption of fixed effects might be violated in this setting as well. Therefore, we focus on the FE estimates.

Column 7 of Table (ref) shows the difference in estimated FE coefficients between the second and first segment ($\hat \beta_2 - \hat \beta_1$). Evidently, the change point is primarily driven by differences in coefficients of the variables 'Change in local house prices', being female, the indicator for living in Arizona, California, Florida and Nevada ('Sand state'), and the health of the home owner.\footnote{The Wald test (Table (ref)) of the $H_0$ hypothesis $\beta_1 = \beta_2$ confirms that there is a significant change in the vector of slope coefficients across the two segments for both FFE and FE estimates.} Interestingly, two of these regressors do not vary over time ('Female' and 'Sand state'), yet their coefficient changes.

The period before the change point (2009 - 2012Q2) represents the direct aftermath of the financial crisis when economic uncertainty was high. In this segment we find a significant positive effect of home owner's subjective economic and financial well-being ('economic sentiment') on house price expectations. In 2012Q3 the Federal Reserve announced its third round of quantitative easing, which was an open-ended bond purchasing program of agency mortgage-backed securities and it was announced that the federal funds rate would be likely maintained near zero for at least the next three years. Overall, uncertainty in the market decreased and the steady recovery of the U.S. housing market began. Our results in Table (ref) suggest that during this recovery period home owners looked for more objective measures of economic performance (such as changes in state-level unemployment rates and state-level house prices) to infer their house values, while previously they relied more on subjective assessments.

table[table omitted — 3,410 chars of source]

\FloatBarrier

Conclusion

In this paper, we proposed a method for estimating short panels subject to multiple change points and unobserved, possibly time-varying individual effects. We propose first estimating by OLS all the change points that occur in the pseudo-true parameters, under relatively general time-variation in the individual effects. Next, we assume that the individual effects either only change at these identified change points, or remain constant over the sample, and contrast the asymptotic properties of two consistent slope estimators, helping to identify the number and location of changes in the slope parameters. We demonstrate the usefulness of our method via two applications: the enviromental Kuznets curve and house price expectations.

Our method can also be used as a diagnostic tool for model specification in short panels: if changes are found at each point in time, the specification should be revisited for more parsimonious modelling of time-variation, and if changes occur rarely, then change point modelling is a better alternative.

Bibliography

Apergis, N., and Payne, J. E. (2009). {CO2 emissions, energy usage, and output in Central America.} Energy Policy 37, 3282-3286.\\ Arouri, M. E. H., Youssef, A. B., M'henni, H., and Rault, C. (2012). {Energy consumption, economic growth and CO2 emissions in Middle East and North African countries.} Energy Policy 45, 342-349.\\ Armona, R., Fuster, A. and Zafar, B. (2018). {Home price expectations and behavior: evidence from a randomized information experiment.} Review of Economic Studies, forthcoming.\\ Aue, A. and Horv\'{a}th, L. (2013). {Structural breaks in time series}, Journal of Time Series Analysis 34, 1-16.\\ Bai, J. (2009). {Panel data models with interactive fixed effects.} Econometrica 77, 1229-1279.\\ Bai, J. (2010). {Common breaks in means and variances for panel data.} Journal of Econometrics 157, 78-92.\\ Bai, J. and Li, K. (2014). {Theory and methods of panel data models with interactive effects.} \textit{Annals of Statistics} 42, 142-170.\\ Bai, J., and Perron, P. (1998). {Estimating and testing linear models With multiple structural changes.} \textit{Econometrica} 66, 47-78.\\ Baier, S. L., and Bergstrand, J. H. (2007). {Do free trade agreements actually increase members' international trade?.} \textit{Journal of International Economics} 71(1), 72-95.\\ Baltagi, B.H., Feng, Q. and Kao, C. (2016). {Estimation of heterogeneous panels with structural breaks}, \textit{Journal of Econometrics} 191, 176-195.\\ Baltagi, B.H., Kao, C. and Liu, L. (2017). {Estimation and identification of change points in panel models with nonstationary or stationary regressors and error term}, \textit{Econometric Reviews 36: 85-102.}\\ Bardwell, L., Fearnhead, P., Eckley, I.A., Smith, S. and Spott, M. (2018). {Most recent changepoint detection in panel data.} \textit{Technometrics}, \href{https://doi.org/10.1080/00401706.2018.1438926}{https://doi.org/10.1080/00401706.2018.1438926}/\\ Blanco, L. and Ruiz, I. (2013). {The Impact of Crime and Insecurity on Trust in Democracy and Institutions,} \textit{American Economic Review: Papers and Proceedings} 103, 284-288.\\ B�rsch-Supan, A., Brandt, M., Hunkler, C., Kneip, T., Korbmacher, J., Malter, F., Schaan, B., Stuck, S. and Zuber, S. (2013). {Data resource profile: the Survey of Health, Ageing and Retirement in Europe (SHARE).} \textit{International Journal of Epidemiology} 42, 992-1001.\\ Chan, J., Horv\'{a}th, L., and Hu\v{s}kov\'{a}, M. (2013). {Darling-Erd\H{o}s limit results for change point detection in panel data.} \textit{Journal of Statistical Planning and Inference} 143, 955-970.\\ Chan, N. H., Yau, C. Y., and Zhang, R. (2014). {Group LASSO for Structural Break Time Series.} \textit{Journal of American Statistical Association} 109, 590-599.\\ Cs\"{o}rg\"{o}, M., and Horv\'{a}th, L. (1997). {Limit theorems in change point Analysis.} \textit{Wiley Series in Probability and Statistics} Vol. 18. John Wiley & Sons Inc.\\ Cho. H. (2016). {change point detection in panel data via double CUSUM statistic.} \textit{Electronic Journal of Statistics} 10, 2000-2038.\\ Cho, H. and Fryzlewicz, P. (2015). {Multiple change point detection for high dimensional time series via sparsified binary segmentation.} \textit{ Journal of the Royal Statistical Society B} 77, 475-507.\\ de Wachter, S. and Tsavalis, E. (2012). {Detection of structural breaks in linear dynamic panel data models}, \textit{Computational Statistics and Data Analysis} 56, 3020-3034. \\ Emerson, J. and Kao, C. (2001). {Testing for structural change of a time trend regression in panel data: Part I.} \textit{Journal of Propagations in Probability and Statistics} 2, 57-75.\\ Emerson, J. and Kao, C. (2002). {Testing for structural change of a time trend regression in panel data: Part II.} \textit{Journal of Propagations in Probability and Statistics} 2, 207-250.\\ Farhani, S., Mrizak, S., Chaibi, A., and Rault, C. (2014). {The environmental Kuznets curve and sustainability: A panel data analysis.} \textit{Energy Policy} 71, 189-198.\\ Geng, N. and van Soest, A. H. O. (2014). {House price expectations.} \textit{IZA Working Paper 8536}, \href{http://ftp.iza.org/dp8536.pdf}{http://ftp.iza.org/dp8536.pdf}.\\ Hall, A. R., Osborn, D. and Sakkas, N. (2013). {Inference about structural breaks using information criteria}, \textit{The Manchester School} 81, 54-81.\\ Harchaoui, Z., and L\'{e}vy-Leduc, C. (2010). {Multiple change point estimation with a total variation enalty.} \textit{Journal of the American Statistical Association} 105, 1480-1493.\\ Helm, D. (2012). {Climate policy: the Kyoto approach has failed.} \textit{Nature} 491, 663-665. \\ Horv\'ath, L., and Hu\v{s}kov\'a, M. (2012). {change point detection in panel data.} \textit{Journal of Time Series Analysis} 33, 631-648.\\ Kim, D. (2011). {Estimating a common deterministic time trend break in large panels with cross sectional dependence.} \textit{Journal of Econometrics} 164(2), 310-330.\\ Kyle, M. and Williams, H. (2016). {Is American health care uniquely inefficient? Evidence from prescription drugs.} \textit{American Economic Review} 107, 486-490.\\ Lean, H. H., and Smyth, R. (2010). {CO2 emissions, electricity consumption and output in ASEAN.} \textit{Applied Energy} 87, 1858-1864.\\ Li, D., Qian, J. and Su, L. (2016). {Panel data models with interactive fixed effects and multiple structural breaks.} \textit{Journal of the American Statistical Association}, 516: 1804-1819.\\ Moon, H. R., and Weidner, M. (2015). {Linear regression for panel With unknown number of factors as interactive fixed effects.} Econometrica 83, 1543-1579.\\ Nimomiya, Y. (2005). {Information criterion for Gaussian change point model.} \textit{Statistics and Probability Letters} 72, 237-247.\\ Perron, P. and Yamamoto, Y. (2015). {Using OLS to estimate and test for structural changes in models with endogenous regressors.} \textit{Journal of Applied Econometrics} 28, 119-144. \\ Pesaran, H. (2006). {Estimation and inference in large heterogeneous panels with a multifactor error structure.} \textit{Econometrica} 74, 967-1012.\\ Okuy, R. and Wang, W. (2018). {"Heterogeneous structural breaks in panel data models.} \textit{Working Paper}, \href{"Heterogeneous structural breaks in panel data models}, \href{https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3031689}{https://papers.ssrn.com/sol3/papers.cfm?abstract$\_$id=3031689}.\\ Qian, J. and Su, L. (2016). {Shrinkage estimation of common breaks in panel data models via adaptive group fused LASSO.} \textit{Journal of Econometrics} 191, 86-109.\\ Qu, Z., and Perron, P. (2007). {Estimating and testing structural changes in multivariate regressions.} \textit{Econometrica} 75, 459-502.\\ Torgovitski, L. (2015). {Panel data segmentation under finite time horizon.} \textit{Journal of Statistical Planning and Inference} 167, 69-89.\\ Vert, J.-P. and Bleakley, K. (2010). {Fast detection of multiple change points shared by many signals using group LARS.} \textit{Proceedings of Advances in Neural Information Processing Systems 23.}\\