EconBase
← Back to paper

Asymptotic properties of Bayesian inference in linear regression with a structural break

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

75,304 characters · 14 sections · 43 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Asymptotic properties of Bayesian inference in linear regression with a structural break

\thanksmarkseries{alph}

The paper was previously titled “Structural break in linear regression models: Bayesian asymptotic analysis” and “Bayesian inference in linear regression with a structural break." } }

\affil{{ Adam Smith Business School \\ University of Glasgow \\ University Ave, Glasgow, G12 8QQ, United Kingdom }}

titlingpage\usethanksrule \begin{abstract} This paper studies large sample properties of a Bayesian approach to inference about slope parameters $\gamma$ in linear regression models with a structural break. In contrast to the conventional approach to inference about $\gamma$ that does not take into account the uncertainty of the unknown break location $\tau$, the Bayesian approach that we consider incorporates such uncertainty. Our main theoretical contribution is a Bernstein-von Mises type theorem (Bayesian asymptotic normality) for $\gamma$ under a wide class of priors, which essentially indicates an asymptotic equivalence between the conventional frequentist and Bayesian inference. Consequently, a frequentist researcher could look at credible intervals of $\gamma$ to check robustness with respect to the uncertainty of $\tau$. Simulation studies show that the conventional confidence intervals of $\gamma$ tend to undercover in finite samples whereas the credible intervals offer more reasonable coverages in general. As the sample size increases, the two methods coincide, as predicted from our theoretical conclusion. Using data from Paye and Timmermann (2006) on stock return prediction, we illustrate that the traditional confidence intervals on $\gamma$ might underrepresent the true sampling uncertainty. \end{abstract}

\onehalfspacing

Introduction

We consider the linear regression with a structural break, following the notations of bai1997:

equation[equation omitted — 245 chars of source]

where $w_t$ and $z_t$ are $d_w \times 1$ and $d_z \times 1$ vectors of covariates, and the random variable $\epsilon_t$ is a regression error. $\floor*{a}$ is the largest integer that is strictly smaller than $a$. The relationship between the outcome $y_t$ and the covariate $z_t$, measured by $\delta$'s, changes across regimes, which are defined by the break location parameter $\tau \in (0,1)$. There can be another set of covariates $w_t$ whose relationship with $y_t$, measured by $\alpha$, stays unchanged across the regimes. The unknown parameters include the break location $\tau$ as well as the slope parameters $\gamma=(\alpha,\delta_1,\delta_2)$. The focus of the current study is on inference about the slope parameter $\gamma$\footnote{ For an extensive review of important aspects in structural break models such as estimation and inference of the number of breaks as well as break locations, see perron2006. }.

The classic literature

In the literature, the conventional least-squares estimators $\left( \hat{\tau}_{LS}, \hat{\gamma}_{LS} \right)$ for $\left( \tau, \gamma \right)$ are computed as follows: for each candidate $\tau$, compute the sum of squared residuals of the regression and denote the minimizing choice by $\hat{\tau}_{LS}$. Plug in the value $\tau = \hat{\tau}_{LS}$ in the model and define $\hat{\gamma}_{LS}=\hat{\gamma}(\hat{\tau}_{LS})$, where $\hat{\gamma}({\tau})$ is the usual OLS estimator of $\gamma$ assuming the break location $\tau$. bai1997 assumes that the true jump size $\delta_0$ is either fixed or shrinks to zero as $T \to \infty$, but at a rate slower than $\sqrt{T} \to \infty$. Bai shows that $\hat{\tau}_{LS}$ converges at the rate $T^{-1}$ in the former case and, in the latter case, finds an asymptotic distribution of $\hat{\tau}_{LS}$ that can be used for constructing confidence intervals for $\tau$. In both cases, Bai proves that the asymptotic distribution of $\hat{\gamma}_{LS}$ is the same as that of $\hat{\gamma}(\tau_0)$, where $\tau_0$ is the true value of $\tau$. This means that one can ignore the very problem of unknown $\tau$ when making inference on $\gamma$.

Figure (ref) displays finite-sample distributions of $\hat{\tau}_{LS}$ (blue solid curves) which are produced based on 1,000 repeated experiments on the following model $ y_t = \delta_0 1\left(t>\floor*{\tau_0T} \right)+\epsilon_t$\footnote{ Figure (ref) also shows distributions (red solid curve with small circles) of a Bayesian point estimator of $\tau$, the posterior mode. We later show that the posterior mode converges to the same limiting distribution as $\hat{\tau}_{LS}$. }. Note that despite the $T$-consistency, $\hat{\tau}_{LS}$ displays significant variation, especially when the true break size $\delta_0$ is small\footnote{In addition, the distributions exhibit three modes as reported in the literature (e.g.,\ baek2021; casini_perron2021continuous). }. In practice, the conventional approach to inference on the slope parameters $\gamma$ would ignore this uncertainty, neglecting all possible values of $\tau$ other than $\hat{\tau}_{LS}$. As a consequence, the corresponding confidence intervals on $\gamma$ tend to undercover since it might not be the case that $\hat{\tau}_{LS} = \tau_0$ in a given sample (See our simulation in Section (ref)).

\FloatBarrier \graphicspath{{Figures_final/tau_posterior/}}

figure[figure omitted — 970 chars of source]

Bayesian perspective

For a Bayesian, this non-standard estimation problem\footnote{ Estimation of structural break models is considered non-standard in a sense that there is a non-regular parameter (e.g.,\ break location) whose point estimator converges faster than $T^{-1/2}$, the rate at which the regular parameters (e.g.,\ slope coefficients) converge. } can be dealt with by placing prior on both $\tau$ and $\gamma$ and by computing corresponding posterior probabilities. The uncertainty of $\tau$ will be automatically reflected on the marginal posterior probability of $\gamma$. This is because the posterior distribution of $\gamma$ given the data $\bm{D}_T$ can be written as a mixture where the weights correspond to the marginal posterior density $\pi_T(\tau)$ for $\tau$:

equation[equation omitted — 130 chars of source]

where $p\left( \gamma |\tau,\bm{D}_T \right) $ is the posterior conditional distribution of $\gamma$ given $\tau$. The posterior density $\pi_T(\tau)$ reflects the uncertainty of $\tau$ given the data set. Figure (ref) shows three realizations of $\pi_T(\tau)$ (gray dashed curves) which are randomly chosen out of the 1,000 repetitions. Compared to the conventional approach, the key difference is that the Bayesian approach (ref) incorporates all possibilities of $\tau$ (not just $\hat{\tau}_{LS}$) and weights them according to the posterior density. As we see in simulation studies, this results in longer lengths of Bayesian credible intervals of $\gamma$ compared to the conventional counterparts. Consequently, the credible intervals tend to avoid undercoverage. See Section (ref) for further discussion. Note that, unlike conventional frequentist methods, Bayesian inference has a valid interpretation even in finite samples as it does not rely on asymptotic theory.

In this study, we examine the asymptotic behavior of Bayesian estimation of the considered model under the fixed jump size framework. Specifically, we prove a Bernstein-von Mises type theorem for the slope parameters $\gamma$ which validates a frequentist interpretation of Bayesian credible regions. A Bayesian researcher can invoke our theorem to convey statistical results to frequentist researchers. A frequentist researcher could look at the credible interval of $\gamma$ to check robustness with respect to the uncertainty of the break location. Such sensitivity analysis is reasonable as our result guarantees the credible interval to converge to the conventional confidence interval. We first establish theoretical results under normal likelihood and natural conjugate prior. We further extend the results to non-conjugate priors using Laplace approximations.

The literature on the theoretical properties of Bayesian approaches in non-regular models such as (ref) is very scarce despite their popularity in applications. To our knowledge, frequentist properties of the Bayesian approach for linear regression models with structural breaks have not been studied in the literature. ghosal_samanta1995 consider a general non-regular estimation problem from a Bayesian perspective and establish conditions under which the Bernstein-von Mises theorem holds for the regular part of the parameter. However, their assumptions are difficult to verify in regard to our model in consideration.

Recently, casini_perron2020generalized propose a generalized Laplace estimator of the break location $\tau$ which is defined by an integration rather than an optimization. Their approach provides a better approximation about the uncertainty in $\tau$ than the conventional method. Although our focus of the current paper is on inference about the slope coefficients $\gamma$ and not $\tau$, our Bayesian approach toward inference shares the same spirit; any statement about $\gamma$ is expressed as a weighted average (ref) over the marginal posterior density of $\tau$.

The paper is organized as follows. Section (ref) introduces the model and lists a set of assumptions. Section (ref) introduces a Bayesian approach based on normal likelihood and conjugate prior. The section then establishes frequentist properties of the approach. Section (ref) extends the results to non-conjugate priors. Section (ref) presents simulation evidence to assess the adequacy of the asymptotic theory and to illustrate that conventional confidence intervals on the slope parameters tend to undercover. Section (ref) reports an empirical application to the stock return prediction model of paye_timmermann2006. Section (ref) concludes the paper. The mathematical proofs and derivations are listed in the Appendix. Additional tables are provided in the online appendix.

The model and data generating process

The model

Using the reparametrization $x_t=(w_t',z_t')'$, $\beta=(\alpha',\delta_1')'$, and $\delta=\delta_2-\delta_1$, the equations (ref) can be rewritten as

equation[equation omitted — 229 chars of source]

Note that $z_t$ is a subvector of $x_t$. More generally, let $z_t=R'x_t$, where $R$ is a $d_x \times d_z$ known matrix with full column rank and hence $z_t$ is defined as a linear transformation of $x_t$. For $R=(0_{d_z\times d_w},I_{d_z})'$, we obtain model (ref). For $R=I_{d_x}$, a pure change model is obtained. To rewrite the model in matrix form, we introduce further notations. Define $Y=(y_1,\ldots,y_T)'$, $\epsilon=(\epsilon_1,\ldots,\epsilon_T)'$, $X=(x_1,\ldots,x_T)'$, $X_{1\tau}=(x_1,\ldots,x_{\floor*{\tau T} },0,\ldots,0)'$, $X_{2\tau}=(0,\ldots,0,x_{\floor*{\tau T} +1},\ldots,x_T)'$. Define $Z, Z_{1\tau}$, and $Z_{2\tau}$ similarly. Then, $Z=XR$, $Z_{1\tau}=X_{1\tau}R$, and $Z_{2\tau}=X_{2\tau}R$. Now, the equations (ref) can be written as

equation[equation omitted — 99 chars of source]

where $\chi_\tau = (X, Z_{2\tau})$ and $\gamma=(\beta',\delta')'$. $S_T(\tau)$ denotes the sum of squared residuals of the regression (ref) given $\tau$. Let $\mathcal{H} \subset (0,1)$ be the space of the break locations. The least-squares estimator of $\tau$ is defined as

equation[equation omitted — 109 chars of source]

and the least-squares estimator for the slope coefficients $\gamma=(\beta',\delta')'$ is

equation[equation omitted — 86 chars of source]

where $\hat{\gamma}(\tau)$ denotes the usual OLS estimator given the value of $\tau$.

Data generating process

The data are assumed to include $T$ observations on a response and a vector of covariates: $\bm{D}_T=( \bm{Y}_T, \bm{X}_T) = (y_1,\ldots,y_T,x_1,\ldots,x_T)$ where $y_t \in \mathbb{R}$ and $x_t \in \mathcal{X} \subset \mathbb{R}^{d_x}$, $t=1,...,T$. $\mathcal{X}$ is assumed to be a convex and bounded set. Conditional on $ \bm{X}_T $, the response is generated according to model (ref) with the true parameters $(\gamma_0',\sigma_0^2, \tau_0)'$. We use $\theta=(\gamma',\sigma^2)'$ to denote the regression parameters. We make the following assumptions about the true data-generating-process (DGP):

assumption\quad \begin{enumerate}[label=(\roman*)] • $\delta_0 \ne 0$. • $\epsilon_t$ is i.i.d. with $E(\epsilon_t|x_t)=0$, $E(\epsilon_t^2|x_t) =\sigma_0^2$, where $\sigma_0^2$ is unknown to the econometrician. • $\Sigma_X =E[x_tx_t']= \operatorname*{plim} \frac{1}{T} \sum_{t=1}^T x_t x_t'$ exists and is positive definite. • For all $\tau_1,\tau_2 \in (0,1)$ with $\tau_1 < \tau_2$, $\frac{1}{T} \sum_{\floor*{\tau_1T}+1}^{\floor*{\tau_2T}} x_t\epsilon_t =O_p(T^{-1/2})$ and $\frac{1}{T} \sum_{\floor*{\tau_1T}+1}^{\floor*{\tau_2T}} x_tx_t' =(\tau_2-\tau_1)\Sigma_X+O_p(T^{-1/2})$ \end{enumerate}

Under the above assumptions, the classical theoretical results apply. bai1997 shows that the convergence rate of $\hat{\tau}_{LS}$ is $T^{-1}$ if $\delta_0$ is fixed with respect to the sample size: \[ \hat{\tau}_{LS} =\tau_0 + O_p(T^{-1}), \] and that the least-squares estimator for $\gamma$ is asymptotically normal with the asymptotic covariance matrix being the same as if $\tau_0$ is known:

equation[equation omitted — 149 chars of source]

where \[ V =\operatorname*{plim} T^{-1}

pmatrix[pmatrix omitted — 159 chars of source]

= \operatorname*{plim} T^{-1} \chi_{\tau_0}'\chi_{\tau_0}. \] This means that $\tau$ can be treated as known for the purpose of inference about $\gamma$. In other words, the uncertainty of the break location is essentially ignored, and thus the confidence interval for $\gamma$ tends to undercover (see Section (ref) for simulation).

There are several comments on Assumption (ref). In threshold regression models (see hansen2000), the threshold variable is often one of the regressors. In this case, sorting the threshold variable leads to a trend in the regressors, which requires an alternative approach for the asymptotic analysis. We do not consider the case with one of the regressors being the threshold variable in this paper. In addition, we require the regression errors to be i.i.d.\ with variance $\sigma^2$. Adding more flexibility such as heteroscedasticity and serial correlation would be an important future direction.

A Bayesian approach under normal likelihood and conjugate prior

The distribution of covariates is assumed to be ancillary and it is not modeled. Throughout this paper, we assume the normal likelihood function\footnote{ Similarly, qu_perron2007 propose a quasi-maximum likelihood estimator assuming normal errors. }

equation[equation omitted — 207 chars of source]

where $\chi_{\tau,t}'$ is the $t$th row of the matrix $\chi_{\tau}$. Note that the normality is not assumed for the true DGP, so the model can be mis-specified.

The break location $\tau$ and the regression parameters $\theta$ are independent a-priori and the prior on $\theta$ is the natural conjugate prior. That is, $ \pi \left(\gamma,\sigma^2, \tau \right) = \pi(\gamma | \sigma^2) \pi(\sigma^2) \pi(\tau)$ where the prior on $\gamma$ conditional on $\sigma^2$ is normal $N_{(d_x+d_z)}(\underline{\mu}, \sigma^2 \underline{H}^{-1}) $ and the prior on $\sigma^2$ is inverse-gamma $InvGamma(\underline{a},\underline{b}) $. Note that by taking $\underline{H} \to 0$, $\underline{a} \to -(d_x+d_z)/2$, and $\underline{b} \to 0$, we have the uninformative improper prior $ \pi \left(\gamma,\sigma^2 \right) \propto \sigma^{-2}$ as a special case. The prior on $\tau$ can be of any form as long as it is positive at $\tau_0$, and $\pi(\tau)$ is finite for all $\tau \in \mathcal{H} $.

The conjugate prior is a popular choice in the Bayesian estimation of linear regression models. Our restriction on the prior for the break location $\tau$ is very mild. For example, the uniform distribution on $\mathcal{H}$ satisfies the requirement. Recently, baek2021 investigates the same model (ref). As the distribution of $\hat{\tau}_{LS}$ might exhibit tri-modality for small jumps, Baek proposes a new point estimator for $\tau$ based on a modified objective function. The proposed modification can be regarded as equivalent to specifying a certain type of prior for $\tau$ and indeed such prior satisfies our restriction.

Under the normal likelihood function and the prior defined above, the posterior distributions are

align[align omitted — 462 chars of source]

where $ \bar{H}_\tau = \underline{H} + \chi_\tau'\chi_\tau$, $\bar{\mu}_{\tau} = \bar{H}_\tau^{-1} \left[ \underline{H}\underline{\mu}+\chi_\tau' Y \right]$, $\bar{b}_\tau = \underline{b} +0.5\left[ \underline{\mu}' \underline{H} \underline{\mu} + Y'Y - \bar{\mu}_\tau' \bar{H}_\tau \bar{\mu}_\tau \right]$, and $\bar{a} = \underline{a}+ T/2$, and $t_k(v,\mu,\Sigma)$ is the $k$-dimensional t-distribution with $v$ degrees of freedom, a location vector $\mu \in \mathbb{R}^k$, and a $k \times k$ shape matrix $\Sigma$. See Appendix (ref) for the derivation.

Due to the availability of the closed-forms for the conditional posteriors given $\tau$, the posterior sampling is simple and fast. One can first draw $\tau_{(1)},\ldots,\tau_{(S)}$ from the marginal posterior of $\tau$ as in (ref) via, for example, the Metropolis-Hastings algorithm, where $S$ is the number of posterior draws. For each $\tau_{(s)}$, one can sample posterior draws of $\sigma^2_{(s)}$ from the posterior conditional on $\tau=\tau_{(s)}$, namely (ref). Conditional on $\tau$ and $\sigma^2$, one can draw $\gamma$ from $p(\gamma | \sigma^2, \tau, \bm{D}_T )$\footnote{ It can be shown that $\gamma \big| \sigma^2, \tau, \bm{D}_T \sim N_{(d_x+d_z)} \left( \bar{\mu}_\tau, \sigma^2 \bar{H}_\tau^{-1} \right) $}. For example, a laptop with a 2.2GHz processor and 8GB RAM takes about 4.1 seconds to draw 10,000 posterior draws in an empirical example in Section (ref) that has ten slope coefficients in total.

Asymptotic theory

We investigate the asymptotic behavior of the Bayesian method under the normal likelihood and the conjugate prior defined above. We do so in two steps. Section (ref) shows that the marginal posterior of the break location $\tau$ contracts to the true value $\tau_0$ at the rate of $T^{-1}$, the same rate at which the least-squares estimator $\hat{\tau}_{LS}$ converges. The proof is based on studying the behavior of the log ratio of the marginal posterior densities of $\tau$. In addition, we establish the limiting distribution of the posterior mode of $\tau$. Section (ref) establishes a Bernstein-von Mises type theorem for the regression slope coefficients $\gamma$. The proof is based on the $T$-consistency of the marginal posterior of $\tau$ and the fact that the conditional posterior for $\sqrt{T}\left( \gamma - \hat{\gamma}_{LS} \right)$ given $\tau$ is asymptotically normal. Proofs of the theorems can be found in Appendix (ref).

Marginal posterior of $\tau$

An intermediate step for proving the Bernstein-von Mises theorem is the marginal posterior consistency of $\tau$ at rate $T^{-1}$. Marginal posteriors have not been studied extensively or systematically in the literature. Here, we directly analyze the form of the marginal posterior of $\tau$. Let $L_T(\tau)$ be the marginal likelihood conditional on $\tau$, that is \[ L_T(\tau)= \int p(\bm{Y}_T |\bm{X}_T, \theta,\tau) \pi(\theta, \tau) d\theta, \] which is available up to a multiplicative constant under the normal likelihood and the conjugate prior as can be seen in (ref). The marginal posterior density $\pi_T(\tau) $ of $\tau$ is defined as \[ \pi_T(\tau) = \frac{ L_T(\tau) }{ \int L_T(\tau) d\tau }. \]

The following theorem establishes the first step for proving the Bernstein-von Mises theorem, the $T$-consistency of the marginal posterior of $\tau$. It states that the posterior mass outside of a ball around $\tau_0$ with radius proportional to $T^{-1}$ will be asymptotically negligible.

theorem[Marginal posterior consistency of $\tau$ at rate $T^{-1}$] Suppose Assumption (ref) holds. Then, under the normal likelihood and the conjugate prior described above, $\forall \eta>0, \epsilon>0$, $\exists M>0$ and $k>0$ such that $T \geq k \implies$ \[ P_{\theta_0, \tau_0} \left( \int_{B^c_{M/T} (\tau_0)} \pi_T(\tau) d\tau < \eta \right) > 1-\epsilon, \] where for any constant $d>0$, $B^c_{d}(\tau_0)$ denotes the set difference $\mathcal{H} \setminus (\tau_0-d, \tau_0+d)$.

The proof of Theorem 1 is built on some intermediate steps, Propositions (ref)-(ref). It can be shown that $ \int_{ B^c_{M/T}(\tau_0) } \pi_T(\tau) d\tau $ is bounded by the product of $ \int_{ B^c_{M/T}(\tau_0) } \frac{ L_T( \tau) }{ L_T(\tau_0) } d\tau $ and the inverse of $ \int_{ B^c_{M_0/T} (\tau_0) } \frac{ L_T(\tau') }{ L_T( \tau_0 )} d\tau' $ for each $T$ and for any $M_0>0$. Proposition (ref) shows that under the normal likelihood and the conjugate prior, due to the availability of the marginal likelihood conditional on $\tau$ up to a normalization constant as in (ref), studying the log marginal likelihood ratio boils down to comparing the sum of squared residuals $S_T(\tau)$. Proposition (ref) establishes the probability limit of $T^{-1} S_T(\tau)$, for which we show examples in Figure (ref). We then show that the limit of $T^{-1} S_T(\tau)$ achieves a unique minimum at $\tau_0$ (Proposition (ref)), and study the modulus of continuity of an appropriate empirical process (Proposition (ref)) in order to derive bounds. The detail of the proof of Theorem 1 can be found in Appendix (ref). \FloatBarrier \graphicspath{{Figures_final/}}

figure[figure omitted — 297 chars of source]

\FloatBarrier

The Bayesian counterpart of the least-squares estimator $\hat{\tau}_{LS}$ would be the posterior mode: \[ \hat{\tau}_{Bayes} = \operatorname*{arg\,max}_{\tau \in \mathcal{H}} \pi_T(\tau). \] bai1997 shows that $ \operatorname*{arg\,max}_{m} W^*(m) $ is the asymptotic distribution of $\hat{\tau}_{LS}$ \footnote{ $W^*(m)$ is a stochastic process defined on the set of integers as follows: $W^*(0)=0$, $W^*(m)=W_1(m)$ for $m<0$, and $W^*(m)=W_2(m)$ for $m>0$, with

alignat{4} W_1(m)& = -\delta_0 \sum_{i=m+1}^0 z_iz_i' \delta_0 &&+ 2\delta_0 \sum_{i=m+1}^0 z_i\epsilon_i, && for m=-1,-2,... \nonumber \\ W_2(m)& = -\delta_0 \sum_{i=1}^m z_iz_i' \delta_0 &&- 2\delta_0 \sum_{i=1}^m z_i\epsilon_i, && for m=1,2,... \nonumber

}. A consequence of the proof of Theorem (ref) is that $\hat{\tau}_{Bayes} $ converges to the same limiting distribution. See Appendix (ref) for a proof.

cor[Limiting distribution of the posterior mode of $\tau$] Suppose Assumption (ref) holds. Then, under the normal likelihood and the conjugate prior described above, \[ \floor*{T\left( \hat{\tau}_{Bayes} - \tau_0 \right) } \overset{d}{\to} \operatorname*{arg\,max}_{m} W^*(m) . \]

Bernstein-von Mises Theorem for $\gamma$

The marginal posterior of $\gamma$ is a mixture with weights corresponding to the marginal posterior density $\pi_T(\tau) $. Furthermore, due to Theorem (ref), we can focus our attention on the values of $\tau$ in a $T^{-1}$ neighborhood of $\tau_0$: \[ \int p( \gamma |\tau,\bm{D}_T ) \pi_T(\tau) d\tau = \int_{B_{M/T}(\tau_0)} p( \gamma|\tau,\bm{D}_T) \pi_T(\tau) d\tau +o_p(1). \] We are now ready to establish the Bernstein-von Mises type result.

theorem[Bernstein-von Mises theorem for the slope coefficients] Suppose Assumption (ref) holds. Then, under the normal likelihood and the conjugate prior described above, \[ d_{TV} \left( \pi \left[ \begin{matrix} \sqrt{T}\left( \gamma- \hat{\gamma}_{LS} \right) \end{matrix} \bigg| \bm{D}_T\right], N_{(d_x+d_z)}\left( 0, \sigma^2_0V^{-1} \right) \right) \to 0 , \] in $P_{\theta_0, \tau_0}-probability$ where $d_{TV}$ is the total variation distance.

The proof of Theorem 2 exploits the fact that the conditional posterior for $\sqrt{T}\left( \gamma - \hat{\gamma}_{LS} \right)$ given $\tau$ is asymptotically normal, which is close to the asymptotic distribution of $ \hat{\gamma}_{LS}$ when $\tau$ is close to $\tau_0$. A bound on the Kullback–Leibler (KL) divergence between two normal densities together with the $T$-consistency is used to make the argument precise. The proof is presented in Appendix (ref).

An extension to non-conjugate priors

The previous section establishes the asymptotic properties of the posterior distributions under the conjugate prior. A natural question is whether these results can be extended to other priors. For example, an independent prior between the slope coefficients $\gamma$ and the error variance $\sigma^2$, e.g.,\ $\pi(\gamma,\sigma^2)=\pi(\gamma)\pi(\sigma^2)$ with $\gamma \sim N_{(d_x+d_z)}(\underline{\mu}, \underline{\Sigma})$ and $\sigma^2 \sim InvGamma(\underline{a},\underline{b})$, is a popular choice for the Bayesian estimation of regression models in practice. Under the normal likelihood and the conjugate prior, the analytical expressions of the marginal posterior of $\tau$ up to a normalization constant (ref) and the conditional posterior of $\gamma$ given $\tau$ (ref) facilitate the theoretical analysis. They are not available, for instance, under the independent prior mentioned above. In this section, we extend the theoretical results by keeping the normal likelihood (ref) but without requiring the conjugate prior on $\theta$. In order to study the asymptotic behavior of the posterior distributions without having their closed-form expressions, we employ Laplace approximation type results in hong_preston2012. To do so, we make an additional assumption as shown below. Let $\hat{\theta}(\tau)$ be the maximum likelihood estimator of $\theta$ conditional on $\tau \in \mathcal{H}$, i.e.,\ $\hat{\theta}(\tau)=\arg\sup_{\theta \in \Theta} \log p(\bm{Y}_T | \bm{X}_T, \theta, \tau)$. Denote by $\theta^*(\tau)$ the corresponding pseudo true parameter value that minimizes the KL divergence between the model $p(\bm{Y}_T | \bm{X}_T, \theta, \tau)$ and the DGP.

assumption\quad \begin{enumerate}[label=(\roman*)] • There is a compact convex subset $\Theta$ of $\mathbb{R}^{d_x+d_z+1}$ such that $\theta^*(\tau) \in int\left(\Theta \right)$ for all $\tau \in \mathcal{H}$. • The prior $\pi(\theta,\tau)$ is supported on $\Theta \times \mathcal{H}$. It is continuous in $\theta$ and bounded away from 0 and $\infty$ around $\left( \theta^*(\tau), \tau \right)$ for all $\tau \in \mathcal{H}$. \end{enumerate}

Under the normal likelihood and Assumption (ref), together with Assumption (ref), we can invoke the Laplace approximation results of hong_preston2012. Note that, under the normal likelihood and Assumption (ref), $\theta^*(\tau)$ exists and is a function of parameters in the DGP. In this section, we no longer assume the natural conjugate prior on $\theta$. For instance, the independent prior $\pi(\gamma,\sigma^2,\tau)=\pi(\gamma)\pi(\sigma^2) \pi(\tau)$ mentioned above satisfies the conditions in (ii) of Assumption (ref) as long as they are truncated on $\Theta$ and $\pi(\tau)$ is positive and finite at all $\tau$.

Theorem 3 below establishes the $T$-consistency of the marginal posterior of $\tau$ under this prior and the additional assumption.

theorem[Marginal posterior consistency of $\tau$ at rate $T^{-1}$, non-conjugate priors] Suppose Assumptions (ref) and (ref) hold. Then, under the normal likelihood, $\forall \eta>0, \epsilon>0$, $\exists M>0$ and $k>0$ such that $T \geq k \implies$ \[ P_{\theta_0, \tau_0} \left( \int_{B^c_{M/T} (\tau_0)} \pi_T(\tau) d\tau < \eta \right) > 1-\epsilon, \] where for any constant $d>0$, $B^c_{d}(\tau_0)$ denotes the set difference $\mathcal{H} \setminus (\tau_0-d, \tau_0+d)$.

Recall that while proving the $T$-consistency under the conjugate prior (i.e.,\ Theorem 1), we utilize the closed-form expression of the marginal posterior of $\tau$ up to a multiplicative constant (ref) in order to study the behavior of the marginal likelihood ratio conditional on $\tau$. Under non-conjugate priors, such expression is not available in general. For this reason, we invoke a Laplace approximation to investigate the quantity $\int p(\bm{Y}_T | \bm{X}_T, \theta, \tau) \pi(\theta,\tau) d\theta$ to prove Theorem 3. See Appendix (ref) for the detail.

As in the previous section, an implication of the $T$-consistency of the marginal posterior of $\tau$ is that the posterior mode converges to the limiting distribution of $\hat{\tau}_{LS}$. Proof is in Appendix (ref).

cor[Limiting distribution of the posterior mode of $\tau$, non-conjugate priors] Suppose Assumptions (ref) and (ref) hold. Then, under the normal likelihood, \[ \floor*{T\left( \hat{\tau}_{Bayes} - \tau_0 \right)} \overset{d}{\to} \operatorname*{arg\,max}_{m} W^*(m), \] where the stochastic process $W^*(m)$ is defined in Section (ref).

Theorem 4 establishes our main theoretical result, the Bernstein-von Mises theorem for $\gamma$, under the prior defined in Assumption (ref) (ii).

theorem[Bernstein-von Mises theorem for the slope coefficients, non-conjugate priors] Suppose Assumptions (ref) and (ref) hold. Then, under the normal likelihood, \[ d_{TV} \left( \pi \left[ \begin{matrix} \sqrt{T}\left( \gamma- \hat{\gamma}_{LS} \right) \end{matrix} \bigg| \bm{D}_T\right], N_{(d_x+d_z)}\left( 0, \sigma^2_0V^{-1} \right) \right) \to 0 , \] in $P_{\theta_0, \tau_0}-probability$ where $d_{TV}$ is the total variation distance.

When proving the corresponding result under the conjugate prior (i.e.,\ Theorem (ref)), we utilize the closed-form expression of the marginal posterior of $\gamma$ given $\tau$ (ref). As this is not available under the prior in this section, we again use a Laplace approximation to study the asymptotic behavior of the marginal posterior. See Appendix (ref) for a proof.

Simulation

The main purpose of the simulation studies below is to compare inference on the slope parameters $\gamma$ between the two methods: the conventional least-squares method in bai1997 and the Bayesian approach described in our paper. For the Bayesian approach, we use the uniform prior for $\tau$ and the conjugate prior for the regression parameters with $\underline{H}=0.1 I_{(d_x+d_z)}$, $\underline{\mu}=0_{(d_x+d_z)}$, and $\underline{a}=\underline{b}=1$. The findings are similar even when we use the uninformative improper prior. Following the literature (e.g.,\ casini_perron2021continuous), we set the range of the candidate values of $\tau$ to be $(\epsilon,1-\epsilon)$ with $\epsilon=0.05$ for all methods\footnote{It prevents the break location estimator from being in the first and last $100\epsilon$% of the sample. The trimming parameter $\epsilon$ should not be chosen too high otherwise it might introduce bias in the break location estimate. casini_perron2021continuous find the choice $\epsilon=0.05$ performs well in general, which we also confirm in our simulation exercises.}.

We consider the following model: $y_t=\delta_01(t>\floor*{\tau_0T})+\epsilon_t$. In order to compare the methods in repeated experiments, for each combination of $\tau_0$, $\delta_0$, and $T$, we generate 1,000 data sets. We consider different values of the break location $\tau_0 \in \{0.3,0.5\}$, the jump size $\delta_0 \in \{0.25,0.5,1.0,2.0\}$, and the sample size $T \in \{20, 50,100,250, 500, 1000\}$. The error $\epsilon_t$ is independently and identically generated from $N(0,1)$. In the online appendix, we present a robustness check with the errors generated from a mixture of two normals $0.5N\left(-1/\sqrt{2},1/2 \right) + 0.5N\left( 1/\sqrt{2},1/2 \right) $ and illustrate that the overall findings are similar to these under the normal DGP.

Table (ref) shows the simulation results concerning $\delta$. The top panel “Coverage" shows empirical coverages of the true jump size $\delta_0$ by the 95% confidence and credible intervals. The frequentist confidence intervals are computed based on the conventional asymptotic theory (ref). For the Bayesian approach, we report the equal-tailed credible intervals. The middle panel “Length" presents the average lengths of the aforementioned intervals. The bottom panel “MSE for $\delta$” shows the mean-squared-errors for the point estimator, which is the least-squares estimator $\hat{\delta}_{LS}$ defined in (ref) for the conventional method and the posterior mean for the Bayesian approach.

There are several significant findings. First, for small $T$ and/or small $\delta_0$, the conventional confidence intervals significantly undercover. Meanwhile, the Bayesian credible intervals have relatively reasonable coverages. Second, the Bayesian intervals tend to be longer than the conventional confidence intervals for small $T$ and/or $\delta_0$. Third, as $T$ increases, the discrepancy between the two methods decreases, as expected from the Bernstein-von Mises theorem that we establish.

\FloatBarrier

table[table omitted — 3,573 chars of source]

\FloatBarrier

Table (ref) shows the results of estimation and inference of the break location $\tau$. Although the main focus of the current paper is on inference about the slope parameters $\gamma$ and not on inference about $\tau$, we report the empirical coverage and the length of the 95% confidence interval of bai1997 and the highest posterior density (HPD) set\footnote{ Note that although we prove that the posterior mode of the break location $\tau$ converges to the limiting distribution of the least-squares estimator, whether the posterior distribution of $\tau$ converges to the same limit or not is still an open question. The Bayesian literature on the Bernstein-von Mises-like result for non-regular parameters is very scarce. To our best knowledge, the only available work is that by kleijn_knapik2012 whose results do not seem to be applicable to the model in consideration in this paper. Hence, it is not guaranteed that a credible set of $\tau$ has frequentist coverage even asymptotically. However, we emphasize that credible sets on $\tau$ still have a statistically valid interpretation even in finite samples.}. We also report the inverted likelihood ratio (ILR) confidence set suggested by eo_morley2015.

Overall, the HPD set and the ILR confidence set of the break location $\tau$ behave similarly although the HPD set slightly undercovers relative to the ILR confidence set for small $T$ and/or $\delta_0$. We confirm several findings of Eo and Morley (2015). First, when $T$ is large, the confidence interval of Bai has longer lengths than the ILR confidence set and the HPD set\footnote{ eo_morley2015 explain that the likelihood ratio test is more powerful than the Wald-type test used to construct the confidence interval of Bai, which results in a shorter length of the ILR confidence set. }. Second, when $T$ and $\delta_0$ are small, the confidence interval of Bai tends to severely undercover compared to the ILR confidence set and the HPD set. The interval of Bai indeed has a shorter length than the other two sets for small $T$, but its undercoverage raises concerns for small samples in practice\footnote{ Eo and Morley (2015) also find that the confidence interval of qu_perron2007 for the break location, which is also based on the Wald-type test as the confidence interval of Bai, tends to undercover in small sample despite having a slightly shorter length than the ILR confidence set. } \footnote{ In addition, as also reported by Eo and Morley (2015), the ILR confidence set tends to slightly overcover even in large sample. }.

The bottom panels of Table (ref) shows the mean-absolute-error (MAE) of the point estimator of $\tau$ which is $\hat{\tau}_{LS}$ defined in (ref) for the conventional method and the posterior mode $\hat{\tau}_{Bayes}$ for the Bayesian approach. It is known that the finite-sample distribution of the least-squares estimator $\hat{\tau}_{LS}$ tends to be trimodal (see baek2021) when the jump size is relatively small. The same seems to be true for the Bayesian point estimator (see Figure (ref)).

\FloatBarrier

table[table omitted — 4,542 chars of source]

\FloatBarrier

To better understand the importance of the uncertainty of the break location $\tau$ for inference on the slope parameters, we conduct a hypothetical experiment. We repeat the simulation exercise but now fixing $\tau$ at the least-squares estimate $\hat{\tau}_{LS}$. Table (ref) displays the results. Note that the results for the least-squares estimator are of course the same as in Table (ref). We however now see that, not only the conventional confidence intervals of $\delta$ but also the credible intervals undercover for small $T$ and/or small $\delta_0$. They also have similar lengths in general. Importantly, the credible intervals when $\tau$ is fixed at $\hat{\tau}_{LS}$ (Table (ref)) have shorter lengths compared to the full Bayesian intervals (Table (ref)). On average, the full Bayesian credible intervals are 17.1% longer\footnote{The difference is larger when $T$ and/or $\delta_0$ are/is smaller.} than the credible intervals produced by fixing the value of $\tau$ at $\hat{\tau}_{LS}$. Note that a Bayesian equivalent of the conventional approach to inference on the slope parameters would be to fix the value of $\tau$ at the posterior mode (whose value is very similar to $\hat{\tau}_{LS}$ as we can see from Figure (ref) and deduce from Corollary (ref)). We can see in Figure (ref) that both $\hat{\tau}_{LS}$ and the posterior mode of $\tau$ display significant amount of variations. Fixing $\tau$ at a point estimate forces the Bayesian approach to ignore this uncertainty of $\tau$; as a result, the credible interval on $\delta$ becomes shorter and hence undercovers. The full Bayesian approach takes into account such uncertainty via marginal posterior of $\tau$ (see examples of the density in Figure (ref)). This results in longer lengths of the full Bayesian intervals on the slope parameters and helps them avoid undercoverage. In contrast, by construction (i.e.,\ Equation (ref)), the conventional confidence intervals do not have this feature.

\FloatBarrier

table[table omitted — 3,594 chars of source]

\FloatBarrier

In summary, the simulation exercises demonstrate that (1) the credible intervals on the slope coefficient tend to have more reasonable coverages than the conventional confidence intervals because of longer lengths, (2) the longer length of the credible intervals is a reflection of the uncertainty of the unknown\footnote{When $\tau_0$ is known, the two intervals behave very similarly. To illustrate this point, we conduct another hypothetical experiment by repeating the simulation exercise as before but now fixing the value of $\tau$ at the true value $\tau_0$ in both conventional and Bayesian approaches. Table (ref) summarizes the results. In this case, we see that both confidence and credible intervals have coverages quite close to 95% in all cases. They also have similar lengths. Note that when the true value $\tau_0$ is given, the usual asymptotic normality and the regular Bernstein-von Mises theorem apply. As a consequence, both frequentist and Bayesian intervals seem to converge faster to the limit compared to the case with unknown $\tau$.} break location $\tau$, and (3) the two intervals converge to each other asymptotically as expected from our Bernstein-von Mises theorem.

\FloatBarrier

table[table omitted — 3,596 chars of source]

\FloatBarrier

Application

In this section, we illustrate difference in estimation and inference of the regression parameters in linear regression models with a structural break between the conventional approach and the Bayesian approach that we consider in this paper. paye_timmermann2006 consider the problem of ex-post prediction in stock returns under a structural break in the coefficients of state variables. Their multivariate model with a structural break is \[ Ret_{t} =

cases\delta^{(1)}_1 + \delta^{(2)}_1 Div_{t-1} + \delta^{(3)}_1 Tbill_{t-1} + \delta^{(4)}_1 Spread_{t-1} + \delta^{(5)}_1 Def_{t-1} + \epsilon_t , & if t \leq \floor{\tau T} \\ \delta^{(1)}_2 + \delta^{(2)}_2 Div_{t-1} + \delta^{(3)}_2 Tbill_{t-1} + \delta^{(4)}_2 Spread_{t-1} + \delta^{(5)}_2 Def_{t-1} + \epsilon_t , & if t > \floor{\tau T} ,

\] where $Ret_t$ is the excess return for the international index in question during month $t$, $Div_{t-1}$ is the lagged dividend yield, $Tbill_{t-1}$ is the lagged local country short interest rate, $Spread_{t-1}$ is the lagged local country term spread, and $Def_{t-1}$ is the lagged U.S. default premium. The authors estimate the model using the conventional frequentist approach: they first compute $\hat{\tau}_{LS}$ and then obtain point estimates as well as confidence intervals for the slope coefficients by fixing $\tau$ at $\hat{\tau}_{LS}$. We examine whether the Bayesian method performs differently from the conventional approach.

Monthly series are collected from Global Financial Data and Federal Reserve Economic Data (FRED). In this paper, we consider estimating the model for the United Kingdom and Japan\footnote{ paye_timmermann2006 conduct the sequential method suggested by bai_perron1998, bai_perron2003, perron2006 for determining the number of breaks and find multiple breaks for some countries. They find single breaks for the U.K. and Japan, but, for example, two breaks for the U.S. A fully Bayesian approach would be to place a prior on the number of breaks and use a trans-dimensional estimation method such as a reversible jump MCMC, which is beyond the scope of this paper. }. The indices to which the total return and the dividend yield correspond are the FTSE All-share for the U.K. and Nikko Securities Composite for Japan. For each country, a 3-month Treasury bill rate is used as a measure of the short interest rate while the yield on a long-term government bond is used as a measure of the long interest rate. Excess returns are computed as the total return on stocks in the local currency minus the local short rate. The dividend yield is expressed as an annual rate and is constructed as the sum of dividends over the preceding 12 months, divided by the current price. A term spread is the difference between the long and short local country interest rates. The U.S. default premium is defined as the difference in yields between Moody's Baa and Aaa rated bonds. For each country, the sample spans between January 1970 and December 2003.

For both approaches, we set the range of the candidate values of $\tau$ to be $(\epsilon,1-\epsilon)$ with $\epsilon=0.05$ as we do in the simulation studies in the previous section. For the Bayesian approach, we use the uniform prior on $(\epsilon,1-\epsilon)$ for $\tau$ and the conjugate prior for the regression parameters with $H=0.1 I_{(d_x+d_z)}$, $\underline{\mu}=0_{(d_x+d_z)}$, and $\underline{a}=\underline{b}=1$. The findings are similar even when we use the uninformative improper prior. For the break date, we compute the least-squares estimator $\hat{\tau}_{LS}$ and the posterior mode $\hat{\tau}_{Bayes}$ of $\tau$ as well as the 95% confidence interval of bai1997, the highest posterior density (HPD) set, and the inverted likelihood ratio (ILR) confidence set of eo_morley2015. For the slope parameters, we compute $\hat{\gamma}_{LS}$ and the posterior mean of $\gamma$ as well as the 90% confidence intervals of bai1997 based on the asymptotic result (ref) and the equal-tailed credible intervals.

table[table omitted — 2,374 chars of source]

When the uncertainty about $\tau$ is small, estimation and inference of the slope parameters roughly match between the conventional least-squares approach and the Bayesian approach, as illustrated by our simulation studies and indicated by our proven Bernstein-von mises theorem. See Table (ref) for the results for the U.K. Both methods estimate a break at 1975:01. The confidence interval of bai1997, the Bayesian highest posterior density (HPD) set, and the inverted likelihood ratio (ILR) confidence set by eo_morley2015 are all similar and narrow, indicating that the uncertainty about $\tau$ is small. This can be seen also from the posterior density on the break date in Panel (a) of Figure (ref), which has a sharp peak around 1975:01\footnote{ The mean and the standard deviation of the excess return of the FTSE All-share index during the sample period are -1.53 and 6.94 respectively. At $t=$1974:12, we have $Ret_t=-9.9$ while at $t=$1975:01, $Ret_t=43.75$, where the change is approximately 7.7 standard deviations. Therefore, the change in the dependent variable is large enough for the break point to be detected with small uncertainty.}. paye_timmermann2006 explain that the break in the mid-1970's might be related to the large macroeconomic shocks reflecting oil price increases. As a result of the small uncertainty about $\tau$, the point estimates of the slope parameters as well as the corresponding confidence/credible intervals are similar between the conventional and the Bayesian approach. Importantly, when the confidence interval of a given slope parameter includes (or does not include) zero, the corresponding credible interval also includes (or does not include) zero. Hence, the conventional approach to inference about the slope parameters for the U.K. sample seems to be robust with respect to the uncertainty on the break date.

\FloatBarrier \graphicspath{{Figures_final/application_oct2021/}}

figure[figure omitted — 403 chars of source]

\FloatBarrier

table[table omitted — 2,382 chars of source]

In contrast, when the uncertainty on $\tau$ is large, the conventional and the Bayesian results on inference about the slope parameters might disagree. Table (ref) shows the results for Japan. Although both $\hat{\tau}_{LS}$ and the posterior mode of $\tau$ are at 1996:05, the HPD set and the ILR confidence set are much wider than the confidence interval of bai1997, indicating a large uncertainty of the break date. The posterior density on $\tau$ in Figure (ref) also illustrates that the uncertainty of the break date is much larger for Japan than for the U.K. during the sample period\footnote{ In addition, the posterior on $\tau$ for Japan exhibits tri-modality, which would be similar to the tendency of a finite-sample distribution of $\hat{\tau}_{LS}$ to have three modes as reported in the literature (e.g.,\ baek2021; casini_perron2021continuous). }. The large uncertainty of $\tau$ is reflected on Bayesian inference on the slope parameters. In the upper panel of Table (ref), we see that in general the Bayesian credible intervals are wider than the confidence intervals. Importantly, this can have a qualitative consequence on statistical importance of some parameters. For seven of the ten slope coefficients, the confidence intervals do not include zero while the the Bayesian credible intervals do. Hence, the conventional approach to inference on the slope parameters might not be robust with respect to the uncertainty of the break date, for the Japanese sample.

Conclusion and future direction

In this paper, we establish a Bernstein-von Mises type theorem for the slope coefficients in linear regression with a structural break. By doing so, we bridge the gap between the frequentist and the Bayesian approaches for inference on this model. On the one hand, a frequentist researcher can look at Bayesian credible intervals for the slope coefficients as a robustness check to see whether the uncertainty of the break location affects inference on the slope parameters. Such sensitivity analysis is natural as our theoretical result guarantees the credible interval to converge to the conventional confidence interval that the frequentist researcher would use otherwise. On the other hand, Bayesian inference can be conveyed to frequentists via our proven result.

Potential extensions include several directions. First, the homoscedasticity assumption could be too strong in some applications, and hence extending the results to the case of heteroscedasticity and autocorrelation would be of interest. Second, a popular Bayesian method of chib1998 is different from the approach we took in this paper in that we place an explicit prior on $\tau$ and that Chib's framework can be naturally extended to the case of multiple breaks. It would be interesting to study frequentist properties of Chib's approach.