EconBase
← Back to paper

Subgeometric ergodicity and $β$-mixing

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

35,874 characters · 6 sections · 74 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Subgeometric ergodicity and $$-mixing

abstractIt is well known that stationary geometrically ergodic Markov chains are $\beta$-mixing (absolutely regular) with geometrically decaying mixing coefficients. Furthermore, for initial distributions other than the stationary one, geometric ergodicity implies $\beta$-mixing under suitable moment assumptions. In this note we show that similar results hold also for subgeometrically ergodic Markov chains. In particular, for both stationary and other initial distributions, subgeometric ergodicity implies $\beta$-mixing with subgeometrically decaying mixing coefficients. Although this result is simple it should prove very useful in obtaining rates of mixing in situations where geometric ergodicity can not be established. To illustrate our results we derive new subgeometric ergodicity and $\beta$-mixing results for the self-exciting threshold autoregressive model. Classifications (MSC2010): 60J05, 37A25. Keywords: Markov chains; rates of convergence, mixing coefficients, subgeometric rate, subexponential rate, polynomial rate, SETAR model.

Introduction

Let $X_{t}$ ($t=0,1,2,\ldots$) be a Markov chain on the state space $\mathsf{X}$ with $n$-step transition probability measure $P^{n}$ and stationary distribution $\pi$. If the $n$-step probability measures $P^{n}$ converge in total variation norm to the stationary probability measure $\pi$ at rate $r^{n}$ (for some $r>1$), that is,

equation[equation omitted — 140 chars of source]

the Markov chain is said to be geometrically ergodic. It is well known that for stationary Markov chains, geometric ergodicity implies that so-called $\beta$-mixing coefficients (or coefficients of absolute regularity) $\beta(n)$, to be defined formally in Section 2, converge to zero at the same rate, $\lim_{n\to\infty}r^{n}\beta(n)=0$ (see, e.g., doukhan1994mixing, bradley2005basic, or bradley2007introduction). For initial distributions other than the stationary one, a similar mixing result has been obtained by liebscher2005towards.

We are interested in counterparts of these mixing results when the convergence in ((ref)) takes place at a rate $r(n)$ slower than geometric, that is,

equation[equation omitted — 142 chars of source]

When ((ref)) holds with suitably defined rates $r(n)$ slower than geometric, the Markov chain is called subgeometrically ergodic. The main result of this note establishes that for both stationary and other initial distributions, subgeometric ergodicity implies $\beta$-mixing with subgeometrically decaying mixing coefficients, that is, $\lim_{n\to\infty}\tilde{r}(n)\beta(n)=0$ for some rate function $\tilde{r}(n)$.

To illustrate some common rate functions, consider the expression \[ r(n)=(1+\ln(n))^{\alpha}\,\cdot\,(1+n)^{\beta}\,\cdot\,e^{cn^{\gamma}}\,\cdot\,e^{dn},\qquad\alpha,\beta,c,d\geq0,\,\,\gamma\in(0,1),\,\,n\geq1. \] In the case $\alpha,\beta,c,d>0$ the four terms above satisfy $e^{dn}/e^{cn^{\gamma}}\to\infty$, $e^{cn^{\gamma}}/(1+n)^{\beta}\to\infty$, and $(1+n)^{\beta}/(1+\ln(n))^{\alpha}\to\infty$ as $n\to\infty$, and this hierarchy can be used to define different growth rates. Ordered from the fastest to the slowest growth rate, a growth rate is called geometric (sometimes also exponential) if the dominant term is $e^{dn}$ (with $d>0$; note that $e^{dn}=r^{n}$ with $r>1$ iff $d>0$), subexponential if the dominant term is $e^{cn^{\gamma}}$ ($c>0$ and above $d=0$), polynomial if the dominant term is $(1+n)^{\beta}$ ($\beta>0$, $c=d=0$), and logarithmic if the dominant term is $(1+\ln(n))^{\alpha}$ ($\alpha>0$, $\beta=c=d=0$).

To provide some brief background on subgeometric ergodicity, we note that the first subgeometric ergodicity results for general state space Markov chains were obtained by nummelin1983rate and tweedie1983criteria; the subgeometric rate functions $r(n)$ considered were introduced by stone1967one. tuominen1994subgeometric gave a set of conditions that imply the convergence in ((ref)) and, in particular, formulated a sequence of so-called drift conditions to establish subgeometric ergodicity. Subsequent work by fort2000vsubgeometric, jarner2002polynomial, fort2003polynomial, and douc2004practical lead to a formulation of a single drift condition to ensure subgeometric ergodicity, paralleling the use of a Foster-Lyapunov drift condition to establish geometric ergodicity (see, e.g., meyn2009markov).

The rest of the paper proceeds as follows. Section 2 contains necessary mathematical preliminaries. Section 3 reviews the relation of geometric ergodicity and $\beta$-mixing, while the corresponding results in the subgeometric case are given in Section 4. The general results obtained are exemplified in Section 5 where subgeometric ergodicity and $\beta$-mixing results for the self-exciting threshold autoregressive model are presented. Section 6 concludes, and all proofs are given in an Appendix.

Preliminaries

To formalize the discussion in the Introduction, consider $X_{t}$ ($t=0,1,2,\ldots$), a time-homogeneous discrete-time Markov chain on a general measurable state space $(\mathsf{X},\mathcal{B}(\mathsf{X}))$. Comprehensive treatments of the relevant Markov chain theory can be found in meyn2009markov or douc2018markov. Let $\mu$ be any initial measure on $\mathcal{B}(\mathsf{X})$, and suppose that $X_{0}$ has distribution $\mu$. Denote the transition probabilities with $P(x\,;\,A)$ ($x\in\mathsf{X}$, $A\in\mathcal{B}(\mathsf{X})$) and let $(\Omega,\mathcal{F},\mathsf{P}_{\mu})$ denote the probability space of the Markov process $\{X_{0},X_{1},\ldots\}$. As usual, $\mathsf{P}_{x}$ denotes the probability measure corresponding to a fixed initial value $X_{0}=x$ and $P^{n}(x\,;\,A)=\mathsf{P}_{x}(X_{n}\in A)$ ($x\in\mathsf{X}$, $A\in\mathcal{B}(\mathsf{X})$) signifies the $n$-step transition probability measure.

Next consider the rate of convergence of the $n$-step probability measures $P^{n}$ to the stationary probability measure $\pi$. To this end, for any two probability measures $\lambda_{1}$ and $\lambda_{2}$ on $(\mathsf{X},\mathcal{B}(\mathsf{X}))$, the total variation distance is defined as $\lVert\lambda_{1}-\lambda_{2}\rVert=2\sup_{B\in\mathcal{B}(\mathsf{X})}\lvert\lambda_{1}(B)-\lambda_{2}(B)\rvert=\sup_{\lvert h\rvert\leq1}\lvert\lambda_{1}(h)-\lambda_{2}(h)\rvert$, where the last supremum runs over all $\mathcal{B}(\mathsf{X})$-measurable functions $h:\mathsf{X}\to\mathbb{R}$ bounded in absolute value by 1 and $\lambda_{i}(h)=\int_{\mathsf{X}}\lambda_{i}(dx)h(x)<\infty$. The $n$-step probability measures $P^{n}$ converge in total variation norm to the stationary probability measure $\pi$ at rate $r(n)$, $n\geq0$, if

equation[equation omitted — 142 chars of source]

If ((ref)) holds we say that the Markov chain $X_{t}$ is ergodic with rate $r(n)$; geometric ergodicity obtains when $r(n)=r^{n}$ for some $r>1$.

To define the $\beta$-mixing coefficients, let $\mathcal{F}_{k}^{l}$, $0\leq k\leq l\leq\infty$, signify the $\sigma$-algebra generated by $\{X_{k},\ldots,X_{l}\}$. For the stochastic process $\{X_{0},X_{1},\ldots\}$ the $\beta$-mixing coefficients $\beta(n)$, $n=1,2,\ldots$, are defined as (doukhan1994mixing; bradley2007introduction)

align*[align* omitted — 350 chars of source]

where $\mathbb{N}=\{0,1,2,\ldots\}$ and in the first expression for $\beta(n)$ the second supremum is taken over all pairs of (finite) partitions $\{A_{1},A_{2},\ldots,A_{I}\}$ and $\{B_{1},B_{2},\ldots,B_{J}\}$ of $\Omega$ such that $A_{i}\in\mathcal{F}_{0}^{m}$ for each $i$ and $B_{j}\in\mathcal{F}_{n+m}^{\infty}$ for each $j$. For our purposes it is convenient to use the following alternative expression obtained by davydov1973mixing:

equation[equation omitted — 192 chars of source]

where $\mu P^{m}(\cdot)=\int_{\mathsf{X}}\mu(dx)P^{m}(x\,;\,\cdot)$ denotes the distribution of $X_{m}$ ($m=1,2,\ldots$; $\mu P^{0}=\mu$). In case of a stationary Markov chain (i.e., one with initial distribution $\pi$), the $\beta$-mixing coefficients can be expressed simply as

equation[equation omitted — 162 chars of source]

Process $X_{t}$ is said to be $\beta$-mixing (or sometimes absolutely regular) if $\lim_{n\rightarrow\infty}\beta(n)=0$. As with the convergence in ((ref)), the rate of this convergence is of interest, and in what follows we seek for results of the form $\lim_{n\rightarrow\infty}r(n)\beta(n)=0$ with some rate function $r(n)$.

The geometric case

We start by briefly discussing the relation of geometric ergodicity and $\beta$-mixing; although these results are well known, comparing them with the subgeometric case will be illuminating. In case of a stationary Markov chain (i.e., one with initial distribution $\pi$), this relation is particularly simple. As was first shown by nummelin1982geometric, a geometrically ergodic Markov chain satisfies, for some $r>1$, $\lim_{n\rightarrow\infty}r^{n}\int\pi(dx)\left\Vert P^{n}(x\,;\,\cdot)-\pi(\cdot)\right\Vert =0$; given expression ((ref)), the $\beta$-mixing property immediately follows and the mixing coefficients satisfy $\lim_{n\rightarrow\infty}r^{n}\beta(n)=0$. Statements of this result can be found for instance in doukhan1994mixing, bradley2005basic, and bradley2007introduction. For initial distributions other than the stationary one, a corresponding result seems to have first appeared in liebscher2005towards.

To facilitate comparison with the subgeometric case, we present the ergodicity and mixing results as consequences of a particular drift criterion; as is discussed in meyn2009markov, this is how geometric ergodicity is often established. We use the following traditional Foster-Lyapunov type geometric drift condition (cf. meyn2009markov).\footnote{As a technical remark, note that in Condition Drift\textendash G we assume the function $V$ to be everywhere finite (i.e., $V\,:\,\mathsf{X}\rightarrow[1,\infty)$) and such that $\sup_{x\in C}V(x)<\infty$. In contrast, in meyn2009markov it is only assumed that $V$ is extended-real-valued (i.e., $V\,:\,\mathsf{X}\rightarrow[1,\infty]$) and finite at some one $x_{0}\in\mathsf{X}$. Our stronger requirements hold in most practical applications and lead to more transparent exposition and proofs.} Here $\boldsymbol{1}_{C}(\cdot)$ signifies the indicator function of a set $C$.

condition*[Drift\textendash G] Suppose there exist a petite set $C$, constants $b<\infty$, $\beta>0$, and a measurable function $V\,:\,\mathsf{X}\rightarrow[1,\infty)$ such that $\sup_{x\in C}V(x)<\infty$, satisfying \[ E\left[V(X_{1})\,\left|\,X{}_{0}=x\right.\right]\leq V(x)-\beta V(x)+b\boldsymbol{1}_{C}(x),\qquad x\in\mathsf{X}. \]

For the definition of a `petite set' appearing in this condition, and for the concepts of irreducibility and aperiodicity in the theorem below, we refer the reader to meyn2009markov. Theorem 1 summarizes the relation between geometric ergodicity and $\beta$-mixing.

thmSuppose $X_{t}$ is a $\psi$-irreducible and aperiodic Markov chain and that Condition Drift\textendash G holds. Then \begin{lyxlist}{(x)} • $X_{t}$ is geometrically ergodic, i.e, for some $r_{1}>1$, $\lim_{n\rightarrow\infty}r_{1}^{n}\left\Vert P^{n}(x\,;\,\cdot)-\pi(\cdot)\right\Vert =0$ for all $x\in\mathsf{X}$. \end{lyxlist} Suppose further that the initial state $X_{0}$ has distribution $\mu$ such that $\int_{\mathsf{X}}\mu(dx)V(x)<\infty$. Then \begin{lyxlist}{(x)} • for some $r_{2}>1$, $\lim_{n\rightarrow\infty}r_{2}^{n}\int_{\mathsf{X}}\mu(dx)\left\Vert P^{n}(x\,;\,\cdot)-\pi(\cdot)\right\Vert =0$, \end{lyxlist} and \begin{lyxlist}{(x)} • $X_{t}$ is $\beta$-mixing and the mixing coefficients satisfy, for some $r_{3}>1$, $\lim_{n\rightarrow\infty}r_{3}^{n}\beta(n)=0$. \end{lyxlist} Moreover: \begin{lyxlist}{(x)} • In the stationary case ($\mu=\pi$) condition $\int_{\mathsf{X}}\pi(dx)V(x)<\infty$ is not needed, (b) and (c) hold with $r_{2}=r_{3}$, and (b) and (c) are equivalent. \end{lyxlist}

Parts (a) and (b) are very well known (see for instance meyn2009markov for part (a) and nummelin1982geometric for part (b)) and so is also the mixing result in the stationary case (see the references given earlier). Part (c) for general initial distributions was obtained by liebscher2005towards, although our formulation is somewhat different from his (our formulation and proof avoid the use of so-called `$Q$-geometric ergodicity' employed by Liebscher; for completeness, our proof of Theorem 1, which may be of independent interest, is provided in a Supplementary Appendix). Part (d) elaborates parts (b) and (c) as well as their relation in the stationary case.

The subgeometric case

We seek a counterpart of Theorem 1 in which the geometric rate $r^{n}$ is replaced by some slower rate function; such rate functions were already exemplified in the Introduction. More formally, the subgeometric rate functions we consider are defined as follows (cf., e.g., nummelin1983rate and douc2004practical). Let $\Lambda_{0}$ be the set of positive nondecreasing functions $r_{0}\,:\,\mathbb{N}\rightarrow[1,\infty)$ such that $\ln[r_{0}(n)]/n$ decreases to zero as $n\rightarrow\infty$. The class of subgeometric rate functions, denoted by $\Lambda$, consists of positive functions $r\,:\,\mathbb{N}\rightarrow(0,\infty)$ for which there exists some $r_{0}\in\Lambda_{0}$ such that

equation[equation omitted — 153 chars of source]

Typical examples are obtained of rate functions $r$ for which these inequalities hold with (for notational convenience, we set $\ln(0)=0$) \[ r_{0}(n)=(1+\ln(n))^{\alpha}\,\cdot\,(1+n)^{\beta}\,\cdot\,e^{cn^{\gamma}},\qquad\alpha,\beta,c\geq0,\,\gamma\in(0,1). \] The rate function $r_{0}(n)$ is called subexponential when $c>0$, polynomial when $c=0$ and $\beta>0$, and logarithmic when $\beta=c=0$ and $\alpha>0$.

In analogy with the geometric case, subgeometric ergodicity and mixing results are most conveniently obtained by verifying an appropriate drift condition. The following drift condition for subgeometric ergodicity is adapted from douc2018markov.\footnote{A somewhat more general drift condition, for instance allowing for $V$ to be extended-real-valued, is given in douc2004practical.}

condition*[Drift\textendash SubG] Suppose there exist a petite set $C$, a constant $b<\infty$, a concave increasing continuously differentiable function $\phi\,:\,[1,\infty)\rightarrow(0,\infty)$ satisfying $\lim_{v\rightarrow\infty}\phi'(v)=0$, and a measurable function $V\,:\,\mathsf{X}\rightarrow[1,\infty)$ such that $\sup_{x\in C}V(x)<\infty$ and \[ E\left[V(X_{1})\,\left|\,X{}_{0}=x\right.\right]\leq V(x)-\phi(V(x))+b\boldsymbol{1}_{C}(x),\qquad x\in\mathsf{X}. \]

Note that if $\phi(v)=\eta v$ ($\eta>0$), one obtains Condition Drift\textendash G (but assumption $\lim_{v\rightarrow\infty}\phi'(v)=0$ rules this out; as we are interested in subgeometric rates of ergodicity, assuming this means no loss of generality, see douc2018markov).

Following douc2004practical we next introduce a rate function, denoted by $r_{\phi}$. First define the function $H_{\phi}(v)=\int_{1}^{v}\frac{dx}{\phi(x)}$, where $\phi$ is as in Condition Drift\textendash SubG. The definition implies that $H_{\phi}$ is a nondecreasing, concave, and differentiable function on $[1,\infty)$, and it has an inverse $H_{\phi}^{-1}\,:\,[0,\infty)\rightarrow[1,\infty)$ which is increasing and differentiable (see douc2004practical). Thus, we can define the rate function \[ r_{\phi}(z)=(H_{\phi}^{-1})'(z)=\phi\circ H_{\phi}^{-1}(z). \] douc2004practical show that this rate function is subgeometric and that Condition Drift\textendash SubG implies the convergence ((ref)) at rate $r_{\phi}(n)$.

Theorem 2 summarizes the relation between subgeometric ergodicity and $\beta$-mixing. Here $\lfloor k\rfloor$ denotes the integer part of the real number $k$.

thmSuppose $X_{t}$ is a $\psi$-irreducible and aperiodic Markov chain and that Condition Drift\textendash SubG holds. Then \begin{lyxlist}{(x)} • $X_{t}$ is subgeometrically ergodic with rate $r_{\phi}(n)$, i.e, $\lim_{n\rightarrow\infty}r_{\phi}(n)\left\Vert P^{n}(x\,;\,\cdot)-\pi(\cdot)\right\Vert =0$ for all $x\in\mathsf{X}$. \end{lyxlist} Suppose further that the initial state $X_{0}$ has distribution $\mu$ such that $\int_{\mathsf{X}}\mu(dx)V(x)<\infty$. Then \begin{lyxlist}{(x)} • $\lim_{n\rightarrow\infty}r_{\phi}(n)\int\mu(dx)\left\Vert P^{n}(x\,;\,\cdot)-\pi(\cdot)\right\Vert =0$, \end{lyxlist} and \begin{lyxlist}{(x)} • $X_{t}$ is $\beta$-mixing and the mixing coefficients satisfy $\lim_{n\rightarrow\infty}\tilde{r}_{\phi}(n)\beta(n)=0$ for any rate function $\tilde{r}_{\phi}(n)$ such that $\limsup_{n\rightarrow\infty}\tilde{r}_{\phi}(n)/r_{\phi}(n_{1})<\infty$ where $n_{1}=\lfloor n/2\rfloor$. \end{lyxlist} Moreover: \begin{lyxlist}{(x)} • In the stationary case ($\mu=\pi$) condition $\int_{\mathsf{X}}\pi(dx)V(x)<\infty$ is not needed, (b) and (c) hold with $r_{\phi}(n)=\tilde{r}_{\phi}(n)$, and (b) and (c) (with $r_{\phi}(n)=\tilde{r}_{\phi}(n)$) are equivalent. • If $r_{\phi}(n)$ satisfies ((ref)) with $r_{\phi,0}(n)=(1+\ln(n))^{\alpha}\cdot(1+n)^{\beta}\cdot e^{cn^{\gamma}}$ and $\tilde{r}_{\phi}(n)$ satisfies ((ref)) with $\tilde{r}_{\phi,0}(n)=(1+\ln(n))^{\alpha}\cdot(1+n)^{\beta}\cdot e^{\tilde{c}n^{\gamma}}$ for some $0<\tilde{c}<c2^{-\gamma}$, then $\limsup_{n\rightarrow\infty}\tilde{r}_{\phi}(n)/r_{\phi}(n_{1})<\infty$. \end{lyxlist}

Of the results in Theorem 2, part (a) is given in Proposition 2.5 of douc2004practical. Part (b) can be obtained by combining Theorem 4.1 of tuominen1994subgeometric and Proposition 2.5 of douc2004practical, but in the proof we make use of the work of nummelin1983rate. Part (c) is new and illuminates the relation between subgeometrically ergodic Markov chains and their $\beta$-mixing properties, thereby providing a counterpart of a result obtained by liebscher2005towards in the case of geometric ergodicity. Part (d) is analogous to its counterpart in Theorem 1 and provides further insight to parts (b) and (c) whereas part (e) makes part (c) more concrete in the case of the most common rate functions. For completeness, we give a detailed proof in the Appendix.

As discussed in douc2004practical and meitz2019subgear, there is a connection between the function $\phi$ and the rate function $r_{\phi}$, which can be used to find out the latter in particular cases. For instance, polynomial rate functions are associated with cases where the function $\phi$ is of the form $\phi(v)=cv^{\alpha}$ with $\alpha\in(0,1)$ and $c\in(0,1]$, and then the rate obtained is $r_{\phi}(n)=n^{\alpha/(1-\alpha)}$ (an alternative form is $r_{\phi}(n)=n^{\kappa-1}$ with $\kappa=1+\alpha/(1-\alpha)$ already given by jarner2002polynomial). In the subexponential case the function $\phi$ is such that $v/\phi(v)$ goes to infinity slower than polynomially so that a possibility, given in meitz2019subgear, is $\phi(v)=c(v+v_{0})/[\ln(v+v_{0})]^{\alpha}$ for some $\alpha,c,v_{0}>0$. This results in the rate $r_{\phi}(n)=(e^{d})^{n^{1/(1+\alpha)}}$ for some $d>0$ which is faster than polynomial. A logarithmic rate is an example of a rate slower than polynomial. Then the function $\phi$ is of the form $\phi(v)=c[1+\ln(v)]^{\alpha}$ for some $\alpha>0$ and $c\in(0,1]$, and the resulting rate is $r_{\phi}(n)=[\ln(v)]^{\alpha}$ (see douc2004practical).

Theorem 2 (or 1) also provides information about the moments of the stationary distribution of $X_{t}$. Specifically, once part (a) of Theorem 2 (or 1) has been established, one can deduce from Condition Drift\textendash SubG (or Drift\textendash G) and Theorem 14.3.7 of meyn2009markov that $\int_{\mathsf{X}}\pi(dx)\phi(V(x))<\infty$ (or $\int_{\mathsf{X}}\pi(dx)V(x)<\infty$). This can be very useful when one aims to apply limit theorems developed for $\beta$-mixing processes where moment conditions are typically assumed.

We close this section by noting that Condition Drift\textendash SubG can also be used to obtain more general ergodicity results than provided in Theorem 2. Without going into details we only mention that Theorem 2.8 of douc2004practical and Theorem 1 of meitz2019subgear show how a stronger form of ergodicity, called ($f,r$)-ergodicity, can be established.

Example

To illustrate our results we consider the self-exciting threshold autoregressive (SETAR) model studied by chan1985multiple. These authors analyzed the model

equation[equation omitted — 105 chars of source]

where $-\infty=r_{0}<\cdots<r_{M}=\infty$ and for each $j=1,\ldots,M$, $\{W_{t}(j)\}$ is an independent and identically distributed mean zero sequence independent of $\{W_{t}(i)\}$, $i\neq j$, and with $W_{t}(j)$ having a density that is positive on the whole real line. They considered the following conditions

subequations\begin{align} & \theta(1)<1,\quad\theta(M)<1,\quad\theta(1)\theta(M)<1,\\ & \theta(1)=1,\quad\theta(M)<1,\quad0<\varphi(1),\\ & \theta(1)<1,\quad\theta(M)=1,\quad\varphi(M)<0,\\ & \theta(1)=1,\quad\theta(M)=1,\quad\varphi(M)<0<\varphi(1),\\ & \theta(1)<0,\quad\theta(1)\theta(M)=1,\quad\varphi(M)+\varphi(1)\theta(M)>0, \end{align}

and showed that the SETAR model is ergodic if and only if one of the conditions ((ref))\textendash ((ref)) holds (chan1985multiple). Moreover, if $E[\vert W_{t}(j)\vert]<\infty$ for each $j$, they showed that condition ((ref)) ensures geometric ergodicity (chan1985multiple). To our knowledge, in the cases ((ref))\textendash ((ref)) no results regarding the rate of ergodicity have as yet appeared in the literature and our Theorem 4(b) below indicates that geometric ergodicity may not always hold without stronger assumptions.\footnote{meyn2009markov also discuss the (geometric) ergodicity of the SETAR model ((ref)), reproducing the ergodicity result of chan1985multiple as their Proposition 11.4.5. On their p. 541, meyn2009markov also state that (our additions in brackets) “in the interior of the parameter space {[}the union of ((ref))\textendash ((ref)){]} we are able to identify geometric ergodicity in Proposition 11.4.5\ \ldots\ the stronger form {[}geometric ergodicity{]} is actually proved in that result” but no formal proof is given for this statement. }

We consider rates of ergodicity and $\beta$-mixing in case ((ref)) when the autoregressive coefficients $\theta(1)$ and $\theta(M)$ equal unity. For intuition, note that due to nonzero intercept terms $\varphi(1)$ and $\varphi(M)$, both the first and the last regimes exhibit nonstationary random walk type behavior with a drift. As the intercept terms satisfy $\varphi(M)<0<\varphi(1)$, the drift is increasing in the first regime and decreasing in the last regime. This feature prevents the process $y_{t}$ from exploding to (plus or minus) infinity, thereby providing intuition why ergodicity can hold true. It is noteworthy that ergodicity is in no way dependent of the behavior of the process in the middle regimes ($2,\ldots,M-1$) which can exhibit stationary, random walk type (with or without drift), or even explosive behavior.

In their results, chan1985multiple allow for regime dependent distributions for the error term $W_{t}(j)$. To obtain our results for the case ((ref)), we strengthen the assumptions on the error term and, in particular, assume that the error distribution is the same in each regime (this stronger assumption is needed to apply the results mentioned in the proof of Theorem 3 below, and relaxing it appears less than straightforward). To compensate, we obtain results for a model more general than the SETAR model ((ref)) with ((ref)). Specifically, we formulate our results in terms of the general nonlinear autoregressive model

equation[equation omitted — 81 chars of source]

where the function $g\,:\,\mathbb{R}\rightarrow\mathbb{R}$ and the error term $\varepsilon_{t}$ satisfy the following conditions:

lyxlist{0000} • $g$ is a measurable function with the property $\left|g(x)\right|\rightarrow\infty$ as $\left|x\right|\rightarrow\infty$ and such that there exist positive constants $r$ and $M_{0}$ such that \[ \left|g(x)\right|\leq\left(1-r/\left|x\right|\right)\left|x\right|\quad\textrm{for }\left|x\right|\geq M_{0}\quad\textrm{and}\quad{\textstyle \sup_{\left|x\right|\leq M_{0}}}\left|g(x)\right|<\infty; \]$\{\varepsilon_{t},\,t=1,2,\ldots\}$ is a sequence of independent and identically distributed mean zero random variables that is independent of $X_{0}$ and the distribution of $\varepsilon_{1}$ has a (Lebesgue) density that is bounded away from zero on compact subsets of $\mathbb{R}$.

Model ((ref)) with conditions A1 and A2 is a special case of models considered by fort2003polynomial, douc2004practical, and meitz2019subgear. These authors consider much more general models but for clarity of presentation we have simplified the model as much as possible while still being able to obtain results for the SETAR model ((ref)) with ((ref)) (the first two of the abovementioned papers consider a multivariate version of ((ref)), whereas the third one considers a higher-order generalization of ((ref)); the inequality constraint for the function $g$ in condition A1 is also more general in these papers where it is only required that $\left|g(x)\right|\leq\left(1-r\left|x\right|^{-\rho}\right)\left|x\right|$ for some $0<\rho\leq2$).

The following Theorem establishes ergodicity and $\beta$-mixing results for model ((ref)) with varying rates of convergence. The proof (in the Appendix) makes use of results in fort2003polynomial, douc2004practical, and meitz2019subgear to obtain rates of ergodicity, as well as Theorems 1 and 2 above to obtain rates of $\beta$-mixing (only the subgeometric mixing results in parts (b) and (c) are new).

thmConsider model ((ref)) with conditions (A1) and (A2). \begin{lyxlist}{(x)} • If $E\bigl[e^{z_{0}\left|\varepsilon_{1}\right|}\bigr]<\infty$ for some $\ensuremath{z_{0}>0}$, then $X_{t}$ is geometrically ergodic with convergence rate $r(n)=r_{1}^{n}$ for some $r_{1}>1$. Moreover, if the initial state $X_{0}$ has a distribution such that $E[e^{z\left|X_{0}\right|}]<\infty$ for some $z>0$, then $X_{t}$ is also $\beta$-mixing and the mixing coefficients satisfy, for some $r_{3}>1$, $\lim_{n\rightarrow\infty}r_{3}^{n}\beta(n)=0$. • If $E\bigl[e^{z_{0}\left|\varepsilon_{1}\right|^{\kappa_{0}}}\bigr]<\infty$ for some $\ensuremath{z_{0}>0}$ and $\ensuremath{\kappa_{0}\in(0,1)}$, then $X_{t}$ is subexponentially ergodic with convergence rate $r(n)=(e^{c})^{n^{\kappa_{0}}}$ (for some $c>0$). Moreover, if the initial state $X_{0}$ has a distribution such that $E[e^{z\left|X_{0}\right|^{\kappa_{0}}}]<\infty$ for some $z>0$, then $X_{t}$ is also $\beta$-mixing and the mixing coefficients satisfy, for some $\tilde{c}>0$, $\lim_{n\rightarrow\infty}(e^{\tilde{c}})^{n^{\kappa_{0}}}\beta(n)=0$. • If $E\left[\left|\varepsilon_{1}\right|^{s_{0}}\right]<\infty$ for either $s_{0}=2$ or $s_{0}\geq4$, then $X_{t}$ polynomially ergodic with convergence rate $r(n)=n^{s_{0}-1}$. Moreover, if the initial state $X_{0}$ has distribution such that $E\left[\left|X_{0}\right|^{s_{0}}\right]<\infty$, then $X_{t}$ is also $\beta$-mixing and the mixing coefficients satisfy $\lim_{n\rightarrow\infty}n^{s_{0}-1}\beta(n)=0$. \end{lyxlist}

Theorem 3 shows that there is a trade-off between rates of ergodicity and $\beta$-mixing and finiteness of moments of the error term. The fastest geometric rate is obtained when $E\bigl[e^{z_{0}\left|\varepsilon_{1}\right|}\bigr]<\infty$ ($z_{0}>0$) so that $\varepsilon_{1}$ has finite moments of all orders and the slowest polynomial rate is obtained when only $E\left[\varepsilon_{1}^{2}\right]<\infty$. As discussed after Theorem 2, we also have $\int_{\mathsf{X}}\pi(dx)\phi(V(x))<\infty$ so that there is a similar trade-off between these convergence rates and finiteness of moments of the stationary distribution (expressions of $V$ and $\phi$ are available in the proof of Theorem 3).

Above it was mentioned that fort2003polynomial, douc2004practical, and meitz2019subgear consider (subgeometric) ergodicity of models more general than ((ref)) with conditions (A1) and (A2). Making use of our Theorems 1 and 2, subgeometric rates of $\beta$-mixing can straightforwardly be obtained also for these more general models. We omit the details for brevity.

In a series of papers, Veretennikov and co-authors also considered the model ((ref)) with function $g$ satisfying $\left|g(x)\right|\leq\left(1-r\left|x\right|^{-\rho}\right)\left|x\right|$ for some $1\leq\rho\leq2$. Using methods very different from ours, they obtained results on subgeometric ergodicity and subgeometric rates for $\beta$-mixing coefficients. The cases $1<\rho<2$ and $\rho=2$ are considered in veretennikov2000polynomial, klokov2004sub,klokov2005subexponential, and klokov2007lower and are shown to lead to subgeometric rates. For the case $\rho=1$ relevant for the SETAR example, these papers refer to veretennikov1988bounds,veretennikov1991estimating and veretennikov1990rate. A result corresponding to our Theorem 3(a) can be found in veretennikov1990rate but subgeometric rates, such as those in our Theorem 3(b) and (c), do not seem to be established in the case $\rho=1$.

We now specialize the results above to the SETAR model ((ref)) with ((ref)). It is easy to see that this model, with the function $g$ in ((ref)) defined as $g(x)=\sum_{j=1}^{M}[\varphi(j)+\theta(j)x]\bm{1}\{x\in(r_{j-1},r_{j}]\}$ (with $\bm{1}\{\cdot\}$ denoting the indicator function), satisfies the condition in A1. Namely, for $x$ large enough and positive we have $\lvert g(x)\rvert=g(x)=x+\varphi(M)=\lvert x\rvert-(-\varphi(M))$ whereas for $x$ small enough and negative we have $\lvert g(x)\rvert=-g(x)=-x-\varphi(1)=\lvert x\rvert-\varphi(1)$, so that the inequality in A1 holds for $M_{0}>\max\{\left|r_{1}\right|,\left|r_{M-1}\right|\}$ and $r=\min\{\varphi(1),-\varphi(M)\}$ (and the supremum condition is obviously satisfied).

Part (a) of the next theorem simply restates the result of Theorem 3 for the SETAR model ((ref)) with ((ref)), whereas part (b) establishes that geometric ergodicity cannot hold under the weaker moment assumptions of Theorem 3(b) and (c).

thmConsider the SETAR model ((ref)) with the parameters satisfying ((ref)) and the error terms satisfying $W_{t}(j)=\varepsilon_{t}$ ($j=1,\ldots,M$) with $\varepsilon_{t}$ as in (A2). \begin{lyxlist}{(x)} • Sufficient conditions for geometric, subexponential, and polynomial ergodicity and $\beta$-mixing of $X_{t}$ are as in parts (a), (b), and (c) of Theorem 3, respectively. • If $E\bigl[e^{z_{0}\left|\varepsilon_{1}\right|}\bigr]=\infty$ for all $\ensuremath{z_{0}>0}$, then $X_{t}$ is not geometrically ergodic. \end{lyxlist}

Theorem 4(b) shows that for the SETAR model ((ref)) with ((ref)), the subgeometric rates of Theorem 3(b) and (c) cannot be improved to a geometric rate unless stronger moment assumptions are made regarding the error term. This result is obtained by making use of a necessary condition for geometric ergodicity of certain specific type of Markov chains in jarner2003necessary (using their necessary condition to obtain this result appears possible only in case ((ref)) out of ((ref))\textendash ((ref))).

Conclusion

In this note we have shown that subgeometrically ergodic Markov chains are $\beta$-mixing with subgeometrically decaying mixing coefficients. Although this result is simple it should prove very useful in obtaining rates of mixing in situations where geometric ergodicity can not be established. An illustration using the popular self-exciting threshold autoregressive model showed how our results can yield new subgeometric rates of mixing.