Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
81,690 characters · 11 sections · 88 citation commands
The Local to Unity Dynamic Tobit Model
\global\long\def\uwrite#1#2{\underset{#2}{\underbrace{#1}} }
\global\long\def\blw#1{\ensuremath{#1}}
\global\long\def\abv#1{\ensuremath{\overline{#1}}}
\global\long\def\vect#1{\mathbf{#1}}
\global\long\def\smlseq#1{\{#1\} }
\global\long\def\seq#1{\left\{ #1\right\} }
\global\long\def\smlsetof#1#2{\{#1\mid#2\} }
\global\long\def\setof#1#2{\left\{ #1\mid#2\right\} }
\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long \global\long \global\long \global\long\def\Ellp#1{\ensuremath{\mathcal{L}^{#1}}}
\global\long \global\long \global\long \global\long \global\long
\global\long\def\abs#1{\ensuremath{\left|#1\right|}}
\global\long\def\smlabs#1{\ensuremath{\lvert#1\rvert}}
\global\long\def\bigabs#1{\ensuremath{\bigl|#1\bigr|}}
\global\long\def\Bigabs#1{\ensuremath{\Bigl|#1\Bigr|}}
\global\long\def\biggabs#1{\ensuremath{\biggl|#1\biggr|}}
\global\long\def\norm#1{\ensuremath{\left\Vert #1\right\Vert }}
\global\long\def\smlnorm#1{\ensuremath{\lVert#1\rVert}}
\global\long\def\bignorm#1{\ensuremath{\bigl\|#1\bigr\|}}
\global\long\def\Bignorm#1{\ensuremath{\Bigl\|#1\Bigr\|}}
\global\long\def\biggnorm#1{\ensuremath{\biggl\|#1\biggr\|}}
\global\long\def\floor#1{\left\lfloor #1\right\rfloor } \global\long\def\smlfloor#1{\lfloor#1\rfloor}
\global\long\def\ceil#1{\left\lceil #1\right\rceil } \global\long\def\smlceil#1{\lceil#1\rceil}
\global\long \global\long \global\long \global\long \global\long \global\long\def\clsr#1{\ensuremath{\overline{#1}}}
\global\long \global\long \global\long \global\long
\global\long\def\smlinprd#1#2{\ensuremath{\langle#1,#2\rangle}}
\global\long\def\inprd#1#2{\ensuremath{\left\langle #1,#2\right\rangle }}
\global\long \global\long
\global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long\def\sigf#1{\mathcal{#1}}
\global\long\global\long \global\long\def\flt#1{\mathcal{#1}}
\global\long\global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long \global\long \def\independenT#1#2{\mathrel{\rlap{$#1#2$}\mkern2mu{#1#2}}}
\global\long \global\long \global\long \global\long \global\long \global\long\def\inprobu#1{\ensuremath{\overset{#1}{\ensuremath{\rightarrow}}}}
\global\long \global\long \global\long\def\inLp#1{\ensuremath{\overset{\Ellp{#1}}{\ensuremath{\rightarrow}}}}
\global\long \global\long \global\long \global\long\def\wkcu#1{\overset{#1}{\ensuremath{\rightsquigarrow}}}
\global\long
\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long
\global\long \global\long \global\long \global\long \global\long \global\long\def\cv#1{\left\langle #1\right\rangle }
\global\long\def\smlcv#1{\langle#1\rangle}
\global\long\def\qv#1{\left[#1\right]}
\global\long\def\smlqv#1{[#1]}
\global\long \global\long \global\long \global\long \global\long\global\long \global\long\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long\def\vec#1{\boldsymbol{#1}}
\global\long \address[Anna Bykhovskaya]{Duke University} \email{[email removed]}
\address[James A.\ Duffy]{University of Oxford} \email{[email removed]}
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{0.5\linespacing} {\normalfont}{Introduction}
Since the 1950s nonlinear models have played an increasingly prominent role in the analysis and prediction of time series data. In many cases, as was noted in early work by Moran_lynx, linear models are unable to adequately match the features of observed time series. The efforts to develop models that enjoy the flexibility afforded by nonlinearities, while retaining the tractability of linear models, have subsequently engendered an enormous literature (see e.g.\ Fan_Yao_nonlinear_book; Gao07; Chan09; and TTG10).
An important instance of nonlinearity arises when data is bounded by, truncated at, or censored below some threshold, since such phenomena cannot be adequately captured -- even approximately -- by a linear model. Many observed series are bounded below by construction, and may spend lengthy periods at or near their lower boundary, such as unemployment rates, prices, gross sectoral trade flows, and nominal interest rates. The non-negativity of interest rates, and the resulting constraints that this may impose on the efficacy of monetary policy, has received particular attention in recent years, as central bank policy rates have remained at or near the zero lower bound for a significant portion of the past two decades, across many economies (see e.g.\ Mav21Ecta, and the works cited therein).
A tractable model for such series, which generates both their characteristic serial dependence and censoring, is the dynamic Tobit model. In its static formulation, the model originates with Tobin_tobit. In its dynamic formulation, the model typically comes in one of two varieties, which we refer to as the latent and censored models. In the latent dynamic Tobit, an unobserved process $\{y_{t}^{\ast}\}$ follows a linear autoregression, with $y_{t}=\max\{y_{t}^{\ast},0\}$ being observed; whereas in the censored dynamic Tobit, $\{y_{t}\}$ is modelled as the positive part of a linear function of its own lags, and an additive error (see e.g.\ Maddala83, or wei1999). In both models the right hand side may be augmented with other explanatory variables. Relative to the latent model, the censored model has the advantage of being Markovian, which greatly facilitates its use in forecasting. It has also been successfully applied to a range of censored series, in both purely time series and panel data settings, including: the open market operations of the Federal Reserve (Demiralp_Jorda2002; stat_paper); household commodity purchases (DSK12AE); loan charge-off rates (LMS19); credit default and overdue loan repayments (BMMV21JBF); and sectoral bilateral trade flows bykh_JBES. Recently, Mav21Ecta proposed the censored and kinked structural VAR model to describe the operation of monetary policy during periods when the zero lower bound may occasionally bind on the policy rate. If only the actual interest rate (rather than some `shadow rate') affects agents' decision making, as assumed in closely related work by AMSV21JoE, then the univariate counterpart of this model is exactly the censored dynamic Tobit.
The present work is concerned with the censored, rather than the latent, dynamic Tobit model. As discussed by stat_paper, the censored model is arguably more appropriate in settings where $y_t = 0$ results not from limits on the observability of some underlying $y_t^{\ast}$, but from an economic constraint on the values taken by $y_t$. For example, bykh_JBES justifies the application of the dynamic Tobit with reference to a game-theoretic model of network formation, in which the zeros correspond to corner solutions of a constrained optimisation problem, i.e.\ where zeros are systematically observed because of a non-negativity constraint on agents' choices. In our illustrative empirical application, to an exchange rate that is subject to a floor engineered by a central bank (Section (ref)), the most recent values of the exchange rate are taken as sufficient to describe the market equilibrium, and thus the conditional distribution of future rates. Our focus on the censored model is also motivated by relatively greater need for the development of relevant econometric theory in this area. In the latent model, the dynamics are simply those of the latent autoregression, and so are readily understood by standard methods; whereas in the censored model, the censoring affects the dynamics of $\{y_{t}\}$ in a non-trivial manner, making the analysis rather more challenging. Indeed, establishing the stationarity or weak dependence of the censored dynamic Tobit is far from trivial, as can be seen from HK10EcLett, stat_paper, MdJ18EcLett, and bykh_JBES. Henceforth, all references to the `dynamic Tobit' are to the censored version of the model.\footnote{stat_paper alternatively refer to this model as a `dynamic censored regression', but the term `dynamic Tobit' appears more commonly in the literature (see e.g.\ HK10EcLett,MdJ18EcLett,bykh_JBES).}
Motivated in part by recent work on modelling nominal interest rates near the zero lower bound, our concern is with the application of this model to series that are highly persistent, so that above the censoring point they exhibit the random wandering that is characteristic of integrated processes. The appropriate configuration of the dynamic Tobit model for such series, in which the autoregressive polynomial has a root local to unity, has not been considered in the literature to date -- apart from the special case of a first-order model with an exact unit root, as in cav_bounded and bykh_JBES. Our results are thus entirely new to the literature.
Our principal technical contribution, within this setting, is to derive the limiting distributions of both the standardised regressor process, and the ordinary least squares (OLS) estimates of the parameters of the dynamic Tobit, when that model has an autoregressive root local to unity. The reader may find our focus on OLS surprising, as this method would usually provide inconsistent estimates in the presence of censoring. However, it turns out that in our setting consistency (for all model parameters) is restored, and we obtain a usable limit theory for the estimated sum of the autoregressive coefficients, which conventionally provides a measure of the overall persistence of a process (cf.\ AC94JBES; Mik07Ecta). While one may contemplate alternatively using maximum likelihood (ML) or least absolute deviations (LAD) to estimate the model, OLS enjoys the advantages of maintaining only weak distributional assumptions on the innovations (unlike ML), and avoiding the numerical minimisation of a non-convex criterion function (unlike LAD).
Our asymptotics provide the basis for practical unit root tests for highly persistent, censored time series. Motivated by our finding that OLS is consistent, we consider a test based on the (constant only) augmented Dickey--Fuller (ADF) $t$ statistic, but which employs critical values modified to reflect the censoring present in the data generating process. We show, via Monte Carlo simulations, that as our critical values are larger than the conventional ADF critical values, their use eliminates the significant over-rejection that may result from the naive application of the ADF test to censored data. (This tendency to over-reject the null of a unit root appears typical of models that incorporate unit roots and nonlinearities, having been also found by e.g.\ HT97EcLett; KLN02JoE; and WDJ13SASA.) Strikingly, the distribution of $t$ statistic, under censoring, is stochastically dominated by that obtained from the linear autoregressive model.\footnote{This holds only with respect to the distributions: it is not true that if one simulates a linear and a Tobit model with the same underlying innovations, then the former $t$ statistic will always be larger than the latter.}
Our work may be construed, more broadly, as extending the analysis of highly persistent time series, and the associated machinery of unit root testing, from a linear setting to a nonlinear setting appropriate to time series that are subject to a lower bound. In doing so, we complement the seminal work of cavaliere2005, which similarly sought to extend this machinery to the setting of bounded time series. Our contribution is to effect this extension within a class of nonlinear autoregressive models that have been widely applied to censored time series (as evinced by the works cited above), and which fall outside his framework.
On a technical level, the most closely related works to our own are those of cav_bounded,cavaliere2005 and CX14JoE, who develop the asymptotics of what they term `limited autoregressive processes' with a near-unit root, which are (one- or two-sided) non-Markovian censored processes constructed by the addition of regulators to a latent linear autoregression. While their (one-sided) model has a superficial resemblance to the dynamic Tobit, there are important, but subtle differences between the two (see Section (ref) for a discussion). Perhaps the most striking similarity is that both models, in the case of an exact unit root, give rise to processes that converge weakly to regulated Brownian motion; but when roots are merely local to unity, the limiting processes associated with these two models are distinct (see the discussion following Theorem (ref)).
Convergence to regulated Brownian motions has also been obtained previously in the setting of first-order threshold autoregressive models with an (exact) unit root regime and a stationary regime, as considered by LLS11Bern and GLY13JoE. However, allowing for both near-unit roots and higher-order autoregressive terms introduces technical challenges that require us to take a markedly different approach from those employed in these earlier works. With respect to higher-order models, a major difficulty relates to the treatment of the differences $\{\Delta y_t\}$. In a linear autoregressive model with a single unit root, and all other roots outside the unit circle, these would follow a stationary autoregression; but in our setting they instead follow a regime-switching autoregression, where the regime depends on the (lagged) level of $y_t$, and so are inherently non-stationary. Accoridingly, standard arguments for controlling the magnitude of $\Delta y_t$, and deriving the limits of functionals thereof, are unavailing. We provide a striking example in which $\Delta y_t$ is explosive even though all but one of the autoregressive roots (the root at unity) lie outside the unit circle, due to the interactions between the two autoregressive regimes (see Appendix (ref)). To preclude such cases, we develop a condition relating to the joint spectral radius of the autoregressive representation for $\Delta y_t$, which is sufficient to control the magnitude of $\Delta y_t$ and plays an essential role in our arguments. (This concept has been previously employed in the context of stationary autoregressive models, see e.g.\ Lieb05JTSA; Saik08ET.)
The remainder of this paper is organised as follows. Section (ref) discusses the model and our assumptions. Asymptotic results and corresponding tests are derived in Section (ref), while supporting Monte Carlo simulations are shown in Section (ref). Section (ref) applies our framework to the exchange rate between the Swiss franc and the euro during a period when this rate was subject to a lower bound. Finally, Section (ref) concludes. All proofs appear in the appendices.
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{0.5\linespacing} {\normalfont}{The dynamic Tobit model with a near-unit root}
Consider a time series $\{y_{t}\}$ generated by the dynamic Tobit model of order $k\geq1$, written in augmented Dickey--Fuller (ADF) form,\footnote{The autoregressive form is $y_{t}=[\alpha+\sum_{i=1}^{k}\beta_{i}y_{t-i}+u_{t}]_{+}$, where $\beta_1 = \beta + \phi_1$, $\beta_k = -\phi_{k-1}$, and $\beta_i = \phi_i - \phi_{i-1}$ for $i \in \{2,\ldots,k-1\}$. In particular $\beta=\sum_{i=1}^{k}\beta_{i}$ corresponds to the sum of the autoregressive coefficients.}
where $\Delta y_{t}:=y_{t}-y_{t-1}$, and $[x]_{+}\coloneqq\max\{x,0\}$ denotes the positive part of $x\in\mathbb{R}$. Let
where $\phi(z):=1-\sum_{i=1}^{k-1}\phi_{i}z^{i}$. We impose the following on the data generating process (ref).
Figure (ref) displays a typical sample path for the dynamic Tobit (ref), under the preceding assumptions.
[4]
The main consequences of our assumptions may be summarised as follows.
Our machinery extends straightforwardly to the case where $y_{t}$ is censored at some $\mathbf{L}\neq0$. Suppose (ref) is modified to
and take $\mathbf{L}=T^{1/2}\ell$ for some $\ell\in\mathbb{R}$, to allow the censoring point to be of the same order of magnitude as $\{y_{t}\}$.\footnote{ We emphasise that this dependence of $\mathbf{L}$ on $T$ should not be interpreted literally as specifying a model in which the censor point is a function of the sample size. Rather, the scaling is a mathematical device that allows us to obtain an improved asymptotic approximation to the finite sample distribution of $T^{-1/2}y_{\smlfloor{rT}}$, and hence of the test statistics that depend upon it. (See e.g.\ Assumption (A4) in cavaliere2005, where a similar device is used for this purpose.) If $\mathbf{L}$ is in fact `small' relative to a given sample size $T$, then this will be accommodated within our framework by $\ell$ being close to zero.} Defining $\tilde{y}_{t}\coloneqq y_{t}-\mathbf{L}$ and subtracting $\mathbf{L}$ from both sides of (ref), it may be verified that
where $\tilde{\alpha}\coloneqq\alpha+(\beta-1)\mathbf{L}$. Thus, $\{\tilde{y}_{t}\}$ follows a dynamic Tobit with censoring at zero, with drift \[ T^{1/2}\tilde{\alpha}=T^{1/2}[\alpha+(\beta-1)T^{1/2}\ell]\ensuremath{\rightarrow} a+c\ell\eqqcolon\tilde{a} \] and initialisation \[ \tilde{b}_{0}=\operatorname*{plim}_{T\ensuremath{\rightarrow}\infty}T^{-1/2}\tilde{y}_{0}=\operatorname*{plim}_{T\ensuremath{\rightarrow}\infty}T^{-1/2}(y_{0}-\mathbf{L})=b_{0}-\ell. \] All our results hold in the setting of (ref), with appropriate modifications. For simplicity, we work with $\mathbf{L}=0$ throughout the rest of the paper, except where otherwise indicated.
It will be occasionally useful to rewrite (ref) in a form that helps to clarify the connections between the dynamic Tobit and the linear autoregressive model. We can do this by defining
where $[x]_{-}\coloneqq\min\{x,0\}$. That is, when $y_{t}=0$, $y_{t}^{-}$ records the value that $y_{t}$ would have taken had it not been censored at zero.
Since $[x]_{+}=x-[x]_{-}$, we may then rewrite (ref) as
or equivalently, letting $L$ denote the lag operator, as
Thus, if we view $y_{t}^{-}$ as an additional noise term, (ref) takes the form of a linear autoregression. The main challenge is that $y_{t}^{-}$ is itself a complicated nonlinear object, whose presence fundamentally alters the dynamics of $\{y_t\}$, even in the long run.
The representation (ref) allows us to draw out the connections between our model and the limited autoregressive processes developed by cav_bounded,cavaliere2005 and CX14JoE. To put their model -- for the special case of a process constrained to lie in $[0,\infty)$ -- in a form comparable to ours, consider a latent process $\{x_{t}^{\ast}\}$,
where $\{\varepsilon_{t}\}$ is stationary. Define an observed process $\{x_{t}\}$, whose increments are related to those of $\{x_{t}^{\ast}\}$ via
where $\blw{\xi}_{t}>0$ if and only if $x_{t-1}+\Delta x_{t}^{\ast}<0$, so as to ensure that $x_{t}\geq0$ for all $t$. In particular, if we set
then $\{x_{t}\}$ will be censored at zero. When $c=0$, by combining (ref)--(ref) we obtain
as a valid representation of a limited autoregressive process censored at zero.
Both (ref) and (ref) describe censored processes, but have some subtle, and yet important differences. Firstly, while $\{y_t\}$ in (ref) is a Markov process (with state vector $(y_{t},\ldots,y_{t-k+1})$), $\{x_t\}$ in (ref) will be Markov only if $\{\varepsilon_t\}$ is i.i.d. The former is thus more suited to forecasting in the presence of higher-order dynamics.
Secondly, at a technical level, the differences between the models can be clearly illustrated by supposing that $\alpha=0$ and $\beta=1$. Then $B(L)=(1-L)\phi(L)$ in ((ref)), which simplifies to
In the dynamic Tobit, higher-order dynamics are captured by the (stationary) autoregressive polynomial $\phi(L)$, whereas in the limited autoregressive model these enter via the weak dependence in $\{\varepsilon_{t}\}$. To facilitate a comparison of the two models, suppose that $\{\varepsilon_{t}\}$ follows the autoregression $\phi(L)\varepsilon_{t}=u_{t}$. Then ((ref)) becomes
We see immediately that if both models have only first-order dynamics ($k=1$), so that $\phi(L)=1$, then they exactly coincide (cf.\ Cavaliere, 2005, Remark 2.3). However, this no longer holds in the presence of higher-order dynamics ($k\geq2$). Suppose e.g.\ that $\phi(L)=1-\phi_1 L$ for some $\smlabs{\phi_1}<1$. Then ((ref)) becomes
whereas ((ref)) yields
Comparing ((ref)) with ((ref)), we see that the censoring affects the dynamics of $\{y_{t}\}$ and $\{x_{t}\}$ in different ways. Lagged values of $y_{t}^{-}$ have a direct effect on future $y_{t}$ (via $\sum_{s=0}^{\infty}\phi^{s}_1 y_{t-s}^{-}$), whereas lagged $x_{t}^{-}$ have no such effect on $x_{t}$.
Thirdly, the differences between the two models also manifests itself through each giving rise to distinct classes of limiting processes. Even when $k=1$, these coincide only in the special case of an exact unit root ($c=0$): see the discussion following Theorem (ref) below.
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{0.5\linespacing} {\normalfont}{Asymptotic results}
In this section we derive the weak limits of the standardised process $T^{-1/2}y_{\smlfloor{rT}}$ and the ordinary least squares (OLS) estimators of the parameters of the dynamic Tobit. The latter provides the basis for a unit root test for non-negative time series, a test that is both is straightforward to compute, and which does not require any assumption to be made on the distribution of the innovations, beyond the existence of sufficient moments. We first provide a separate treatment of the first-order model ($k=1$), before progressing to the model with higher-order dynamics ($k\geq 1$). This facilitates a simplified exposition of the former case, which avoids the additional assumptions ((ref) and (ref)) and technical concepts -- notably the joint spectral radius -- that are required when $k\geq 2$.
Let $\theta\coloneqq(a,b_{0},c)$, define the process
and denote its `regulated' counterpart by
Here the supremum on the r.h.s.\ regulates $J_{\theta}(r)$ to ensure that it is always non-negative: if $K_{\theta}(r)$ is negative, $[-K_{\theta}(r)]_{+}=-K_{\theta}(r)>0$, so that $K_{\theta}(r)+\sup_{r^{\prime}\leq r}[-K_{\theta}(r^{\prime})]_{+}\geq0$.
We first provide a result for the case $k=1$, under which the model (ref) reduces to
The preceding is a new result, which relates to some of the previous literature as follows.
When $k>1$, the other roots of the lag polynomial $B(L)$ affect the behaviour of $\{y_{t}\}$, and we need a further condition to ensure that the first differences $\{\Delta y_{t}\}$ are well behaved. Let
Under an appropriate condition on the matrices $\{F_{\delta}\mid\delta\in[0,1]\}$, we can ensure $\{\Delta y_{t}\}$ is stochastically bounded. To state that, we need the following (cf.\ Jungers09, Defn.\ 1.1):
Control over the JSR has been previously used to ensure the stationarity of regime-switching autoregressive models (e.g. Lieb05JTSA; Saik08ET), and we shall utilise it in a similar manner here.
Approximate upper bounds for the JSR can be computed numerically, to an arbitrarily high degree of accuracy, via semidefinite programming PJ08LAA, making it reasonably straightforward to verify whether this condition is satisfied by given parameter values. The following result, which is proved in Appendix (ref), provides a sufficient condition for (ref), which may be checked even more simply.
The principal difference between Theorems (ref) and (ref) is that when $k>1$, the stationary dynamics appear in the limit via the factor $\phi(1)$. Notably, the local autoregressive parameter $c$ is replaced by $\phi(1)^{-1}c$ -- exactly as it would be if $\{y_{t}\}$ were generated by a linear autoregression with a root local to unity (cf.\ Hansen99REStat, p.\ 599). Indeed, $\phi(1)=1$ when $k=1$, so in this case the two results coincide.
We first consider the case where $k=1$, as in the model (ref), to develop intuition for our results.
When estimating (ref) by OLS, we need to decide which deterministic terms should be included in the regression. In the absence of censoring, i.e. if the data generating process were simply a linear autoregression, the inclusion of a constant and a linear trend would render the distribution of the OLS estimator of $\beta$ free of any nuisance parameters related to the deterministic components.\footnote{Strictly speaking, this is true only if the autoregressive model is formulated in `unobserved components' form (see e.g.\ AC94JBES as
so that the presence (or absence) of a linear drift in $y_{t}$ is independent of the value of $\beta$, and so can always be removed by deterministic detrending. By contrast, if the model is formulated `directly' as \[ y_{t}=\alpha+\beta y_{t-1}+u_{t}, \] then the linear trend that is present when $\beta=1$ becomes an exponential trend when $\beta$ is local to unity. In the present (censored) setting, we may note that (ref) is not equivalent to
(This model is in fact the latent dynamic Tobit referred to in Section (ref).)} Unfortunately, the nonlinearity introduced by the censoring entails that $\alpha$ -- or rather, the local parameter $a$ -- will show up in the limiting distribution of $\hat{\beta}_{T}$, irrespective of which deterministics are included in the regression. To permit inferences to also be drawn on $a$, if required, we consider the OLS regression of $y_{t}$ on a constant and $y_{t-1}$, i.e.
In the stationary dynamic Tobit model, OLS is inconsistent (see e.g. bykh_JBES). However, as the following shows, when $\beta$ is local to unity, consistency is restored. The reason is that observations in the vicinity of zero accumulate only at rate $T^{1/2}$, so that a vanishingly small fraction of the sample is affected by the censoring.
For the case of general $k\geq1$, let $\vec{\phi}\coloneqq(\phi_{1},\ldots,\phi_{k-1})^{\mathsf{T}}$, and
denote the OLS estimators of the parameters of (ref). Since, as the next results shows, the limiting distributions of $(\hat{\alpha}_{T},\hat{\beta}_{T})$ depend on $\phi(1)$, a consistent estimate of that quantity is needed to compute valid critical values for test statistics based on these estimators. The following also guarantees the consistency of $\hat{\phi}(1)\coloneqq1-\sum_{i=1}^{k-1}\hat{\phi}_{i,T}$.
The theorem shows that, depite the nonlinearity of the Tobit model, OLS is consistent for all parameters in the presence of a near unit root. The associated limit distribution theory for $\hat{\beta}_T$ provides a basis for the unit root tests developed in the next section, yielding a test that is both easy to compute, and semiparametric with respect to the distribution of the innovations. (For examples in which the erroneous imposition of normality can have undesirable consequences, in the context of a dynamic Tobit model, see the empirical illustrations in stat_paper and bykh_JBES.)
The preceding results allow us to conduct asymptotically valid hypothesis tests on key parameters of the dynamic Tobit: in particular, to test the hypothesis of a unit root in this setting. This may be of interest for several reasons. For example, whether the variance of the errors made in forecasting $y_t$ remains bounded, or grows without bound at progressively longer forecast horizons, depends crucially on the presence of a unit root. In a setting with multiple series, one or more of which are non-negative, the presence of unit roots may also lead to spurious regressions or, more constructively, allow long-run equilibrium relationships to be identified from the (nonlinear) cointegrating relationships between the series (see DMW22).
In a linear autoregressive model, the presence of a unit root -- equivalently, the sum of the autoregressive coefficients being unity ($\beta=1$) -- necessarily imparts a stochastic trend to $\{y_t\}$. However, in the dynamic Tobit the value of the intercept also matters. In particular, a fixed negative intercept ($\alpha<0$ and not drifting toward zero) would continually push the process back towards the censoring point, thereby rendering it stationary (for $k=1$, see bykh_JBES). Thus to the extent that the purpose of a test for a unit root is to test for the presence of a stochastic trend in $\{y_t\}$, rather than to detect a unit root per se, it may be considered more appropriate to test the null that $\alpha=0$ and $\beta=1$, as opposed to merely the restriction that $\beta=1$, with it being desirable to reject this null in favour of a stationary alternative, when either $\beta<1$ (exactly as in a linear model), or when $\beta=1$ but $\alpha<0$.\footnote{When $k > 1$, the above must be qualified somewhat, because of the possibility that the higher-order nonlinear dynamics of the system may generate explosive trajectories even when $\beta < 1$. For stationary alternatives that are local to unity in the sense that $\beta = \exp(c/T)$ for some $c<0$, this is excluded by (ref). For non-local alternatives, some further condition (i.e.\ in addition to $\beta < 1$) on the autoregressive system is needed to ensure stationarity: see e.g.\ stat_paper or DMW22.}
To construct our test statistics, we need an estimate of the error variance $\sigma^{2}$. We use $\hat{\sigma}_{T}^{2}\coloneqq\frac{1}{T}\sum_{t=1}^{T}\hat{u}_{t}^{2}$, where
That is, $\{\hat{u}_{t}\}$ are the OLS residuals, computed as if $y_{t}$ were not subject to censoring. Let $\mathcal{M}_{T}\coloneqq\sum_{t=1}^{T}\vec x_{t}\vec x_{t}^{\mathsf{T}}$, where $\vec x_{t}\coloneqq(1,y_{t-1},\Delta y_{t-1},\ldots,\Delta y_{t-k+1})^{\mathsf{T}}$.
This result allows us to conduct a one-sided test of a unit root versus a stationary alternative, which rejects when $t_{\beta,T}\leq c$, where $c$ is drawn from an appropriate quantile of the asymptotic distribution of $t_{\beta,T}$. Under the null of a unit root $a=c=0$, the limiting distribution of $t_{\beta,T}$ in (ref) depends (continuously) on the model parameters only through $b_0 \phi(1)/\sigma$. This follows from the fact that for $\theta_{\phi}=(0,\phi(1)b_0,0)$,
and thus $\sigma\sqrt{\mathcal{J}_{\theta_{\phi}}^{-1}(2,2)}$ and $\mathfrak{b}_{\theta_{\phi}}$ (as defined in (ref) above) do not change so long as $b_0 \phi(1)/\sigma$ remains fixed. Table (ref) tabulates the critical values corresponding to the relevant quantiles of the asymptotic distribution of $t_{\beta,T}$, as a function of $b_0 \phi(1)/\sigma \in [0,2.5]$. One can see that the smaller the ratio of the parameters, the larger the corresponding critical values, for every significance level. Further, as the final two lines of the table and the further discussion in Section (ref) indicate, for values of $b_0 \phi(1)/\sigma$ in excess of $2.5$, the critical values coincide (to within two decimal places) with those of a conventional ADF $t$ test (for a linear autoregression).
With the aid of the tabulated critical values, a test of $H_0 : a=c=0$ may thus be carried out as follows. (For a general lower bound $\mathbf{L} \neq 0$, first subtract it from the data, as described in Section (ref).)
As the simulations in the following section illustrate, our test indeed has the desirable properties outlined above, in the sense of tending to reject both when either $\beta<1$, or when $(\beta=1 ,\ \alpha<0)$, i.e.\ it has power to reject the null whenever $\{y_t\}$ is stationary. In the event of a rejection, the reason for that rejection -- i.e.\ whether this is due to $\alpha<0$ or $\beta<1$ -- may be further investigated with the aid of LAD estimates of the model parameters, which by bykh_JBES are consistent and asymptotically normal in the stationary region. (By contrast, due to the relatively greater frequency of censoring, OLS is not consistent in the stationary region, as discussed in Section 4.3 of that work.)
An implication of the arguments given in the proof of Corollary (ref), which is in line with alternative Tobit representation (ref), is that the OLS residuals $\{\hat{u}_t\}$ are consistent for $\{u_t-y_t^{-}\}$. However, because $y_t=0$ occurs relatively infrequently, $\{u_t-y_t^{-}\}$ and $\{u_t\}$ coincide sufficiently closely that e.g.\ $T^{-1}\sum_{t=1}^{T} (u_t-y_t^{-})^2 = T^{-1}\sum_{t=1}^{T} u_t^2 + o_p(1)$. In view of this, there appears to be some justification for continuing to use standard residual-based specification tests, such as for heteroskedasticity, serial correlation, or non-normality, when the dynamic Tobit is estimated by OLS (for an overview, see KL17book, Sec. 2.6--2.7). The proper second order asymptotics for such tests is outside of the focus of the present paper and is left for future research.
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{0.5\linespacing} {\normalfont}{Simulations}
We now illustrate how the values of the initial condition $b_{0}$ and the localising parameters $a$ and $c$ affect the distribution of $t_{\beta}$, and compare the performance of a test based on critical values derived from Corollary (ref) with one based on the conventional ADF critical values (DF_test), when the data is subject to censoring.
Figure (ref) depicts how a change in $b_{0}$ shifts the probability density (PDF) and the cumulative distribution (CDF) of $t_{\beta}$. As $b_{0}$ moves further above zero, the density shifts progressively to the right, as the probability that any trajectory of $K_{\theta}$ (initialised at $b_{0}$) will reach zero, and so be subject to censoring, correspondingly declines. Indeed, once $b_{0}$ is sufficiently large to make this probability negligible, the density becomes visually indistinguishable from that generated by a linear model (the solid green line in Figure (ref)), which is invariant to $b_{0}$. (Under the parametrisation used in the figure, this occurs when $b_{0}=2$; in general, this will depend on the magnitude of $\phi(1)b_{0}/\sigma$, in accordance with Theorem (ref).) The rightward shift of these distributions, as $b_0$ grows, is similarly manifest in the critical values given in Table (ref) above.
Surprisingly, the CDFs exhibit stochastic ordering, in the sense that the distributions corresponding to higher values of $b_{0}$ first-order stochastically dominate those with lower values of $b_0$ (holding all other parameters constant). This is in line with the monotonically increasing (in $b_0\phi(1)/\sigma$) quantiles in Table (ref). In particular, the null distribution of the conventional ADF $t$ test -- i.e.\ that appropriate to a linear autoregression -- stochastically dominates the null distribution appropriate to a dynamic Tobit. Therefore, when data is generated by a dynamic Tobit with a unit root, the conventional ADF test will tend to over-reject: in the worst case, which occurs when $b_0=0$, we find that a nominal $5$ per cent test will in fact reject $20$ per cent of the time. The intuition is that the censoring causes the trajectories of $\{y_{t}\}$ to appear stationary, masking the presence of a unit root.
For the remainder of this section, all simulations are conducted with $y_{0}=b_{0}=0$.
Figure (ref) shows how a change in local intercept $a$ (left panel) and local slope coefficient $c$ (right panel) affects the density of $t_{\beta}$. The means of these distributions across a range of values for $a$ and $c$ are also reported in Table (ref). We can see that as $a$ or $c$ fall further below zero, the distribution of $t_{\beta}$ (both its mean and its entire probability mass) shifts leftward -- with the opposite effect being observed when these parameters are progressively raised above zero.
The preceding illustrates how changes in $a$ and/or $c$ may shift the distribution of $t_{\beta}$ in either direction, and so will affect the ability of the test to reject the null of a unit root (i.e.\ $H_{0}:\alpha=0,\beta=1$). Power envelopes (rejection probabilities for a nominal 5 per cent, one-sided test), are displayed in Figure (ref). These show that, on the one side, more negative values of $a$ and/or $c$ make it easier for the test to reject the null in favour of a stationary alternative. The power eventually reaches $100$ per cent, indicating the consistency of the test against fixed alternatives in this region. This tendency to reject the null, as $a$ falls below zero, is in fact a desirable property of the test in this setting, since having $\alpha<0$ in the dynamic Tobit implies (for $k=1$) that $\{y_{t}\}$ is stationary, even when $\beta=1$ bykh_JBES.
On the other side, positive values of $c$ move $\{y_{t}\}$ into the explosive region, and our tendency to not reject in these cases is entirely consistent with the use of a one-sided test, as it is in a linear model (Figure (ref)). Although we also fail to reject for sufficiently positive values of $a$ (Figure (ref)), in such cases an upward trend in $\{y_{t}\}$ would become discernable, and would carry the process away from the censoring point. Since $\{y_{t}\}$ would then make few (if any) visits to the censoring point, if one were interested in testing the null of a unit root against a trend stationary alternative, in this case, a conventional ADF test with intercept and trend would be appropriate.
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{0.5\linespacing} {\normalfont}{Empirical illustration}
In this section we illustrate the use of our methods through an application to testing for unit roots in nominal exchange rates, when these are subject to a one-sided bound. Unit roots are routinely detected in these series by conventional tests in empirical work (see e.g.\ BB89JFE; SV06JBF, p. 3156; HP10JBES, p. 107). Their presence is also manifested in the robust performance of exchange rate forecasts based on random walks, which more elaborate models have struggled to beat consistently (Ros13JEL). (For a discussion of the theoretical basis for the presence of a unit root in exchange rates, in the context of open economy New Keynesian models, we refer the reader to Section 2.1 of engel2014exchange.) Here we examine how censoring, as introduced by the deliberate action of a central bank to keep exchange rates above (or below) a nominated threshold, may alter our assessment of the evidence for or against a unit root, depending on whether that censoring is accounted for (cf. cavaliere2005).
In September 2011, in response to the ongoing appreciation of the Swiss franc, and with the policy rate effectively at the zero lower bound, the Swiss National Bank (SNB) instituted a floor on the euro--Swiss franc exchange rate of 1.20 francs per euro (Jor16; Her22RIE). With the exchange rate well below the floor on the previous day, the immediate intervention of the SNB was required to make the floor effective upon its introduction on 6 September. The floor remained in effect until the end of 15 January 2015. As can be seen from Figure (ref), during that period, the exchange rate spent most of its time well above the the floor (reaching a peak of 1.26 in May 2013), with two notable exceptions: the periods of April to August 2012, and from November 2014 until the end of the policy, both of which triggered action by the SNB to prevent the floor from being violated (see Figure 9 in Her22RIE). (For a further discussion of the floor and its aftermath, see also vSTB21IRF.) The observed trajectory of the exchange rate, during these episodes, is thus more plausibly consistent with a dynamic Tobit than it is with a linear autoregression.
In this setting, the presence of a unit root would entail that the exchange rate tends to make extended sojourns away from the floor, visits to which accumulate only at rate $T^{1/2}$, where $T$ denotes the duration of the policy. By contrast, in the stationary dynamic Tobit, visits to the floor would accumulate at rate $T$. Thus the failure to reject the null of a unit root would signal to the central bank that it may need to intervene relatively infrequently in the foreign exchange market, in order to make the floor effective, something that is in turn highly relevant to the future cost and viability of maintaining that floor.
Our data is drawn from the European Central Bank's (ECB) daily reference exchange rate series (code: EXR.D.CHF.EUR.SP00.A) from the period during which the floor was operational (6 Sep 2011 to 15 Jan 2015), transformed by taking logarithms. To select the lag order $k$ used to compute $t_{\beta}$, we evaluated autoregressive models with $k\in\{1,\ldots,15\}$ using the Akaike and Bayesian information criteria; both selected a model with only one lag. For this model, we obtained an ADF statistic slightly above $-2.87$: so that if the censoring were ignored, and this statistic referred to the conventional ADF critical values (Table (ref)), there would be just sufficient evidence to reject the null of a unit root at the five per cent level. On the other hand, the unit root is not rejected at conventional significance levels under the dynamic Tobit. By simulating the asymptotic distribution of $t_{\beta}$, given in Corollary (ref), we compute the $p$-value appropriate to the dynamic Tobit as either $0.2$ (with $b_{0}=0$ imposed) or $0.18$ (using $\hat{b}_{0}=T^{-1/2}(y_{0}-\mathbf{L})$ with $\mathbf{L}=\log(1.20)$); similar results also obtain when $k=2,3$. This places $t_{\beta,T}$ much further from the critical region, thereby lending support to the hypothesis that the data is consistent with a dynamic Tobit model with a unit root.
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{0.5\linespacing} {\normalfont}{Conclusion}
This paper extends local to unity asymptotics to the setting of a dynamic Tobit model. Censoring fundamentally changes the analysis and requires new tools to derive the asymptotics. We obtain novel limit theorems for convergence to regulated processes, that is, to processes constrained to lie above a threshold. The effect of that censoring on the limiting distribution of our test statistics varies according to the proximity of the initialisation to the censoring point, with the distributions associated with the linear model re-emerging as that initialisation moves sufficiently far from the censoring point.
Our results underpin the development of a unit root test appropriate to censored data generated by a dynamic Tobit model. In contrast to the setting of a linear model, here the presence of a stochastic trend entails restrictions on both $\alpha$ and $\beta$, since $\alpha<0$ (with $\beta=1$) can be consistent with stationarity (and is consistent with stationarity for $k=1$). Nonetheless, a test of this null can still be effected by the usual $t_{\beta}$ statistic, using adjusted critical values. We provide an empirical illustration of our methods to testing for a unit root in nominal exchange rates, when these are subject to a one-sided bound.
The results of this paper could be developed further in a number of directions. One possibility would be to extend our results beyond the local to unity setting, by allowing for moderate deviations from a unit root (GP06JTSA), with the aim of establishing (uniformly) valid confidence intervals for $\beta$, as per Mik07Ecta,mikusheva2012one in the linear model. As we depart from the linear model, new types of asymptotics emerge, which may involve $c$ in different ways (cf. bykh_boundary where $c$ is no longer a constant but varies with time).
The analysis of $c\to\pm\infty$ for the censored model may be undertaken in conjunction with the derivation of the asymptotics of least absolute deviations regression or maximum likelihood estimators of the model, which would enjoy consistency across a wider domain than does OLS. For such extensions, the results of \prettyref{sec:yt_asymptotics}, regarding the asymptotics of $\{y_{t}\}$, are likely to be of fundamental importance.
For another direction in which our analysis could be extended, recall from Section (ref) that the dynamic Tobit provides the kernel of more elaborate, multivariate models that allow for the possibility of censoring and/or some other threshold-related nonlinearity. The analysis of (exact) unit roots and cointegration, in the setting of such a model, is developed by DMW22, with the aid of our results.