EconBase
← Back to paper

Estimation of a Dynamic Tobit Model with a Unit Root

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

47,549 characters · 12 sections · 60 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimation of a Dynamic Tobit Model with a Unit Root

\global\long\def\uwrite#1#2{\underset{#2}{\underbrace{#1}} }

\global\long\def\blw#1{\ensuremath{#1}}

\global\long\def\abv#1{\ensuremath{\overline{#1}}}

\global\long\def\vect#1{\mathbf{#1}}

\global\long\def\smlseq#1{\{#1\} }

\global\long\def\seq#1{\left\{ #1\right\} }

\global\long\def\smlsetof#1#2{\{#1\mid#2\} }

\global\long\def\setof#1#2{\left\{ #1\mid#2\right\} }

\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long

\global\long \global\long \global\long \global\long \global\long \global\long\def\Ellp#1{\ensuremath{\mathcal{L}^{#1}}}

\global\long \global\long \global\long \global\long \global\long

\global\long\def\abs#1{\ensuremath{\left|#1\right|}}

\global\long\def\smlabs#1{\ensuremath{\lvert#1\rvert}}

\global\long\def\bigabs#1{\ensuremath{\bigl|#1\bigr|}}

\global\long\def\Bigabs#1{\ensuremath{\Bigl|#1\Bigr|}}

\global\long\def\biggabs#1{\ensuremath{\biggl|#1\biggr|}}

\global\long\def\norm#1{\ensuremath{\left\Vert #1\right\Vert }}

\global\long\def\smlnorm#1{\ensuremath{\lVert#1\rVert}}

\global\long\def\bignorm#1{\ensuremath{\bigl\|#1\bigr\|}}

\global\long\def\Bignorm#1{\ensuremath{\Bigl\|#1\Bigr\|}}

\global\long\def\biggnorm#1{\ensuremath{\biggl\|#1\biggr\|}}

\global\long\def\floor#1{\left\lfloor #1\right\rfloor } \global\long\def\smlfloor#1{\lfloor#1\rfloor}

\global\long\def\ceil#1{\left\lceil #1\right\rceil } \global\long\def\smlceil#1{\lceil#1\rceil}

\global\long \global\long \global\long \global\long \global\long \global\long\def\clsr#1{\ensuremath{\overline{#1}}}

\global\long \global\long \global\long \global\long

\global\long\def\smlinprd#1#2{\ensuremath{\langle#1,#2\rangle}}

\global\long\def\inprd#1#2{\ensuremath{\left\langle #1,#2\right\rangle }}

\global\long \global\long

\global\long \global\long \global\long \global\long \global\long \global\long \global\long

\global\long \global\long \global\long\def\sigf#1{\mathcal{#1}}

\global\long\global\long \global\long\def\flt#1{\mathcal{#1}}

\global\long\global\long \global\long \global\long \global\long \global\long \global\long

\global\long \global\long \global\long \global\long \global\long \global\long \global\long

\global\long \global\long \global\long \global\long \def\independenT#1#2{\mathrel{\rlap{$#1#2$}\mkern2mu{#1#2}}}

\global\long \global\long \global\long \global\long \global\long \global\long\def\inprobu#1{\ensuremath{\overset{#1}{\ensuremath{\rightarrow}}}}

\global\long \global\long \global\long\def\inLp#1{\ensuremath{\overset{\Ellp{#1}}{\ensuremath{\rightarrow}}}}

\global\long \global\long \global\long \global\long\def\wkcu#1{\overset{#1}{\ensuremath{\rightsquigarrow}}}

\global\long

\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long

\global\long \global\long \global\long

\global\long \global\long \global\long \global\long \global\long \global\long\def\cv#1{\left\langle #1\right\rangle }

\global\long\def\smlcv#1{\langle#1\rangle}

\global\long\def\qv#1{\left[#1\right]}

\global\long\def\smlqv#1{[#1]}

\global\long \global\long \global\long \global\long \global\long\global\long \global\long\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long

comment\global\long\def\objlabel#1#2{\{\}} \global\long\def\objref#1{\{\}}

\global\long\def\vec#1{\boldsymbol{#1}}

\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long

\setcounter{footnote}{0}

abstractThis paper studies robust estimation in the dynamic Tobit model under local-to-unity (LUR) asymptotics. We show that both Gaussian maximum likelihood (ML) and censored least absolute deviations (CLAD) estimators are consistent, extending results from the stationary case where ordinary least squares (OLS) is inconsistent. The asymptotic distributions of MLE and CLAD are derived; for the short-run parameters they are shown to be Gaussian, yielding standard normal $t$-statistics. In contrast, although OLS remains consistent under LUR, its $t$-statistics are not standard normal. These results enable reliable model selection via sequential $t$-tests based on ML and CLAD, paralleling the linear autoregressive case. Applications to financial and epidemiological time series illustrate their practical relevance. Keywords: dynamic Tobit, unit root, MLE, CLAD.

\thispagestyle{plain}

\pagenumbering{roman}

\thispagestyle{plain}

\setcounter{tocdepth}{2}

\pagenumbering{arabic}

Introduction

Many macroeconomic and financial time series are highly persistent, often exhibiting the random wandering that is the hallmark of a stochastic trend. While it is relatively straightforward to accommodate these series within the framework of a linear (vector) autoregressive model, even in this setting such persistence creates significant challenges for inference, giving rise to limiting distributions that are non-standard and dependent on nuisance parameters. Accordingly, a substantial literature has developed on inference methods that are broadly robust to temporal dependence, remaining valid for both stationary and stochastically trending data (see e.g.\ Hansen99REStat; JM06Ecta; Mik07Ecta,mikusheva2012one; EMW15Ecta).

Though this literature has now attained a high degree of maturity, a significant limitation is that it is entirely devoted to linear models. Although nonlinear autoregressive models have been widely studied, this has almost exclusively been in the context of stationary processes (as in e.g.\ Tong90; Chan09; TTG10). Indeed, for a long time the only available dynamic models that permitted some conjunction of nonlinearities with stochastic trends were those in which the nonlinearities were confined to the short-run dynamics, i.e.\ to first differences and equilibrium error components (as in the nonlinear VECM models considered by BF97IER; HS02JoE; KR10JoE). This excluded, for example, nonlinear autoregressive models with regimes endogenously dependent on the level of a stochastically trending series. Only recently has it been shown possible to configure nonlinear (vector) autoregressive models so as to accommodate both nonlinearity in levels and stochastic trends (cavaliere2005; BD22; DM24; DMW22).

A major impetus for this recent work has come from the desire to adequately model highly persistent series that are subject to occasionally binding constraints, of which the zero lower bound on nominal interest rates is a leading example (Mav21Ecta, AMSV21JoE). For such series, a plausible descriptive model is the dynamic Tobit, which in a stationary setting has a lengthy pedigree in the empirical analysis of constrained, i.e.\ censored, time series (see e.g.\ Demiralp_Jorda2002, stat_paper, DSK12AE, LMS19, BMMV21JBF, bykh_JBES). BD22 recently determined conditions under which this model may also generate stochastically trending, $I(1)$-like series, and obtained the asymptotic distribution of OLS estimators in a local-to-unity framework. However, while their results are sufficient for the purposes of conducting unit root tests, there is no possibility of extending the validity of OLS beyond their setting, since even its consistency is known to fail for the stationary dynamic Tobit.

The present work is thus motivated by the need to develop methods of estimation and inference, in a nonlinear time series setting, that are robust to the degree of persistence of the data generating mechanism. More specifically, we seek methods that enjoy validity for the dynamic Tobit model on a wider domain that allows for both stochastically trending and stationary data, exactly as OLS enjoys in a linear autoregressive model. This leads us to consider two alternative estimators that are known to be consistent and asymptotically normal for the stationary dynamic Tobit: Gaussian maximum likelihood (ML) and Pow84JoE's Pow84JoE censored least absolute deviations (CLAD; see stat_paper, and bykh_JBES).

The principal contribution of this paper is to derive the asymptotic distributions of the ML and CLAD estimators for a dynamic Tobit model, in a local-to-unity framework. When the model is rendered in ADF form, the limiting distribution of the estimator of the sum of the autoregressive coefficients is non-standard, but takes a known form that is suitable for inference. Moreover, estimates of `short run' coefficients (i.e.\ those on lagged differences) have an asymptotically Gaussian distribution, so that inferences on these and on certain related functionals, such as short-horizon impulse responses, remain standard irrespective of the degree of persistence of the regressor (exactly as is the case for OLS estimation in a linear autoregressive model). These results facilitate the determination of the model lag order on the basis of sequential $t$-tests.

At a technical level, our results contribute to the literature on the asymptotics of nonlinear extremum estimators with highly peristent processes (see e.g.\ PP00Ecta,PP01Ecta; PJH07JoE; Xiao09JoE; CW15JoE). Relative to previous work, the analysis is complicated by two aspects of our problem. Firstly, we consider a nonlinear autoregressive model, whereas the literature has been concerned with regression (or discrete choice) models where the regressor $x_{t}$ is a linear unit root process, and the object is to estimate some (possibly) nonlinear transformation of $x_{t}$. Secondly, in our analysis of the CLAD estimator, we cannot rely on the convexity of the criterion function, something which greatly facilitates the derivation of the asymptotics of the uncensored LAD estimator, as in Poll91ET, Herce96ET, LL09ET and Xiao09JoE. (By contrast, convexity may be fruitfully exploited in the analysis of the Gaussian MLE.) Nor may approaches employed in the i.i.d.\ or stationary cases Pow84JoE,stat_paper,bykh_JBES be readily transposed to our setting. Thus deriving the asymptotics of the (centred) CLAD criterion function, in particular, requires some delicate and novel arguments. We expect these arguments will also prove useful for the analysis of other nonlinear extremum estimators, in the presence of highly persistent data.

The remainder of this paper is organised as follows. Section (ref) introduces the dynamic Tobit model and our assumptions. Section (ref) derives the asymptotic distributions of ML and CLAD estimators. A discussion of the results, and an empirical illustration, is presented in Section (ref). Finally, Section (ref) concludes. All proofs appear in the appendices.

notation*All limits are taken as $T\ensuremath{\rightarrow}\infty$ unless otherwise stated. $\ensuremath{\overset{p}{\ensuremath{\rightarrow}}}$ and $\ensuremath{\overset{d}{\ensuremath{\rightarrow}}}$ respectively denote convergence in probability and in distribution (weak convergence). We write `$X_{T}(\lambda)\ensuremath{\overset{d}{\ensuremath{\rightarrow}}} X(\lambda)$ on $D_{\mathbb{R}^{m}}[0,1]$' to denote that $\{X_{T}\}$ converges weakly to $X$, where these are considered as random elements of $D_{\mathbb{R}^{m}}[0,1]$, the space of cadlag functions $[0,1]\ensuremath{\rightarrow}\mathbb{R}^{m}$, equipped with the uniform topology; we denote this as $D[0,1]$ whenever the value of $m$ is clear from the context. $\smlnorm{\cdot}$ denotes the Euclidean norm on $\mathbb{R}^{m}$, and the matrix norm that it induces. For $X$ a random vector and $p\geq1$, $\smlnorm X_{p}\coloneqq(\ensuremath{\mathbb{E}}\smlnorm X^{p})^{1/p}$.

Model

We suppose $\{y_t\}$ is generated by a dynamic Tobit model of order $k$, written in augmented Dickey--Fuller (ADF) form as

equation[equation omitted — 153 chars of source]

where $[x]_+:=\max\{x,0\}$, $\vec y_{t-1}:=(y_{t-1},\ldots, y_{t-k+1})^{\mathsf{T}}$, $\vec{\phi}_0=(\phi_{1,0},\ldots,\phi_{k-1,0})^{\mathsf{T}}$ and $\Delta y_{t}:=y_t-y_{t-1}$. The model has two attractive features. Firstly, it is Markovian -- with the state vector defined by $k$ consecutive lags of $y_t$ -- making it well-suited for forecasting. Secondly, the presence of the positive part on the right of the equality enforces a lower bound of zero on $y_t$, making the model appropriate for series that are constrained to take non-negative values.

We assume that the parameters $\rho_0:=(\alpha_0,\beta_0,\vec{\phi}_0^{\mathsf{T}})^{\mathsf{T}}\in\mathbb{R}^{k+1}$, innovations $\{u_t\}$ and initial conditions satisfy the following assumptions.

{{{{\scalefont{0.76}A1}}}}

assumption$\{y_{t}\}$ is initialised by (possibly) random variables $\{y_{-k+1},\ldots,y_{0}\}$. Moreover, $T^{-1/2}y_{0}\ensuremath{\overset{p}{\ensuremath{\rightarrow}}} b_{0}$ for some $b_{0}\geq0$.

{{{{\scalefont{0.76}A2}}}}

assumption$\{y_{t}\}$ is generated according to ((ref)), where: \begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}] • $\{u_{t}\}_{t\in\mathbb{Z}}$ is independently and identically distributed with $\ensuremath{\mathbb{E}} u_{t}=0$ and $\ensuremath{\mathbb{E}} u_{t}^{2}=\sigma_{0}^{2}$. • $\alpha_0=\alpha_{T,0}\coloneqq T^{-1/2}a_{0}$ and $\beta_0=\beta_{T,0}=1+T^{-1}c_{0}$, for some $a_{0},c_{0}\in\mathbb{R}$. \end{enumerate}

{{{{\scalefont{0.76}A3}}}}

assumptionThere exist $\delta_{u}>0$ and $C<\infty$ such that: \begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}] • $\ensuremath{\mathbb{E}}\smlabs{u_{t}}^{2+\delta_{u}}<C$. • $\ensuremath{\mathbb{E}}\smlabs{T^{-1/2}y_{0}}^{2+\delta_{u}}<C$, and $\ensuremath{\mathbb{E}}\smlabs{\Delta y_{i}}^{2+\delta_{u}}<C$ for $i\in\{-k+2,\ldots,0\}$. \end{enumerate}

Assumption (ref) allows the initialisation to be of the same order of magnitude as the process $y_t$. This is appropriate because the first observation in any sample is unlikely to be the `true' starting point of the process, but simply the point at which data collection begins. Therefore, it is natural to allow $y_0$ to be of the same scale as any later $y_t$. Assumptions (ref)(ref) and (ref) are standard conditions on the innovations. Assumption (ref)(ref) ensures that the data is highly persistent, as a consequence of the autoregressive polynomial having a root local to unity.

As shown in BD22, owing to the nonlinearity of the model additional technical conditions -- which go beyond the classical requirements on the stability of the roots of the autoregressive polynomial -- are needed to rule out explosive behaviour of the first differences $\{\Delta y_t\}$. Here we ensure this through a constraint on the joint spectral radius of two auxiliary companion-form matrices. For $\delta\in[0,1]$, define

equation[equation omitted — 217 chars of source]

and let $\lambda_{{\scriptstyle \mathrm{JSR}}}({\cal A})$ denote the joint spectral radius of a bounded collection of matrices ${\cal A}$, which may be defined as \[ \lambda_{{\scriptstyle \mathrm{JSR}}}(\mathcal{A})\coloneqq\limsup_{n\ensuremath{\rightarrow}\infty}\sup_{M\in\mathcal{A}^{n}}\lambda(M)^{1/n} \] where $\lambda(M)$ denotes the spectral radius of $M$, and $\mathcal{A}^{n}\coloneqq\{\prod_{i=1}^{n}A_{i}\mid A_{i}\in\mathcal{A}\}$ (cf., Jungers09).

{{{{\scalefont{0.76}A4}}}}

assumption$\lambda_{{\scriptstyle \mathrm{JSR}}}(\{F_{0},F_{1}\})<1$.

Approximate upper bounds for the joint spectral radius (JSR) can be computed numerically to an arbitrarily high degree of accuracy using semidefinite programming PJ08LAA, making it possible to verify whether the condition is satisfied for a given set of parameter values (see DMW23stat, for a further discussion). Assumption (ref) ensures that the process $y_t$ does not exhibit explosive behavior by jointly controlling the dynamics across both the censored ($\delta=0$) and uncensored ($\delta=1$) `regimes'. Without this restriction the interaction between the two regimes could generate a `bouncing' effect, where hitting the lower bound triggers a strong rebound, potentially leading to exponential growth.

BD22 show that under Assumptions (ref)-(ref)

equation[equation omitted — 272 chars of source]

on $D[0,1]$, where

equation*[equation* omitted — 208 chars of source]

for $\phi(1)=1-\sum_{i=1}^{k-1}\phi_{i,0}>0$, and $W(\cdot)$ a standard Brownian motion. The limiting distribution of the estimators of $\alpha$ and $\beta$, though not of $\vec{\phi}$, will be shown below to depend on the process $Y(\cdot)$ -- a dependence similar to that of the OLS estimators on the limiting Brownian Motion (or Ornstein--Uhlenbeck process) in a linear autoregressive model with a root at (or local to) unity. In fact, the process $K(\cdot)$ -- which $Y(\cdot)$ is the constrained counterpart of -- corresponds exactly to the limit of a local-to-unity autoregressive process. So surprisingly, despite the Tobit model's added complexity due to censoring, a similar asymptotic structure emerges: the limiting distributions of our estimators will closely resemble those of a linear autoregression, albeit with a nonlinear limiting process $Y(\cdot)$.

Estimation and inference

This section introduces two estimation methods -- Gaussian maximum likelihood (ML) and censored least absolute deviations (CLAD) -- and establishes their asymptotic properties. We show that the estimators of the coefficients $\vec{\phi}_0$ on the stochastically bounded (though not stationary) lags $\Delta \vec {y}_{t-1}$ are asymptotically normal under both methods, permitting standard inferences to be drawn. As discussed further in Section (ref), this contrasts sharply with the behaviour of ordinary least squares estimators, which have some limiting bias (of order $T^{-1/2}$) due to the censoring, even in the presence of a local to unit root.

Maximum likelihood

To permit estimation by maximum likelihood, we need to make a parametric assumption regarding the distribution of the innovations $\{u_{t}\}$. A conventional choice -- which is particularly attractive for computational reasons, because it yields a loglikelihood with a convex reparametrisation -- is for $u_{t}$ to be Gaussian, as per the following.

{{{{\scalefont{0.76}B}}}}

assumption$u_{t}\sim N[0,\sigma_{0}^{2}]$, and $\delta_{u}>2$ in (ref).

Note that for $\{u_{t}\}$ itself, the maintained Gaussianity makes the condition on $\delta_{u}$ redundant; this condition is nonetheless still needed instead to ensure that the initial conditions also have a little more than four finite moments.

Under assumption (ref), when $\{y_{t}\}$ is generated according to ((ref)), the conditional density of $y_{t}$ given $(y_{t-1},\ldots,y_{t-k})$, evaluated at the values $y_{t-i}=\mathsf{y}_{t-i}$ for $i \in {0, \ldots, k}$, is

equation[equation omitted — 450 chars of source]

where $\mathsf{y}_{t-1} = (\mathsf{y}_{t-1},\ldots,\mathsf{y}_{t-k+1})^{\mathsf{T}}$, and $\Phi(x)$ and $\varphi(x)$ denote the standard normal cdf and pdf functions, i.e.\ $\varphi(x)\coloneqq\tfrac1{\sqrt{2\pi}}\mathrm{e}^{-x^{2}/2}$. We consider a conditional ML estimator that maximises the loglikelihood of the $\{y_{t}\}_{t=1}^{T}$ conditional on the initial values $\{y_{i}\}_{i=-k+1}^{0}$, i.e.\ which maximises

equation*[equation* omitted — 159 chars of source]

with respect to $\alpha,\beta,\vec{\phi},\sigma$. Define \[ (\hat{\alpha}_{T}^{{\scriptscriptstyle \mathsf{M}}},\hat{\beta}_{T}^{{\scriptscriptstyle \mathsf{M}}},\hat{\vec{\phi}}{}_{T}^{{\scriptscriptstyle \mathsf{M}}},\hat{\sigma}_{T}^{{\scriptscriptstyle \mathsf{M}}})\in\operatorname*{argmax}_{(\alpha,\beta,\vec{\phi},\sigma)\in\mathbb{R}^{k+2}\times\mathbb{R}_{+}}{\cal L}_{T}(\alpha,\beta,\vec{\phi},\sigma) \] to be a sequence of maximisers of ${\cal L}_{T}$. Let $\Omega\coloneqq\operatorname*{plim}_{T\to\infty}\frac1{T}\sum_{t=1}^{T}\Delta\mathbf{y}_{t-1}\Delta\mathbf{y}_{t-1}^{\mathsf{T}}$, which exists as a consequence of Lemma B.4 in BD22.

thmSuppose (ref)--(ref) and (ref) hold. Then \begin{equation} \begin{bmatrix}T^{1/2}(\hat{\alpha}_{T}^{{\scriptscriptstyle \mathsf{M}}}-\alpha_{0})\\ T(\hat{\beta}_{T}^{{\scriptscriptstyle \mathsf{M}}}-\beta_{0}) \end{bmatrix}\xrightarrow[T\to\infty]{d}\sigma_0\begin{bmatrix}1 & \int_{0}^{1}Y(\tau)\ensuremath{\,\ensuremath{\mathrm{d}}}\tau\\ \int_{0}^{1}Y(\tau)\ensuremath{\,\ensuremath{\mathrm{d}}}\tau & \int_{0}^{1}Y^{2}(\tau)\ensuremath{\,\ensuremath{\mathrm{d}}}\tau \end{bmatrix}^{-1}\begin{bmatrix}W(1)\\ \int_{0}^{1}Y(\tau)\ensuremath{\,\ensuremath{\mathrm{d}}} W(\tau) \end{bmatrix} \end{equation} jointly with \begin{equation} T^{1/2}\begin{bmatrix}\hat{\vec{\phi}}_{T}^{{\scriptscriptstyle \mathsf{M}}}-\vec{\phi}_{0}\\ (\hat{\sigma}_{T}^{{\scriptscriptstyle \mathsf{M}}})^{2}-\sigma_{0}^{2} \end{bmatrix}\xrightarrow[T\to\infty]{d}\begin{bmatrix}\sigma_{0}^{2}\Omega^{-1} & 0\\ 0 & (2\sigma_{0}^{2})^{-1} \end{bmatrix}^{1/2}\xi \end{equation} where $\xi\sim \mathcal{N}[0,I_{k}]$ is independent of $W$ (and therefore also $Y$).

In contrast to the classical autoregressive setting, where the Gaussian MLE is numerically identical to OLS, and therefore has the same asymptotic distribution, Theorem (ref) and simulations in Section (ref) show that censoring breaks this equivalence (see BD22 for the OLS limit theory). Interestingly, though, up to a change of the rescaled limit of $y_t$, the ML distributions in Theorem (ref), match their linear analogues, see, e.g., Hamilton94.

The asymptotic variance in (ref) is given by the inverse of the information matrix, i.e., the inverse of the negative Hessian (see Proposition (ref)), matching the corresponding result in the stationary case stat_paper. Unlike in the stationary case, where censoring binds much more frequently ($O(T)$ rather than $O(\sqrt{T})$ times), here we can obtain an explicit form for the Hessian, which is unaffected by censored observations apart from their impact on the asymptotic distribution of $y_t$ in (ref).

Censored least absolute deviations

The censored least absolute deviations (CLAD) estimator of ((ref)), which we denote as $(\hat{\alpha}_{T}^{{\scriptscriptstyle \mathsf{L}}}, \hat{\beta}_{T}^{{\scriptscriptstyle \mathsf{L}}}, \hat{\vec{\phi}}{}_T^{{\scriptscriptstyle \mathsf{L}}})$, minimises

equation[equation omitted — 187 chars of source]

as per Pow84JoE. The presence of the positive part -- reflecting Tobit-type censoring as in (ref) -- within the objective function (ref) renders the criton function non-convex, making the analysis of the CLAD estimator rather more challenging than that of its uncensored counterpart. A further complication arises from the high persistence in $\{y_t\}$, which prevents the application of maximal inequalities or uniform central limit theorems appropriate to stationary processes.

{{{{\scalefont{0.76}C}}}}

assumption\begin{enumerate}[label={{{\scalefont{0.76}\arabic*.}}}, ref={{{\scalefont{0.76}.\arabic*}}}] • $u_{t}$ is continuously distributed, with density $f_{u}$ that is positive at zero, bounded, and continuously differentiable with bounded derivatives, $\operatorname{med}(u_{t})=0$ and $\delta_{u}>2$ in (ref); and • $(\alpha_0,\beta_0,\vec{\phi}^{\mathsf{T}}_0)^{\mathsf{T}}\in\Pi$, for some compact $\Pi\subset\mathbb{R}^{k+2}$, and $(0,1,\vec{\phi}_{0}^{\mathsf{T}})^{\mathsf{T}}\in\operatorname{int}\Pi$. \end{enumerate}

We expect that our requirement that $u_t$ have a little more than four finite moments ($\delta_{u}>2$) could be relaxed to two finite moments. Obtaining the limiting distribution in the latter case would necessitate a sharper bound on the order of

equation*[equation* omitted — 137 chars of source]

which is closely related to the rate at which censored observations occur. The arguments used to prove the following theorem rely on an estimate of $o_p(T)$ for such terms, but this could likely be improved to $O_p(T^{1/2})$.

thmSuppose that Assumptions (ref)-(ref) and (ref) hold. Then \begin{equation} \begin{bmatrix}T^{1/2}(\hat{\alpha}_{T}^{{\scriptscriptstyle \mathsf{L}}}-\alpha_{0})\\ T(\hat{\beta}_{T}^{{\scriptscriptstyle \mathsf{L}}}-\beta_{0}) \end{bmatrix} \xrightarrow[T\to\infty]{d} \frac1{2f_u(0)}\begin{bmatrix}1 & \int_{0}^{1}Y(\tau)\ensuremath{\,\ensuremath{\mathrm{d}}}\tau\\ \int_{0}^{1}Y(\tau)\ensuremath{\,\ensuremath{\mathrm{d}}}\tau & \int_{0}^{1}Y^{2}(\tau)\ensuremath{\,\ensuremath{\mathrm{d}}}\tau \end{bmatrix}^{-1}\begin{bmatrix}\widetilde{W}(1)\\ \int_0^1 Y(\tau)d\widetilde{W}(\tau) \end{bmatrix} \end{equation} jointly with, and independently of \begin{equation} \sqrt{T}(\hat{\vec{\phi}}_T^{{\scriptscriptstyle \mathsf{L}}}-\vec{\phi}_0) \xrightarrow[T\to\infty]{d} \frac1{2f_u(0)}\mathcal{N}(0,\Omega^{-1}), \end{equation} where $\Omega=\operatorname*{plim}_{T\to\infty}\frac1{T}\sum_{t=1}^{T}\Delta\mathbf{y}_{t-1}\Delta\mathbf{y}_{t-1}^{\mathsf{T}}$ and $\widetilde{W}(\cdot)$ is a $1$-dimensional Brownian motion, such that the covariance matrix of $(\sigma_0 W(\cdot),\widetilde{W}(\cdot))^{\mathsf{T}}$ is $$\Sigma=\begin{pmatrix} \sigma_0^2 & \mathbb{E}|u_t| \\ \mathbb{E}|u_t| & 1 \end{pmatrix}.$$

Comparing our asymptotic results with those for (uncensored) LAD in the linear regression setting of Herce96ET, we see that the structure of our limting distributions closely mirrors its linear counterpart, with the only difference being the limiting processes for $\{y_t\}$ (upon rescaling) that appears in the two distributional limits. This resemblance is particularly surprising given that the observations with $y_t=0$ cannot be ignored; rather, they play a crucial role in the derivation of Theorem (ref).

There, the key step involves deriving the first- and second-order terms in the expansion of the CLAD objective function evaluated at an arbitrary parameter vector relative to its value at the true parameters. The first-order term converges to a quadratic form in the parameters, while the second-order term converges to a linear function. This structure enables us to characterize the solution to the limiting optimization problem.

It is also worth comparing the results of Theorem (ref) with their stationary counterparts stat_paper and bykh_JBES. Both share the same factor $\frac1{2f_u(0)}$, where the density at zero arises from a Taylor expansion and governs the behavior when $y_{t-1}$ is close to zero. Moreover, the asymptotics for the short-run coefficients $\hat{\vec{\phi}}^{{\scriptscriptstyle \mathsf{L}}}$ match their stationary analogue, since the number of observations with $y_t=0$ is small enough that $\frac1{T}\sum_t\Delta\mathbf{y}_{t-1}\Delta\mathbf{y}_{t-1}^{\mathsf{T}}$ and $\frac1{T}\sum_t\ensuremath{\mathbf{1}}\{y_{t}>0\}\Delta\mathbf{y}_{t-1}\Delta\mathbf{y}_{t-1}^{\mathsf{T}}$ are asymptotically identical.

Discussion

Pros and cons of CLAD and MLE

When $u_t$ follows a standard normal distribution, both Theorems (ref) and (ref) are applicable, allowing us to compare the variances of the short-run coefficients in MLE and CLAD. In this case, $\tfrac{1}{2f_u(0)} = \sigma_0\sqrt{\pi/2} \approx 1.25\sigma_0$, implying roughly a $25\%$ increase in the standard deviation of CLAD relative to the (optimal) ML estimator. The increased asymptotic variance of all CLAD coefficients compared to MLE is illustrated in Figure (ref). We observe that, because the asymptotic distributions of both CLAD and MLE estimators of $\alpha$ and $\beta$ involve the non-standard limiting process $Y(\cdot)$, they are skewed and centered at positive values for $\alpha$ and negative values for $\beta$, with the MLE exhibiting tighter concentration than CLAD. In contrast, the short-run estimators are all centered at zero and approximately normal, though again with CLAD displaying higher variance than MLE.

figure[figure omitted — 751 chars of source]
figure[figure omitted — 816 chars of source]

On the other hand, relative to the (Gaussian) MLE, a major advantage of the CLAD estimator is that it is semiparametric, insofar as it does not require a parametric assumption on the distribution of the errors $u_t$. As illustrated in Section (ref) and previously in stat_paper,bykh_JBES, real-world data can deviate substantially from Gaussianity, in which case the CLAD estimator is likely to yield more reliable inferences.

Our simulations suggest that the asymptotic behavior of MLE changes little when the errors are non-Gaussian, so long as they have sufficient moments, consistent with the well-attested good performance of quasi-MLEs. However, depending on the value of $f_u(0)$, CLAD may now be more efficient than the quasi-MLE. For example, if $u_t$ follows a Laplace$(0,1/\sqrt{2})$ distribution, so that $\mathbb{E}u_t=0$ and $\mathbb{E} u_t^2=1$, we have $f_u(0) = 1/\sqrt{2}$ and $1/2f_u(0)= 1/\sqrt{2} \approx 0.7 < 1$. In this case, the MLE standard deviation for the short-run coefficients is almost $1.5$ times larger than that of CLAD. This is illustrated in Figure (ref), where the behaviour of the MLE estimator is very similar to that observed for Gaussian errors in Figure (ref), while the CLAD estimator is now much more tightly distributed, reflecting the higher value of the density at zero ($1/\sqrt{2}$ for the Laplace distribution versus $1/\sqrt{2\pi}$ for the standard normal). The same pattern is observed when the errors follow a $t$-distribution with $\nu$ degrees of freedom, where $\nu \in (2, 4.6]$ to ensure that variance is well-defined and $2f_u(0) > 1$.

The main drawbacks of the CLAD estimator come from two aspects. First, the CLAD optimisation problem is non-convex, potentially causing numerical optimisers to converge local rather than global optima, and thus to be sensitive to their initialisation. The Gaussian loglikelihood, by contrast, may be reparameterised so as to render the objective function concave, thereby avoiding this problem.\footnote{As per Ols78Ecta, the new parameters are $\alpha/\sigma,\, \beta/\sigma,\,\vec{\phi}/\sigma$, and $\sigma^{-1}$.} (This may in turn be used to provide reasonable starting values for the optimisation of the CLAD criterion function.) The second drawback stems from the necessity of estimating $f_u(0)$ for the purposes of inference (i.e.\ to construct confidence intervals or conduct hypothesis tests). We tackle this by means of kernel density estimation, as proposed in Pow84JoE, though this has the undesirable consequence of rendering inferences dependent on the choice of a bandwidth parameter.

Comparison with OLS

It is noteworthy that all three estimators share a common structure in their asymptotic distributions: the inverse of $$

bmatrix[bmatrix omitted — 211 chars of source]

$$ appears in the asymptotics of the the estimators of $(\alpha,\beta)$, while the limiting variance of the estimators of $\vec{\phi}$ depends on the inverse of $\Omega$. This structure mirrors the limit of the properly rescaled OLS regressor signal matrix $X^{\mathsf{T}}X$, for the regressors $(1,y_{t-1},\Delta \vec{y}_{t-1}^\mathsf{T})^\mathsf{T}$. The similarity arises because, up to proportionality, the first-order Taylor expansions of the estimators coincide; the differences emerge in the behavior of fluctuations, that is, in the second-order terms.

There are two significant advantages to using MLE or CLAD relative to OLS. The first advantage is that the former two are consistent both under stationary and (near) unit root regimes, see e.g., stat_paper,bykh_JBES, while OLS is consistent only under the (near) unit root regime (see, e.g., bykh_JBES for inconsistency results and BD22 for consistency results). Since an empirical researcher is typically a priori uncertain about the level of persistence in the data, using MLE or LAD permits valid inferences to be drawn more robustly.

The second advantage is related to the estimation of the `short-run' parameters $\phi_i$, for $i \in \{1,\ldots,k-1\}$, i.e.\ the coefficients on the lagged differences $\Delta y_{t-i}$. As is shown in Theorems (ref) and (ref), both the MLE and CLAD estimators of theses are asymptotically normal, leading to standard normal t-statistics. By contrast, as illustrated in Figure (ref), the OLS asymptotic distribution of $\hat{\phi}_{1,OLS}$ and its corresponding $t$-statistic (solid blue curves) are not centered at zero, as a consequence of the censoring. In particular, the mean of the OLS $t$-statistic in Figure (ref) is approximately $-0.4$. This negative mean represents the non-degenerate limit of $\frac1{\sqrt{T}}\sum_t y_t^{-}\Delta y_{t-1}$, where $y_t^{-}=\min\{0,\alpha_0+\beta_0 y_{t-1}+\vec{\phi}_0^{\mathsf{T}}\Delta\vec y_{t-1}+u_{t}\}$. The intuition is that the positive-part binds $O(\sqrt{T})$ times, so there are $O(\sqrt{T})$ nonzero $y_t^{-}$, which, when scaled by $\tfrac{1}{\sqrt{T}}$, generate a negative bias. In contrast, the simulated distributions of the $t$-statistics for the ML (red dashed curve) and CLAD (black dash-dotted curve) estimators align closely with the standard normal density.

figure[figure omitted — 669 chars of source]

Model selection and illustrations

The fact that the $t$-statistics for $\phi_i$ are asymptotically standard normal, for both MLE and CLAD, can be used for the purposes of model selection. In particular, one can determine the appropriate lag order $k$ for the model via a sequential test of the null $H_0 : \phi_{k-1} = 0$, starting from some relatively high value $k=k_0$, and then reducing $k$ (by one) so long as the associated $t$-statistic does not exceed the $\alpha$-level normal critical value. In contrast, since the OLS $t$-statistics for the $\phi_i$ parameters have a non-standard limiting distribution, the naive use of these in such a model selection procedure may lead to an inappropriate choice for $k$.

This is illustrated below with two data sets. For each we estimate a $k$th order dynamic Tobit, for various values of $k$, using MLE and CLAD. For each $k$, and each estimation method, we compute the corresponding $t$-statistic for $H_0: \phi_{k-1} = 0$. The standard errors of the estimators are obtained from Theorems (ref) and (ref). The matrix $\Omega$ is estimated by its sample analogue, while the density at zero, required for the CLAD estimator, is estimated using a uniform kernel, as recommended by Pow84JoE.

US Treasury bill rate

figure[figure omitted — 514 chars of source]

Treasury bill rates in the United States have been near the zero lower bound (ZLB) during three historical periods: during the $1930$s and $1940$s following the Great Depression, during and after the $2007$ Global Financial Crisis, and more recently during the COVID-$19$ pandemic. We use quarterly interest rates on $3$-month Treasury bills, with data extending back to the $1920$s. The data set is based on ramey2018government and extended using FRED data through the first quarter of $2025$, so that $T=421$. We treat values below $0.5$ as effectively at the zero lower bound, so that our identified ZLB periods closely align with those in ramey2018government. Figure (ref) displays the interest rate data, while Figure (ref) presents the corresponding $t$-statistics computed for $k \in \{2,\ldots,20\}$ (i.e., up to $5$ years of quarterly lags).

Comparing the $t$-statistics in Figure (ref) with the $5\%$ normal critical value for a two-sided test ($1.96$), we find that in most cases both estimation methods yield the same results, and point to the model with $k=8$ as the preferred specification. There are, however, three values of $k$, $3$, $6$, and $14$, where the methods yield slightly different conclusions, as the CLAD $t$-statistic falls below $1.96$ for $k=3,6$ and above it for $k=14$. We report the estimation results for $k=8$ in Table (ref).

table[table omitted — 669 chars of source]
comment\begin{figure}[t] \begin{subfigure}{.49\textwidth} \caption{MLE.} \end{subfigure} \begin{subfigure}{.49\textwidth} \caption{LAD} \end{subfigure} \caption{3-month Tbills: $90\%$ confidence sets for $\alpha$ and $\beta$.} \end{figure}

Countywide Covid-19 cases

figure[figure omitted — 518 chars of source]

Understanding and predicting the evolution of pandemic case counts is critical for effective public health planning and response. Accurate forecasts can inform timely policy decisions such as implementing social distancing measures, allocating medical resources, or guiding vaccine distribution. This stresses the importance of valid model selection. To illustrate our approach to pandemic model selection, we use Covid-19 case data for Rowan County, NC, obtained from covid_USfacts. Rowan County provides an instructive example due to its moderate population size and relatively localised outbreaks.

We use daily data for $351$ consecutive days, stopping before reporting irregularities become prevalent: when missing cases began to be accumulated into subsequent days. The sample begins on March 19, 2020 (the date of the first reported case) and ends on March 4, 2021. Unlike the previous example, we do not apply artificial censoring; the $22$ zero observations correspond to days when no positive cases were recorded.

Figure (ref) summarizes the data and the associated $t$-statistics. In contrast to the previous example, we now observe substantial disagreement between LAD and MLE, with LAD favoring more complex models. This divergence suggests that the data may not be well approximated by a normal distribution, consistent with findings in the literature showing that Covid-19 case counts deviate significantly from both normality and log-normality (e.g., gonccalves2021covid,cohen2022covid).

Conclusion

This paper develops asymptotic theory for maximum likelihood (ML) and censored least absolute deviations (CLAD) estimation in dynamic Tobit models under local unit root (LUR) asymptotics. While earlier work had established consistency (and asymptotic normality) of these estimators in the stationary setting, we show that both ML and CLAD remain consistent in the LUR regime and derive their asymptotic distributions. A key finding is that the short-run parameters are asymptotically normal under both methods, facilitatig standard inference and model selection via sequential $t$-testing.

These results contrast sharply with the behavior of ordinary least squares (OLS), which, although consistent under LUR, yields $t$-statistics for the short-run parameters that have an asymptotic bias and non-standard distribution. Consequently, model selection procedures based on OLS can be misleading, especially in settings with censoring and persistent dynamics.

Overall, our results suggest that MLE and LAD offer viable and robust alternatives to OLS in dynamic censored models, with significant advantages for inference and model selection. Future work could extend the analysis to accommodate time-varying volatility, covariates, or alternative forms of censoring commonly encountered in economic and financial data.