Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
47,549 characters · 12 sections · 60 citation commands
Estimation of a Dynamic Tobit Model with a Unit Root
\global\long\def\uwrite#1#2{\underset{#2}{\underbrace{#1}} }
\global\long\def\blw#1{\ensuremath{#1}}
\global\long\def\abv#1{\ensuremath{\overline{#1}}}
\global\long\def\vect#1{\mathbf{#1}}
\global\long\def\smlseq#1{\{#1\} }
\global\long\def\seq#1{\left\{ #1\right\} }
\global\long\def\smlsetof#1#2{\{#1\mid#2\} }
\global\long\def\setof#1#2{\left\{ #1\mid#2\right\} }
\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long \global\long \global\long \global\long\def\Ellp#1{\ensuremath{\mathcal{L}^{#1}}}
\global\long \global\long \global\long \global\long \global\long
\global\long\def\abs#1{\ensuremath{\left|#1\right|}}
\global\long\def\smlabs#1{\ensuremath{\lvert#1\rvert}}
\global\long\def\bigabs#1{\ensuremath{\bigl|#1\bigr|}}
\global\long\def\Bigabs#1{\ensuremath{\Bigl|#1\Bigr|}}
\global\long\def\biggabs#1{\ensuremath{\biggl|#1\biggr|}}
\global\long\def\norm#1{\ensuremath{\left\Vert #1\right\Vert }}
\global\long\def\smlnorm#1{\ensuremath{\lVert#1\rVert}}
\global\long\def\bignorm#1{\ensuremath{\bigl\|#1\bigr\|}}
\global\long\def\Bignorm#1{\ensuremath{\Bigl\|#1\Bigr\|}}
\global\long\def\biggnorm#1{\ensuremath{\biggl\|#1\biggr\|}}
\global\long\def\floor#1{\left\lfloor #1\right\rfloor } \global\long\def\smlfloor#1{\lfloor#1\rfloor}
\global\long\def\ceil#1{\left\lceil #1\right\rceil } \global\long\def\smlceil#1{\lceil#1\rceil}
\global\long \global\long \global\long \global\long \global\long \global\long\def\clsr#1{\ensuremath{\overline{#1}}}
\global\long \global\long \global\long \global\long
\global\long\def\smlinprd#1#2{\ensuremath{\langle#1,#2\rangle}}
\global\long\def\inprd#1#2{\ensuremath{\left\langle #1,#2\right\rangle }}
\global\long \global\long
\global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long\def\sigf#1{\mathcal{#1}}
\global\long\global\long \global\long\def\flt#1{\mathcal{#1}}
\global\long\global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long \global\long \def\independenT#1#2{\mathrel{\rlap{$#1#2$}\mkern2mu{#1#2}}}
\global\long \global\long \global\long \global\long \global\long \global\long\def\inprobu#1{\ensuremath{\overset{#1}{\ensuremath{\rightarrow}}}}
\global\long \global\long \global\long\def\inLp#1{\ensuremath{\overset{\Ellp{#1}}{\ensuremath{\rightarrow}}}}
\global\long \global\long \global\long \global\long\def\wkcu#1{\overset{#1}{\ensuremath{\rightsquigarrow}}}
\global\long
\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long \global\long \global\long
\global\long \global\long \global\long \global\long \global\long \global\long\def\cv#1{\left\langle #1\right\rangle }
\global\long\def\smlcv#1{\langle#1\rangle}
\global\long\def\qv#1{\left[#1\right]}
\global\long\def\smlqv#1{[#1]}
\global\long \global\long \global\long \global\long \global\long\global\long \global\long\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long
\global\long\def\vec#1{\boldsymbol{#1}}
\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long
\setcounter{footnote}{0}
\thispagestyle{plain}
\pagenumbering{roman}
\thispagestyle{plain}
\setcounter{tocdepth}{2}
\pagenumbering{arabic}
Many macroeconomic and financial time series are highly persistent, often exhibiting the random wandering that is the hallmark of a stochastic trend. While it is relatively straightforward to accommodate these series within the framework of a linear (vector) autoregressive model, even in this setting such persistence creates significant challenges for inference, giving rise to limiting distributions that are non-standard and dependent on nuisance parameters. Accordingly, a substantial literature has developed on inference methods that are broadly robust to temporal dependence, remaining valid for both stationary and stochastically trending data (see e.g.\ Hansen99REStat; JM06Ecta; Mik07Ecta,mikusheva2012one; EMW15Ecta).
Though this literature has now attained a high degree of maturity, a significant limitation is that it is entirely devoted to linear models. Although nonlinear autoregressive models have been widely studied, this has almost exclusively been in the context of stationary processes (as in e.g.\ Tong90; Chan09; TTG10). Indeed, for a long time the only available dynamic models that permitted some conjunction of nonlinearities with stochastic trends were those in which the nonlinearities were confined to the short-run dynamics, i.e.\ to first differences and equilibrium error components (as in the nonlinear VECM models considered by BF97IER; HS02JoE; KR10JoE). This excluded, for example, nonlinear autoregressive models with regimes endogenously dependent on the level of a stochastically trending series. Only recently has it been shown possible to configure nonlinear (vector) autoregressive models so as to accommodate both nonlinearity in levels and stochastic trends (cavaliere2005; BD22; DM24; DMW22).
A major impetus for this recent work has come from the desire to adequately model highly persistent series that are subject to occasionally binding constraints, of which the zero lower bound on nominal interest rates is a leading example (Mav21Ecta, AMSV21JoE). For such series, a plausible descriptive model is the dynamic Tobit, which in a stationary setting has a lengthy pedigree in the empirical analysis of constrained, i.e.\ censored, time series (see e.g.\ Demiralp_Jorda2002, stat_paper, DSK12AE, LMS19, BMMV21JBF, bykh_JBES). BD22 recently determined conditions under which this model may also generate stochastically trending, $I(1)$-like series, and obtained the asymptotic distribution of OLS estimators in a local-to-unity framework. However, while their results are sufficient for the purposes of conducting unit root tests, there is no possibility of extending the validity of OLS beyond their setting, since even its consistency is known to fail for the stationary dynamic Tobit.
The present work is thus motivated by the need to develop methods of estimation and inference, in a nonlinear time series setting, that are robust to the degree of persistence of the data generating mechanism. More specifically, we seek methods that enjoy validity for the dynamic Tobit model on a wider domain that allows for both stochastically trending and stationary data, exactly as OLS enjoys in a linear autoregressive model. This leads us to consider two alternative estimators that are known to be consistent and asymptotically normal for the stationary dynamic Tobit: Gaussian maximum likelihood (ML) and Pow84JoE's Pow84JoE censored least absolute deviations (CLAD; see stat_paper, and bykh_JBES).
The principal contribution of this paper is to derive the asymptotic distributions of the ML and CLAD estimators for a dynamic Tobit model, in a local-to-unity framework. When the model is rendered in ADF form, the limiting distribution of the estimator of the sum of the autoregressive coefficients is non-standard, but takes a known form that is suitable for inference. Moreover, estimates of `short run' coefficients (i.e.\ those on lagged differences) have an asymptotically Gaussian distribution, so that inferences on these and on certain related functionals, such as short-horizon impulse responses, remain standard irrespective of the degree of persistence of the regressor (exactly as is the case for OLS estimation in a linear autoregressive model). These results facilitate the determination of the model lag order on the basis of sequential $t$-tests.
At a technical level, our results contribute to the literature on the asymptotics of nonlinear extremum estimators with highly peristent processes (see e.g.\ PP00Ecta,PP01Ecta; PJH07JoE; Xiao09JoE; CW15JoE). Relative to previous work, the analysis is complicated by two aspects of our problem. Firstly, we consider a nonlinear autoregressive model, whereas the literature has been concerned with regression (or discrete choice) models where the regressor $x_{t}$ is a linear unit root process, and the object is to estimate some (possibly) nonlinear transformation of $x_{t}$. Secondly, in our analysis of the CLAD estimator, we cannot rely on the convexity of the criterion function, something which greatly facilitates the derivation of the asymptotics of the uncensored LAD estimator, as in Poll91ET, Herce96ET, LL09ET and Xiao09JoE. (By contrast, convexity may be fruitfully exploited in the analysis of the Gaussian MLE.) Nor may approaches employed in the i.i.d.\ or stationary cases Pow84JoE,stat_paper,bykh_JBES be readily transposed to our setting. Thus deriving the asymptotics of the (centred) CLAD criterion function, in particular, requires some delicate and novel arguments. We expect these arguments will also prove useful for the analysis of other nonlinear extremum estimators, in the presence of highly persistent data.
The remainder of this paper is organised as follows. Section (ref) introduces the dynamic Tobit model and our assumptions. Section (ref) derives the asymptotic distributions of ML and CLAD estimators. A discussion of the results, and an empirical illustration, is presented in Section (ref). Finally, Section (ref) concludes. All proofs appear in the appendices.
We suppose $\{y_t\}$ is generated by a dynamic Tobit model of order $k$, written in augmented Dickey--Fuller (ADF) form as
where $[x]_+:=\max\{x,0\}$, $\vec y_{t-1}:=(y_{t-1},\ldots, y_{t-k+1})^{\mathsf{T}}$, $\vec{\phi}_0=(\phi_{1,0},\ldots,\phi_{k-1,0})^{\mathsf{T}}$ and $\Delta y_{t}:=y_t-y_{t-1}$. The model has two attractive features. Firstly, it is Markovian -- with the state vector defined by $k$ consecutive lags of $y_t$ -- making it well-suited for forecasting. Secondly, the presence of the positive part on the right of the equality enforces a lower bound of zero on $y_t$, making the model appropriate for series that are constrained to take non-negative values.
We assume that the parameters $\rho_0:=(\alpha_0,\beta_0,\vec{\phi}_0^{\mathsf{T}})^{\mathsf{T}}\in\mathbb{R}^{k+1}$, innovations $\{u_t\}$ and initial conditions satisfy the following assumptions.
{{{{\scalefont{0.76}A1}}}}
{{{{\scalefont{0.76}A2}}}}
{{{{\scalefont{0.76}A3}}}}
Assumption (ref) allows the initialisation to be of the same order of magnitude as the process $y_t$. This is appropriate because the first observation in any sample is unlikely to be the `true' starting point of the process, but simply the point at which data collection begins. Therefore, it is natural to allow $y_0$ to be of the same scale as any later $y_t$. Assumptions (ref)(ref) and (ref) are standard conditions on the innovations. Assumption (ref)(ref) ensures that the data is highly persistent, as a consequence of the autoregressive polynomial having a root local to unity.
As shown in BD22, owing to the nonlinearity of the model additional technical conditions -- which go beyond the classical requirements on the stability of the roots of the autoregressive polynomial -- are needed to rule out explosive behaviour of the first differences $\{\Delta y_t\}$. Here we ensure this through a constraint on the joint spectral radius of two auxiliary companion-form matrices. For $\delta\in[0,1]$, define
and let $\lambda_{{\scriptstyle \mathrm{JSR}}}({\cal A})$ denote the joint spectral radius of a bounded collection of matrices ${\cal A}$, which may be defined as \[ \lambda_{{\scriptstyle \mathrm{JSR}}}(\mathcal{A})\coloneqq\limsup_{n\ensuremath{\rightarrow}\infty}\sup_{M\in\mathcal{A}^{n}}\lambda(M)^{1/n} \] where $\lambda(M)$ denotes the spectral radius of $M$, and $\mathcal{A}^{n}\coloneqq\{\prod_{i=1}^{n}A_{i}\mid A_{i}\in\mathcal{A}\}$ (cf., Jungers09).
{{{{\scalefont{0.76}A4}}}}
Approximate upper bounds for the joint spectral radius (JSR) can be computed numerically to an arbitrarily high degree of accuracy using semidefinite programming PJ08LAA, making it possible to verify whether the condition is satisfied for a given set of parameter values (see DMW23stat, for a further discussion). Assumption (ref) ensures that the process $y_t$ does not exhibit explosive behavior by jointly controlling the dynamics across both the censored ($\delta=0$) and uncensored ($\delta=1$) `regimes'. Without this restriction the interaction between the two regimes could generate a `bouncing' effect, where hitting the lower bound triggers a strong rebound, potentially leading to exponential growth.
BD22 show that under Assumptions (ref)-(ref)
on $D[0,1]$, where
for $\phi(1)=1-\sum_{i=1}^{k-1}\phi_{i,0}>0$, and $W(\cdot)$ a standard Brownian motion. The limiting distribution of the estimators of $\alpha$ and $\beta$, though not of $\vec{\phi}$, will be shown below to depend on the process $Y(\cdot)$ -- a dependence similar to that of the OLS estimators on the limiting Brownian Motion (or Ornstein--Uhlenbeck process) in a linear autoregressive model with a root at (or local to) unity. In fact, the process $K(\cdot)$ -- which $Y(\cdot)$ is the constrained counterpart of -- corresponds exactly to the limit of a local-to-unity autoregressive process. So surprisingly, despite the Tobit model's added complexity due to censoring, a similar asymptotic structure emerges: the limiting distributions of our estimators will closely resemble those of a linear autoregression, albeit with a nonlinear limiting process $Y(\cdot)$.
This section introduces two estimation methods -- Gaussian maximum likelihood (ML) and censored least absolute deviations (CLAD) -- and establishes their asymptotic properties. We show that the estimators of the coefficients $\vec{\phi}_0$ on the stochastically bounded (though not stationary) lags $\Delta \vec {y}_{t-1}$ are asymptotically normal under both methods, permitting standard inferences to be drawn. As discussed further in Section (ref), this contrasts sharply with the behaviour of ordinary least squares estimators, which have some limiting bias (of order $T^{-1/2}$) due to the censoring, even in the presence of a local to unit root.
To permit estimation by maximum likelihood, we need to make a parametric assumption regarding the distribution of the innovations $\{u_{t}\}$. A conventional choice -- which is particularly attractive for computational reasons, because it yields a loglikelihood with a convex reparametrisation -- is for $u_{t}$ to be Gaussian, as per the following.
{{{{\scalefont{0.76}B}}}}
Note that for $\{u_{t}\}$ itself, the maintained Gaussianity makes the condition on $\delta_{u}$ redundant; this condition is nonetheless still needed instead to ensure that the initial conditions also have a little more than four finite moments.
Under assumption (ref), when $\{y_{t}\}$ is generated according to ((ref)), the conditional density of $y_{t}$ given $(y_{t-1},\ldots,y_{t-k})$, evaluated at the values $y_{t-i}=\mathsf{y}_{t-i}$ for $i \in {0, \ldots, k}$, is
where $\mathsf{y}_{t-1} = (\mathsf{y}_{t-1},\ldots,\mathsf{y}_{t-k+1})^{\mathsf{T}}$, and $\Phi(x)$ and $\varphi(x)$ denote the standard normal cdf and pdf functions, i.e.\ $\varphi(x)\coloneqq\tfrac1{\sqrt{2\pi}}\mathrm{e}^{-x^{2}/2}$. We consider a conditional ML estimator that maximises the loglikelihood of the $\{y_{t}\}_{t=1}^{T}$ conditional on the initial values $\{y_{i}\}_{i=-k+1}^{0}$, i.e.\ which maximises
with respect to $\alpha,\beta,\vec{\phi},\sigma$. Define \[ (\hat{\alpha}_{T}^{{\scriptscriptstyle \mathsf{M}}},\hat{\beta}_{T}^{{\scriptscriptstyle \mathsf{M}}},\hat{\vec{\phi}}{}_{T}^{{\scriptscriptstyle \mathsf{M}}},\hat{\sigma}_{T}^{{\scriptscriptstyle \mathsf{M}}})\in\operatorname*{argmax}_{(\alpha,\beta,\vec{\phi},\sigma)\in\mathbb{R}^{k+2}\times\mathbb{R}_{+}}{\cal L}_{T}(\alpha,\beta,\vec{\phi},\sigma) \] to be a sequence of maximisers of ${\cal L}_{T}$. Let $\Omega\coloneqq\operatorname*{plim}_{T\to\infty}\frac1{T}\sum_{t=1}^{T}\Delta\mathbf{y}_{t-1}\Delta\mathbf{y}_{t-1}^{\mathsf{T}}$, which exists as a consequence of Lemma B.4 in BD22.
In contrast to the classical autoregressive setting, where the Gaussian MLE is numerically identical to OLS, and therefore has the same asymptotic distribution, Theorem (ref) and simulations in Section (ref) show that censoring breaks this equivalence (see BD22 for the OLS limit theory). Interestingly, though, up to a change of the rescaled limit of $y_t$, the ML distributions in Theorem (ref), match their linear analogues, see, e.g., Hamilton94.
The asymptotic variance in (ref) is given by the inverse of the information matrix, i.e., the inverse of the negative Hessian (see Proposition (ref)), matching the corresponding result in the stationary case stat_paper. Unlike in the stationary case, where censoring binds much more frequently ($O(T)$ rather than $O(\sqrt{T})$ times), here we can obtain an explicit form for the Hessian, which is unaffected by censored observations apart from their impact on the asymptotic distribution of $y_t$ in (ref).
The censored least absolute deviations (CLAD) estimator of ((ref)), which we denote as $(\hat{\alpha}_{T}^{{\scriptscriptstyle \mathsf{L}}}, \hat{\beta}_{T}^{{\scriptscriptstyle \mathsf{L}}}, \hat{\vec{\phi}}{}_T^{{\scriptscriptstyle \mathsf{L}}})$, minimises
as per Pow84JoE. The presence of the positive part -- reflecting Tobit-type censoring as in (ref) -- within the objective function (ref) renders the criton function non-convex, making the analysis of the CLAD estimator rather more challenging than that of its uncensored counterpart. A further complication arises from the high persistence in $\{y_t\}$, which prevents the application of maximal inequalities or uniform central limit theorems appropriate to stationary processes.
{{{{\scalefont{0.76}C}}}}
We expect that our requirement that $u_t$ have a little more than four finite moments ($\delta_{u}>2$) could be relaxed to two finite moments. Obtaining the limiting distribution in the latter case would necessitate a sharper bound on the order of
which is closely related to the rate at which censored observations occur. The arguments used to prove the following theorem rely on an estimate of $o_p(T)$ for such terms, but this could likely be improved to $O_p(T^{1/2})$.
Comparing our asymptotic results with those for (uncensored) LAD in the linear regression setting of Herce96ET, we see that the structure of our limting distributions closely mirrors its linear counterpart, with the only difference being the limiting processes for $\{y_t\}$ (upon rescaling) that appears in the two distributional limits. This resemblance is particularly surprising given that the observations with $y_t=0$ cannot be ignored; rather, they play a crucial role in the derivation of Theorem (ref).
There, the key step involves deriving the first- and second-order terms in the expansion of the CLAD objective function evaluated at an arbitrary parameter vector relative to its value at the true parameters. The first-order term converges to a quadratic form in the parameters, while the second-order term converges to a linear function. This structure enables us to characterize the solution to the limiting optimization problem.
It is also worth comparing the results of Theorem (ref) with their stationary counterparts stat_paper and bykh_JBES. Both share the same factor $\frac1{2f_u(0)}$, where the density at zero arises from a Taylor expansion and governs the behavior when $y_{t-1}$ is close to zero. Moreover, the asymptotics for the short-run coefficients $\hat{\vec{\phi}}^{{\scriptscriptstyle \mathsf{L}}}$ match their stationary analogue, since the number of observations with $y_t=0$ is small enough that $\frac1{T}\sum_t\Delta\mathbf{y}_{t-1}\Delta\mathbf{y}_{t-1}^{\mathsf{T}}$ and $\frac1{T}\sum_t\ensuremath{\mathbf{1}}\{y_{t}>0\}\Delta\mathbf{y}_{t-1}\Delta\mathbf{y}_{t-1}^{\mathsf{T}}$ are asymptotically identical.
When $u_t$ follows a standard normal distribution, both Theorems (ref) and (ref) are applicable, allowing us to compare the variances of the short-run coefficients in MLE and CLAD. In this case, $\tfrac{1}{2f_u(0)} = \sigma_0\sqrt{\pi/2} \approx 1.25\sigma_0$, implying roughly a $25\%$ increase in the standard deviation of CLAD relative to the (optimal) ML estimator. The increased asymptotic variance of all CLAD coefficients compared to MLE is illustrated in Figure (ref). We observe that, because the asymptotic distributions of both CLAD and MLE estimators of $\alpha$ and $\beta$ involve the non-standard limiting process $Y(\cdot)$, they are skewed and centered at positive values for $\alpha$ and negative values for $\beta$, with the MLE exhibiting tighter concentration than CLAD. In contrast, the short-run estimators are all centered at zero and approximately normal, though again with CLAD displaying higher variance than MLE.
On the other hand, relative to the (Gaussian) MLE, a major advantage of the CLAD estimator is that it is semiparametric, insofar as it does not require a parametric assumption on the distribution of the errors $u_t$. As illustrated in Section (ref) and previously in stat_paper,bykh_JBES, real-world data can deviate substantially from Gaussianity, in which case the CLAD estimator is likely to yield more reliable inferences.
Our simulations suggest that the asymptotic behavior of MLE changes little when the errors are non-Gaussian, so long as they have sufficient moments, consistent with the well-attested good performance of quasi-MLEs. However, depending on the value of $f_u(0)$, CLAD may now be more efficient than the quasi-MLE. For example, if $u_t$ follows a Laplace$(0,1/\sqrt{2})$ distribution, so that $\mathbb{E}u_t=0$ and $\mathbb{E} u_t^2=1$, we have $f_u(0) = 1/\sqrt{2}$ and $1/2f_u(0)= 1/\sqrt{2} \approx 0.7 < 1$. In this case, the MLE standard deviation for the short-run coefficients is almost $1.5$ times larger than that of CLAD. This is illustrated in Figure (ref), where the behaviour of the MLE estimator is very similar to that observed for Gaussian errors in Figure (ref), while the CLAD estimator is now much more tightly distributed, reflecting the higher value of the density at zero ($1/\sqrt{2}$ for the Laplace distribution versus $1/\sqrt{2\pi}$ for the standard normal). The same pattern is observed when the errors follow a $t$-distribution with $\nu$ degrees of freedom, where $\nu \in (2, 4.6]$ to ensure that variance is well-defined and $2f_u(0) > 1$.
The main drawbacks of the CLAD estimator come from two aspects. First, the CLAD optimisation problem is non-convex, potentially causing numerical optimisers to converge local rather than global optima, and thus to be sensitive to their initialisation. The Gaussian loglikelihood, by contrast, may be reparameterised so as to render the objective function concave, thereby avoiding this problem.\footnote{As per Ols78Ecta, the new parameters are $\alpha/\sigma,\, \beta/\sigma,\,\vec{\phi}/\sigma$, and $\sigma^{-1}$.} (This may in turn be used to provide reasonable starting values for the optimisation of the CLAD criterion function.) The second drawback stems from the necessity of estimating $f_u(0)$ for the purposes of inference (i.e.\ to construct confidence intervals or conduct hypothesis tests). We tackle this by means of kernel density estimation, as proposed in Pow84JoE, though this has the undesirable consequence of rendering inferences dependent on the choice of a bandwidth parameter.
It is noteworthy that all three estimators share a common structure in their asymptotic distributions: the inverse of $$
$$ appears in the asymptotics of the the estimators of $(\alpha,\beta)$, while the limiting variance of the estimators of $\vec{\phi}$ depends on the inverse of $\Omega$. This structure mirrors the limit of the properly rescaled OLS regressor signal matrix $X^{\mathsf{T}}X$, for the regressors $(1,y_{t-1},\Delta \vec{y}_{t-1}^\mathsf{T})^\mathsf{T}$. The similarity arises because, up to proportionality, the first-order Taylor expansions of the estimators coincide; the differences emerge in the behavior of fluctuations, that is, in the second-order terms.
There are two significant advantages to using MLE or CLAD relative to OLS. The first advantage is that the former two are consistent both under stationary and (near) unit root regimes, see e.g., stat_paper,bykh_JBES, while OLS is consistent only under the (near) unit root regime (see, e.g., bykh_JBES for inconsistency results and BD22 for consistency results). Since an empirical researcher is typically a priori uncertain about the level of persistence in the data, using MLE or LAD permits valid inferences to be drawn more robustly.
The second advantage is related to the estimation of the `short-run' parameters $\phi_i$, for $i \in \{1,\ldots,k-1\}$, i.e.\ the coefficients on the lagged differences $\Delta y_{t-i}$. As is shown in Theorems (ref) and (ref), both the MLE and CLAD estimators of theses are asymptotically normal, leading to standard normal t-statistics. By contrast, as illustrated in Figure (ref), the OLS asymptotic distribution of $\hat{\phi}_{1,OLS}$ and its corresponding $t$-statistic (solid blue curves) are not centered at zero, as a consequence of the censoring. In particular, the mean of the OLS $t$-statistic in Figure (ref) is approximately $-0.4$. This negative mean represents the non-degenerate limit of $\frac1{\sqrt{T}}\sum_t y_t^{-}\Delta y_{t-1}$, where $y_t^{-}=\min\{0,\alpha_0+\beta_0 y_{t-1}+\vec{\phi}_0^{\mathsf{T}}\Delta\vec y_{t-1}+u_{t}\}$. The intuition is that the positive-part binds $O(\sqrt{T})$ times, so there are $O(\sqrt{T})$ nonzero $y_t^{-}$, which, when scaled by $\tfrac{1}{\sqrt{T}}$, generate a negative bias. In contrast, the simulated distributions of the $t$-statistics for the ML (red dashed curve) and CLAD (black dash-dotted curve) estimators align closely with the standard normal density.
The fact that the $t$-statistics for $\phi_i$ are asymptotically standard normal, for both MLE and CLAD, can be used for the purposes of model selection. In particular, one can determine the appropriate lag order $k$ for the model via a sequential test of the null $H_0 : \phi_{k-1} = 0$, starting from some relatively high value $k=k_0$, and then reducing $k$ (by one) so long as the associated $t$-statistic does not exceed the $\alpha$-level normal critical value. In contrast, since the OLS $t$-statistics for the $\phi_i$ parameters have a non-standard limiting distribution, the naive use of these in such a model selection procedure may lead to an inappropriate choice for $k$.
This is illustrated below with two data sets. For each we estimate a $k$th order dynamic Tobit, for various values of $k$, using MLE and CLAD. For each $k$, and each estimation method, we compute the corresponding $t$-statistic for $H_0: \phi_{k-1} = 0$. The standard errors of the estimators are obtained from Theorems (ref) and (ref). The matrix $\Omega$ is estimated by its sample analogue, while the density at zero, required for the CLAD estimator, is estimated using a uniform kernel, as recommended by Pow84JoE.
Treasury bill rates in the United States have been near the zero lower bound (ZLB) during three historical periods: during the $1930$s and $1940$s following the Great Depression, during and after the $2007$ Global Financial Crisis, and more recently during the COVID-$19$ pandemic. We use quarterly interest rates on $3$-month Treasury bills, with data extending back to the $1920$s. The data set is based on ramey2018government and extended using FRED data through the first quarter of $2025$, so that $T=421$. We treat values below $0.5$ as effectively at the zero lower bound, so that our identified ZLB periods closely align with those in ramey2018government. Figure (ref) displays the interest rate data, while Figure (ref) presents the corresponding $t$-statistics computed for $k \in \{2,\ldots,20\}$ (i.e., up to $5$ years of quarterly lags).
Comparing the $t$-statistics in Figure (ref) with the $5\%$ normal critical value for a two-sided test ($1.96$), we find that in most cases both estimation methods yield the same results, and point to the model with $k=8$ as the preferred specification. There are, however, three values of $k$, $3$, $6$, and $14$, where the methods yield slightly different conclusions, as the CLAD $t$-statistic falls below $1.96$ for $k=3,6$ and above it for $k=14$. We report the estimation results for $k=8$ in Table (ref).
Understanding and predicting the evolution of pandemic case counts is critical for effective public health planning and response. Accurate forecasts can inform timely policy decisions such as implementing social distancing measures, allocating medical resources, or guiding vaccine distribution. This stresses the importance of valid model selection. To illustrate our approach to pandemic model selection, we use Covid-19 case data for Rowan County, NC, obtained from covid_USfacts. Rowan County provides an instructive example due to its moderate population size and relatively localised outbreaks.
We use daily data for $351$ consecutive days, stopping before reporting irregularities become prevalent: when missing cases began to be accumulated into subsequent days. The sample begins on March 19, 2020 (the date of the first reported case) and ends on March 4, 2021. Unlike the previous example, we do not apply artificial censoring; the $22$ zero observations correspond to days when no positive cases were recorded.
Figure (ref) summarizes the data and the associated $t$-statistics. In contrast to the previous example, we now observe substantial disagreement between LAD and MLE, with LAD favoring more complex models. This divergence suggests that the data may not be well approximated by a normal distribution, consistent with findings in the literature showing that Covid-19 case counts deviate significantly from both normality and log-normality (e.g., gonccalves2021covid,cohen2022covid).
This paper develops asymptotic theory for maximum likelihood (ML) and censored least absolute deviations (CLAD) estimation in dynamic Tobit models under local unit root (LUR) asymptotics. While earlier work had established consistency (and asymptotic normality) of these estimators in the stationary setting, we show that both ML and CLAD remain consistent in the LUR regime and derive their asymptotic distributions. A key finding is that the short-run parameters are asymptotically normal under both methods, facilitatig standard inference and model selection via sequential $t$-testing.
These results contrast sharply with the behavior of ordinary least squares (OLS), which, although consistent under LUR, yields $t$-statistics for the short-run parameters that have an asymptotic bias and non-standard distribution. Consequently, model selection procedures based on OLS can be misleading, especially in settings with censoring and persistent dynamics.
Overall, our results suggest that MLE and LAD offer viable and robust alternatives to OLS in dynamic censored models, with significant advantages for inference and model selection. Future work could extend the analysis to accommodate time-varying volatility, covariates, or alternative forms of censoring commonly encountered in economic and financial data.