EconBase
← Back to paper

Testing for Threshold Effects in Presence of Heteroskedasticity and Measurement Error with an application to Italian Strikes

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

58,061 characters · 10 sections · 67 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing for Threshold Effects in Presence of Heteroskedasticity and Measurement Error with an application to Italian Strikes

\def\spacingset#1{ {#1}} \spacingset{1}

\if00 \fi

\if10 {

center[center omitted — 158 chars of source]

} \fi

abstractMany macroeconomic time series are characterised by nonlinearity both in the conditional mean and in the conditional variance and, in practice, it is important to investigate separately these two aspects. Here we address the issue of testing for threshold nonlinearity in the conditional mean, in the presence of conditional heteroskedasticity. We propose a supremum Lagrange Multiplier approach to test a linear ARMA-GARCH model against the alternative of a TARMA-GARCH model. We derive the asymptotic null distribution of the test statistic and this requires novel results since the difficulties of working with nuisance parameters, absent under the null hypothesis, are amplified by the non-linear moving average, combined with GARCH-type innovations. We show that tests that do not account for heteroskedasticity fail to achieve the correct size even for large sample sizes. Moreover, we show that the TARMA specification naturally accounts for the ubiquitous presence of measurement error that affects macroeconomic data. We apply the results to analyse the time series of Italian strikes and we show that the TARMA-GARCH specification is consistent with the relevant macroeconomic theory while capturing the main features of the Italian strikes dynamics, such as asymmetric cycles and regime-switching.

{\it Keywords:} Threshold autoregressive moving-average model; GARCH model; Lagrange multiplier test.

\spacingset{2}

Introduction

Strike activity of labour unions is linked to the business cycle and to workers' organizational and political power. Hence, understanding strikes' dynamics can shed light on related social and political issues and support the economic analysis of industrial conflicts in several countries. It is acknowledged that the dynamics of strikes presents complex features, but to the best of our knowledge, both the economic and the econometric literature appears to lack a comprehensive analysis that considers most of these aspects in a unique framework pal1982,paldam2022,franzosi1989,hundley1987. The use of strikes as a tool of political and organizational power differs from country to country depending on their specific institutional context franzosi1989,castellani2013. In some countries, unions use strikes to influence public policies and the reasons behind them are mostly political, while in other unionised economies, such as Italy, strikes are also used as a bargaining tool in the labour market. Due to its historical and institutional background, the Italian strikes time series is well-suited to investigate the dynamics of this economic and political variable lange1990,corneo1997,quaranta2012. Strike activity has a very long history in Western countries and presents some stylized facts, i.e. they appear cyclically and occur in waves. For example, during the 1970s and 1980s, OECD countries experienced a long strike wave, while a sharp decline in strikes occurred in the last decades godard2011. During the 1990s most European countries and the United States witnessed a significant decrease in strike activity, together with a strong shift towards the “tertiarisation of conflict” bordogna2002. More recently, strikes against governments have been increasingly engaged by unions across Western Europe hamann2013. A traditional explanation connects the time series of strikes with the business cycle dynamics. ree1952 argued that strikes should increase during booms and decrease during the down-swing phase of the business cycle. Wage claims are another cause of strikes.\footnote{hic1932 pointed out that industrial conflict is costly for both workers and firms. Thus, if agents were rational and fully informed, there would be no reason to strike. Although employees usually strike to gain higher wages or better working conditions, its efficacy remains controversial. sha1980 surveyed the literature on industrial relations theory.} However, the empirical evidence for OECD countries seems to indicate that strikes are procyclical and wage increases have a negative effect on strikes. Qualitative analyses suggest that strikes dynamics is characterized by volatility with clustering effects, primarily due to certain types of strikes that can attract more workers or have prolonged durations vandaele2016. Together with the observed asymmetric cyclical behaviour, this hints at the presence of a regime-switching mechanism with conditional heteroskedasticity. We advocate the threshold autoregressive-moving average (TARMA) model with GARCH innovations as a novel and appropriate specification to deal with the aforementioned features of strikes dynamics. Threshold nonlinearity offers a feasible approximation of general complex dynamics while retaining a good interpretability. Threshold autoregressive models Ton78,Ton80 have been widely applied in Economics hansen2011,Cha17b and Finance chen2011. Threshold models are particularly suitable to describe the phenomenon of regulation which plays a fundamental role in Finance and Economics. For instance, in financial time series it is common to observe a “band of inaction” random walk regime, where arbitrage does not occur, and other regimes where mean reversion takes place so that the model is globally stationary, see e.g. Cha20b. Moreover, pesaran1997 used the TAR model to show that the U.S. GDP is also subject to floor and ceiling effects. koop1999 estimate TAR models for U.S. unemployment rates and altissimo2001 propose a VTAR to study the joint dynamics of U.S. GDP and unemployment rates. Note that, none of these studies included a moving-average component in their specifications, most probably due to the lack of developments on non-linear ARMA models. TARMA models combine the well-known threshold autoregressive (TAR) model and the threshold moving-average (TMA) model Lin05. The incorporation of the moving-average component in a non-linear parametric framework achieves a great approximating capability with few parameters Gor21. They allow to interpret phenomena that change qualitatively across regimes and react differently to shocks, which is a key aspect in macroeconomic dynamics, as also pointed out in Gon21. Moreover, as shown in Cha20b, they naturally account for the presence of measurement error. Despite these advantages, the theoretical development of TARMA models has been stuck for many years due to unsolved theoretical problems, mainly due to their non-Markovian nature. The impasse has been overcome by Cha19, which solved the long-standing open problem regarding the probabilistic structure of the first-order TARMA model. This paved the way for substantial theoretical inferential developments and practical applications.

The main aim of this paper is to introduce a test for threshold ARMA effects with GARCH innovations and ascertain whether the dynamics of the Italian strikes can be adequately modeled using a TARMA-GARCH specification. Existing tests for threshold effects on the conditional mean are negatively affected by the presence of conditional heteroscedasticity. A partial solution to this problem is to use threshold tests sequentially on the residuals of a GARCH model but this is likely to affect non-trivially the overall significance level of the tests. Moreover, the presence of measurement error would require a full TARMA specification but existing tests do not account for this. In particular, Won97 test an AR-ARCH versus a TAR-ARCH specification, while Li08 compare an MA-GARCH against a TMA-GARCH. We fill the gap and test the ARMA-GARCH against the TARMA-GARCH specification using a sup-Lagrange Multiplier test tailored to this scope. We derive the asymptotic null distribution of the test statistic and this requires novel results since the inherent difficulties of working with nuisance parameters absent under the null hypothesis are amplified by the non-linear moving average setting combined with GARCH-type innovations. We show that, in presence of heteroskedasticity, tests that assume i.i.d. innovations can be severely biased, whereas our test can be used successfully in all those cases where the aim is testing for a non-linear (threshold) effect in the conditional mean but the series also presents conditional heteroskedasticity. We use the novel results for the analysis of strikes dynamics, by using Italian data covering the period between January 1949 and December 2009. We identify a regime-switching dynamics with GARCH innovations. We discuss the shortcomings of the linear approach and showcase the adequacy of our non-linear specification, which turns out to be consistent both with observed stylized facts and macroeconomic theories, rooted in the Italian labour market history and on more general social dynamics of labour markets. The remainder of the paper is organized as follows. In Section (ref) we introduce our sup-Lagrange multiplier test for threshold effects with GARCH innovations. We present the asymptotic derivations of the null distribution in Section (ref) and, in Section (ref), we assess the performance of the test in finite samples by comparing it with the test for threshold nonlinearity that assumes i.i.d. innovations Gor23. In Section (ref) we study the dynamics of Italian strikes time series. We apply our test to the monthly series of non-worked hours due to strikes and propose and validate the TARMA-GARCH specification. Finally, Section (ref) concludes the paper. The Supplementary Material contains all the proofs, additional Monte Carlo results and further analyses of the Italian strikes series.

Test for Threshold Effects with GARCH Innovations

Notation and preliminaries

Assume the time series $\{X_t:t=0,\pm 1,\pm2,\dots\}$ to follow the threshold autoregressive moving-average model with GARCH errors, TARMA$(p,q)$-GARCH$(u,v)$, defined by the system of difference equations:

align[align omitted — 413 chars of source]

The process $\{z_t\}$ is a sequence of i.i.d. random variables with zero mean, unit variance and finite fourth moment. $p$ and $q$ are the autoregressive and moving-average orders, respectively; $d$ is the delay parameter; $u$ and $v$ are the ARCH and GARCH orders, respectively. We assume $p,q,d,u,v$ to be known positive integers. Moreover, $I(\cdot)$ is the indicator function and $r\in\mathds{R}$ is the threshold parameter. For notational convenience, we abbreviate $I(X_t\leq r)$ by $I_r(X_t)$. We define the following vectors containing the parameters of Model ((ref)): $\boldsymbol{\phi} = \left(\phi_{0},\phi_{1},\dots,\phi_{p}\right)^\intercal\in\Theta_\phi,$ $\boldsymbol\varphi=\left(\varphi_{0},\varphi_{1},\dots,\varphi_{p}\right)^\intercal\in\Theta_\varphi,$ $\boldsymbol{\theta} = \left(\theta_{1},\dots,\theta_{q}\right)^\intercal\in\Theta_\theta,$ $\boldsymbol\vartheta=\left(\vartheta_{1},\dots,\vartheta_{q}\right)^\intercal\in\Theta_\vartheta,$ $\mathbf{a} = \left(a_0,a_1,\dots,a_u\right)^\intercal\in\Theta_a,$ $\mathbf{b} = \left(b_1,\dots,b_v\right)^\intercal\in\Theta_b,$ with $\Theta_\phi,\Theta_\varphi\subseteq\mathds{R}^{p+1}$; $\Theta_\theta,\Theta_\vartheta\subseteq\mathds{R}^{q}$; $\Theta_a\subseteq\mathds{R}^{u+1}$ and $\Theta_b\subseteq\mathds{R}^{v}$. Moreover let

align[align omitted — 546 chars of source]

$\boldsymbol{\Psi}_2$ and $\boldsymbol{\Psi}_1$ contain the ARMA parameters to be tested and those who are not, respectively; $\boldsymbol{\lambda}$ is the parametric vector for the ARMA-GARCH model whereas $\boldsymbol\eta$ is the vector of all the parameters (excluding the threshold $r$) in Model ((ref)). We use $\boldsymbol{\Psi}$, $\boldsymbol{\phi}$, etc.$\dots$ to refer to unknown parameters, whereas the true parameters are obtained by adding the $0$ subscript, i.e.: $\boldsymbol{\Psi}_0=(\boldsymbol{\Psi}_{0,1}^\intercal,\boldsymbol{\Psi}_{0,2}^\intercal)^\intercal$, $\boldsymbol{\lambda}_0^\intercal$ and $\boldsymbol\eta_0^\intercal$. Also, we assume $\boldsymbol\eta_0$ to be an interior point of the parameter space. We test whether a TARMA$(p,q)$-GARCH$(u,v)$ model provides a significantly better fit than the linear ARMA$(p,q)$-GARCH$(u,v)$ model by developing a Lagrange multiplier test statistics. Letting $\boldsymbol 0$ be the vector of all zeroes, the system of hypothesis results:

equation[equation omitted — 151 chars of source]

Under $H_0$ the process follows a linear ARMA$(p,q)$-GARCH$(u,v)$ model:

align[align omitted — 287 chars of source]

Define the polynomials $\phi(z)=1-\phi_{1}z-\phi_{2}z^2-\dots-\phi_{p}z^p$, $\theta(z)=1-\theta_{1}z-\theta_{2}z^2-\dots-\theta_{q}z^q,$ $a(z)=1-a_{1}z-a_{2}z^2-\dots-a_{u}z^u$, $b(z)=1-b_{1}z-b_{2}z^2-\dots-b_{v}z^v,$ $\varphi(z)=1-\varphi_{1}z-\varphi_{2}z^2-\dots-\varphi_{p}z^p,$ $\vartheta(z)=1-\vartheta_{1}z-\vartheta_{2}z^2-\dots-\vartheta_{q}z^q.$ We assume the following:

description\phantom{bla} \begin{itemize} • $\phi(z)\neq0$ and $\theta(z)\neq 0$ for all $z\in\mathds{C}$ such that $|z|\leq 1$ and they do not share common roots. $\varphi(z)$ and $\vartheta(z)$ are also coprime. • $a_i>0$, $i=0,1,\dots,u$; $b_j>0$, $j=1,\dots,v$; $|\sum_{i=1}^{u}a_i+\sum_{j=1}^{v}b_j|<1$; $a(z)$ and $b(z)$ are coprime. • $\{z_t\}$ is a sequence of i.i.d. random variables with $E[z_t]=0$, $E[z_t^2]=1$ and $E[z_t^4]<\infty$. Moreover, $z_t$ has a continuous and positive density function, say $f_z(x)$. • The sequence $\{\varepsilon_t\}$ is strictly stationary and ergodic with a finite fourth moments. • The process $\{X_t\}$ is ergodic and invertible under $H_0$. \end{itemize}

These assumptions are common in deriving the asymptotic behaviour of test for threshold nonlinearity. See, inter alia, Li08, Gor23, Cha90a. In particular, Assumptions A.1 and A.2 allow to estimate and identify uniquely the parameter of the ARMA and GARCH part, respectively. These assumptions also imply that the process $\{h_t\}$ is strictly stationary and ergodic with $E[h_t^2]<\infty$ and it is bounded away from zero with probability 1, see Li08 and Li11 for further details. Suppose we observe $X_1,\dots,X_n$. Omitting a negative constant, the Gaussian log-likelihood conditional on the initial values $X_0,X_{-1},\dots$ is:

align[align omitted — 699 chars of source]

Also:

align[align omitted — 429 chars of source]

and, under the null hypothesis, $\varepsilon_t(\boldsymbol{\eta}_0,r)=\varepsilon_t$ and $h_t(\boldsymbol{\eta}_0,r)=h_t$.

The derivation of the Lagrange multipliers test is based upon the first and second partial derivatives $\ell_n(\boldsymbol{\eta},r)$ with respect to $\boldsymbol{\Psi}$. Let

align*[align* omitted — 686 chars of source]

with $\partial\varepsilon_t(\boldsymbol{\Psi},r)/\partial\boldsymbol{\Psi}$ (respectively $\partial h_t(\boldsymbol{\Psi},r)/\partial\boldsymbol{\Psi}$) being the partial derivative of $\varepsilon_t(\boldsymbol{\Psi},r)$ ($h_t(\boldsymbol{\Psi},r)$) with respect to $\boldsymbol{\Psi}$:

align*[align* omitted — 744 chars of source]

Lastly, define the block matrix $\mathcal{I}_n(\boldsymbol{\eta},r)$:

equation[equation omitted — 896 chars of source]

The Lagrange multiplier approach requires estimating the model under the null hypothesis. Hence, let $\hat{\boldsymbol{\lambda}}=(\hat{\boldsymbol{\phi}}^\intercal, \hat{\boldsymbol{\theta}}^\intercal, \hat{\mathbf{a}}^\intercal, \hat{\mathbf{b}}^\intercal)^\intercal=\arg\min_{\boldsymbol{\lambda}}\ell_n(\boldsymbol{\lambda})$, with $\ell_n(\boldsymbol{\lambda})=\ell_n(\boldsymbol{\eta},-\infty)$. Hence, $\hat{\boldsymbol{\lambda}}$ is the Maximum Likelihood Estimator (hereafter MLE) of the ARMA-GARCH coefficients in Eq. ((ref)) and we define $\hat{\boldsymbol{\eta}}=(\hat{\boldsymbol{\lambda}}^\intercal,\boldsymbol 0^\intercal)^\intercal$ to be the so called restricted MLE, i.e. under the null hypothesis. We write $\partial \hat{\ell}_n(r)/\partial \boldsymbol{\Psi}_2$ and $\hat{\mathcal{I}}_{n}(r)$ to refer to $\partial \ell_n(\boldsymbol{\eta},r)/\partial \boldsymbol{\Psi}$ and $\mathcal{I}_n(\boldsymbol{\eta},r)$ evaluated at the restricted MLE $\hat{\boldsymbol{\eta}}$ , i.e.:

align*[align* omitted — 429 chars of source]

Under the null hypothesis, the threshold parameter $r$ is absent thereby the standard asymptotic theory is not applicable. To cope with this issue, we firstly develop the Lagrange multiplier test statistic as a function of $r$ ranging in a data-driven set $\mathcal{R}=[r_L,r_U]$, with $r_L$ and $r_U$ being, e.g., some percentiles of the data. Then, we take the overall test statistic as the supremum on $\mathcal{R}$. This approach has become widely used in the literature of tests involving nuisance parameters. This was first proposed in the seminal work of Cha90a, and followed by And93 and Han96. Within the threshold setting, the idea was deployed in Won97, Lin05, Li08, Li11, Cha20b, Gor21. Recently, Gia23 adapted it to prove the validity of a bootstrap scheme in testing threshold nonlinearity. The test statistic is

align[align omitted — 368 chars of source]

In Eq. ((ref)), besides taking the supremum, other convenient functions can be used to derive an overall test statistic.

The Null Distribution

In this section we derive the asymptotic distribution of $T_n$ under the null hypothesis that $\{X_t\}$ follows the ARMA$(p,q)$-GARCH$(u,v)$ process defined in Eq. ((ref)). Hereafter, all the expectations are taken under the true probability distribution for which $H_0$ holds. We use $\|\cdot\|$ to refer to the $\mathcal{L}^2$ matrix norm (the Frobenius' norm, i.e. $\|A\|=\sqrt{\sum_{i=1}^{n}\sum_{j=1}^{m}|a_{ij}|^2}$, where $A$ is a $n\times m$ matrix). Also, $o_p(1)$ indicates the convergence in probability to zero as $n$ increases. $\mathcal{D}_\mathds{R}(a,b)$, $a<b$, is the space of functions from $(a,b)$ to $\mathds{R}$ that are right continuous with left-hand limits. We assume $\mathcal{D}_\mathds{R}(a,b)$ to be equipped with the topology of uniform convergence on compact sets, see Bil68 for further details. In order to obtain its asymptotic distribution in the main theorem, we derive, under the null hypothesis, an (asymptotic) uniform approximation of the test statistic depending on the true parameters $\boldsymbol{\eta}_0=(\boldsymbol{\lambda}_0^\intercal,\boldsymbol 0^\intercal)^\intercal$. In this respect, define the vector $ \nabla_n(r)=\left(\nabla^\intercal_{n,1},\nabla^\intercal_{n,2}(r)\right)^\intercal$, with

align*[align* omitted — 658 chars of source]

and the matrix

align*[align* omitted — 545 chars of source]

In the following proposition, under the null hypothesis and Assumption A, we provide a uniform approximation that will allow us to derive the asymptotic null distribution of the test statistic $T_n$ in the main theorem.

propositionUnder Assumptions A.1--A.5 and under $H_0$, we have the following: \begin{description} • \begin{align*} \sup_{r\in[a,b]} \left\|\left(\frac{\hat{\mathcal{I}}_{n,22}(r)}{n} - \frac{\hat{\mathcal{I}}_{n,21}(r)}{n} \left(\frac{\hat{\mathcal{I}}_{n,11}}{n}\right)^{-1} \frac{\hat{\mathcal{I}}_{n,12}(r)}{n}\right)^{-1} -\left(\Lambda_{22}(r)- \Lambda_{21}(r)\Lambda_{11}^{-1}\Lambda_{12}(r)\right)^{-1}\right\| = o_p(1). \end{align*} • $$\sup_{r\in[a,b]}\left\|\frac{1}{\sqrt{n}}\frac{\partial\hat\ell_n(r)}{ \partial\boldsymbol\Psi_2}- \left(\nabla_{n,2}(r)-\Lambda_{21}(r)\Lambda_{11}^{-1}\nabla_{n,1}\right)\right\| =o_p(1).$$ \end{description}

Define the process: $\{Q(r),r\in\mathds{R}\}$, with $Q(r)=\left(\nabla_{n,2}(r)-\Lambda_{21}(r)\Lambda_{11}^{-1}\nabla_{n,1}\right)$. Note that $\{Q(r)\}$ is a marked empirical process with infinitely many markers. In the next theorem we derive a novel Functional Central Limit Theorem (hereafter FCLT) for $\{Q(r)\}$.

theoremLet $\left\{\xi(r),\;r\in\mathds{R}\right\}$ be a centered Gaussian vector process of dimension $(p+q+1)$, with covariance kernel $\Sigma(r,s) = \Lambda_{22}(r\wedge s)-\Lambda_{21}(r)\Lambda_{11}^{-1}\Lambda_{12}(s).$ Under Assumptions A.1--A.5 and $H_0$, $Q(r)$ converges weakly to $\xi(r)$ in $D_{p+q+1}(-\infty,+\infty)$.

By invoking the continuous mapping theorem, the asymptotic null distribution of our Lagrange multiplier test statistic readily follows.

theoremUnder Assumptions A.1--A.5 and $H_0$, asymptotically, the Lagrange multiplier test statistic $T_n$ has the same distribution of \begin{equation} \sup_{r\in[r_L,r_U]}\xi(r)^\intercal\Sigma(r,r)^{-1}\xi(r), \end{equation} where $\xi(r)$ and $\Sigma(r,s)$ are defined in Theorem (ref).

Finite Sample Performance

In this section, we study the finite sample performance of our test, which we denote by sLMg, and compare it with the sLM test of Gor23, which assumes i.i.d. innovations. The length of the series is $n=100,200,500$ and $z_t$, $t=1,\dots,n$ is generated from a standard Gaussian white noise. The nominal size of the tests is $\alpha=5\%$ and the number of Monte Carlo replications is 10000. Furthermore, we use the tabulated critical values of And03 and the threshold is searched from percentile 25th to 75th of the sample distribution.

Size

We study the empirical size by simulating from the following ARMA$(1,1)$-GARCH$(1,1)$ model:

align[align omitted — 197 chars of source]

where $\phi_1=(-0.9,-0.6,-0.3,0.0,0.3,0.6,0.9)$ and $\theta_1=(-0.8,-0.4,0.0,0.4,0.8)$. We combine these with the following parameters for the GARCH specification: $(a_0,a_1,b_1) = (1,0.04,0.95)$ (case A), $(1,0.3,0.0)$ (case B), $(1,0.4,0.4)$ (case C), so as to obtain 35 different parameter settings for each case. Notice that case B corresponds to an ARCH(1) process. Also, all three cases fulfil the condition of finite fourth moments. The results are presented in Figure (ref), where the boxplots group together the 35 configurations for each of the three cases. The full results are reported in Tables (ref)--(ref) of the Supplementary Material. Clearly, the standard sLM test is oversized and, especially for case C, the bias increases with the sample size and can be severe. On the contrary, our TARMA-GARCH test has correct size in almost every setting for a sample size as low as $n=100$. Note that this holds also for corner cases, such as near integrated cases or when near cancellation of the AR and MA polynomials occurs. As the sample size increases, the empirical size tends to concentrate around the nominal $5\%$ for case B, whereas it seems to settle around lower values for cases A and C. As we will show in the next section, this slight undersize does not impinge negatively upon the power and results in a conservative test, which is generally appealing in practical applications. The size for the three cases grouped together is reported in Figure (ref) of the Supplementary Material.

figure[figure omitted — 410 chars of source]

Power

In order to study the power of the tests we simulate from the following TARMA$(1,1)$-GARCH$(1,1)$ model:

align[align omitted — 288 chars of source]

where $\varphi_{0}=\varphi_1=\vartheta_1 = 0.5 + \Psi$, where $\Psi = (0.00, -0.15, -0.30, -0.45, -0.60, -0.75, -0.90, -1.05).$ Hence, the system of hypotheses of Eq. ((ref)) becomes

$$

casesH_0&:\boldsymbol{\Psi}= \mathbf{0} \\ H_1&:\boldsymbol{\Psi}\neq\mathbf{0},

$$ \noindent where $\boldsymbol{\Psi} = (\Psi,\Psi,\Psi)$, so that the parameter $\Psi$ represents the departure from the null hypothesis. As above, we combine these with the following parameters for the GARCH specification: $(a_0,a_1,b_1) = (1,0.1,0.8)$ (case A), $(1,0.4,0.4)$ (case B), $(1,0.8,0.1)$ (case C). The size-corrected power of the tests (in percentage) is presented in Figure (ref), where each panel corresponds to a different sample size. The blue and orange lines correspond to the sLMg and sLM test, respectively, whereas the different points correspond to cases A (filled round dot), B (empty round dot) and C (filled triangle). The behaviour of the two tests is very similar, and the power loss incurred by estimating the GARCH parameters in the sLMg test is very small and is overwhelmingly compensated by the correct size in presence of heteroskedasticity. The full results are reported in the Supplementary Material, Table (ref) (size corrected power) and Table (ref) (raw power). The general conclusion that can be drawn is that, even in absence of information regarding the presence of heteroskedasticity, it is safe to use the sLMg test as it will lower the risk of a false rejection while retaining a good discriminating power.

figure[figure omitted — 407 chars of source]

Measurement Error

We assess the effect of measurement error on the size of the tests. We simulate $X_t$ from the following AR(1)-GARCH(1,1) model

align[align omitted — 167 chars of source]

where, as above, $\phi_1=(-0.9,-0.6,-0.3,0.0,0.3,0.6,0.9)$ and $(a_0,a_1,b_1) = (1,0.04,0.95)$ (case A), $(1,0.3,0.0)$ (case B), $(1,0.4,0.4)$ (case C). We add measurement noise as follows: $Y_t = X_t + \eta_t$, where the measurement error $\eta_t\sim N(0,\sigma^2_\eta)$ is such that the signal to noise ratio SNR $=\sigma^2_X/\sigma^2_\eta$ is equal to $\{\infty, 50,10,5\}$. Here, $\sigma^2_X$ is the unconditional variance of $X_t$ computed by means of simulation. The case without noise (SNR $=\infty$) is taken as the benchmark. The empirical size (rejection percentages) for the three sample sizes is presented in Table (ref). Clearly, the size of the sLMg test is either minimally affected or not affected at all by the presence of measurement error, even for high levels of noise (signal to noise ratio = 5).

table[table omitted — 2,487 chars of source]

As for the power of the test, the presence of high levels of measurement noise impinge negatively upon it so that larger sample sizes are required to compensate for this. The study is reported in Section (ref) of the Supplementary Material.

Testing and Modelling the Italian Strikes Time Series

figure[figure omitted — 313 chars of source]

In this section, we study the dynamics of the monthly series of hours not worked due to strikes in Italy between January 1949 and December 2009 ($n=732$). The data were obtained from the Italian Statistical Institute (ISTAT). Despite their importance, to the best of our knowledge, this is the first time that these data are used within a macroeconometric approach. Note that, also for Italy, ISTAT has suspended the collection of monthly data from 2009 onwards. The time plot is characterized by cyclical oscillations and a structural change in variance starting from the mid-1980s, see Figure (ref) (left). In the Italian history, the amplitude of such oscillation seems to reduce starting from the 1980s, without getting back to the previous variability in successive years. This can be explained by a series of historical occurrences such as the increase in the workers' opportunity cost associated to the decision to strike, which reduced the incentive to strike for economic reasons, e.g. wage and/or contractual claims. We consider the logarithm (in base 10) as a variance-stabilising transformation and show the result in Figure (ref) (right). The month plot is shown in Figure (ref) of the Supplementary Material. The series has seasonal oscillations and August is the month where the strikes undergo a consistent drop, followed by September, January, and December, where the phenomenon is less pronounced. August is a vacation month in Italy, and September is a post-vacation month, while January and December are characterized by less working days due to Christmas holidays. The series also presents some seasonality in the minima and the maxima, due to the strikes being more likely to be observed in those months when they are more effective. As strikes are costly for both firms and workers, the seasonality of economic activity, as well as the economic cycle, make the strike strategy more effective during the high seasons or expansionary phases while discouraging workers during low seasons or recessionary phases. The autocorrelation of the series decays slowly hinting at the presence of a trend and a strong and seemingly nearly integrated seasonal component.\footnote{The correlograms are reported in Supplementary Material, Section (ref).} The spectral density function (smoothed periodogram with a Daniell window) of Figure (ref) shows the main periodicities of the series. Besides the yearly periodicity, the 7-year peak together with its harmonics (3.5, 1.7, and 0.85 years) are related to the business cycle. The 2.1-year periodicity could be a first indication of a non-linear dynamics emerging as a resonance (non-trivial combination of the main frequencies). This periodicity is also in line with the existence of a cycle in Italian strikes recalled by franzosi1980 based on the spectral analysis of this same series, using data predating the 1980s. Indeed, strike activity concentrates at the expiration of the contracts, whose average duration is, 2-3 years myers2016.

figure[figure omitted — 320 chars of source]

Labour strikes are a tool for bargaining better working conditions, both for those who strike and for those who do not, and this poses a threat to firms barrett1989. Strikes have an upper bound in the number of hours they can be used, though, since above that level the tool can backfire and damage the workers themselves. This can be due, for example, to the firm being so damaged by the strikes that it has to reduce its production and hence lay off some workers. This implies that the series cannot present a unit root, be it regular or seasonal. Moreover, cointegrating relations are not expected. We assess this conjecture by applying a battery of unit root tests to the series of the Italian strikes and to the following covariates: salary: Monthly index of salaries of industrial workers; price: Monthly index of consumer prices\footnote{All the series have been obtained from the Italian National Institute of Statistics (ISTAT).} Following the labor supply models, we have chosen these two economic variables because they represent a measure of the opportunity cost of strikes and one of the main economic motivations behind strikes for wage demands. The results are shown in Table (ref). The first row presents the test $\bar MZ_{\alpha}^{\text{GLS}}$ as proposed by Per07, which is essentially the same as the $MZ_{\alpha}^{\text{GLS}}$ test of Ng01 but the lag of the ADF regression is selected on OLS detrended data. The second row shows the results for the $\bar MP_{t}^{\text{GLS}}$, the modified feasible point optimal test (see also Eq. (9) in Ng01). The third row lists the GLS detrended version of the ADF test (denoted by ADF$^{\text{GLS}}$). Following Cha20b, we have chosen them since they are the best performers among all those proposed in Ng01 and Per07.\footnote{We have ported to R the original Gauss routines of Ng01, which are available at \url{https://drive.google.com/open?id=0B-aG4lrQrBYsazN1RktHX2dfZkU}. The results for the remaining tests are similar and are available upon request.} The last column contains the critical values of the null asymptotic distribution at the 5% level. Clearly, none of the three tests manages to reject the null hypothesis of a unit-root in any of the series.

table[table omitted — 606 chars of source]

For this reason, we perform a pairwise cointegration analysis. We regress the series of strikes on the three covariates and apply the above tests to the residuals of the fitted models. The results are presented in Table (ref). Again, none of the tests is able to reject the null hypothesis and this renders the whole analysis inconclusive in that all the series appear to be integrated but no cointegrating relationships can be found.

table[table omitted — 579 chars of source]

One possible reason for the above results is the lack of power of unit root tests against a non-linear alternative. This could be due to them not explicitly encompassing the nonlinearity within their specification. On the other hand, tests for a unit root against a threshold autoregressive alternative suffer from the presence of MA components and are severely biased, leading to over-rejecting. This is discussed in Cha20b,Cha24, which solves the problem by proposing a unit root test where the null hypothesis entails an integrated MA against the alternative of a stationary threshold ARMA model, possessing a unit root regime. The application of such a test to the time series of Italian strikes produces a test statistic equal to 13.95, with an associated (heteroskedastic robust) wild bootstrap $p$-value of 0.012, and this points to an alternative explanation of the strikes dynamics based upon a regime-switching mechanism.

In Section (ref) of the Supplementary Material, we adopt a linear modelling approach based on a two-step ARIMA-GARCH fit. The results show that both regular and seasonal differences are needed and this is not consistent with stylized economic facts and difficult to interpret. Moreover, the residual analysis hints at the presence of unaccounted non-linear dependence (see Figure (ref) of the Supplementary Material). For these reasons, we specify a TARMA-GARCH model and test it against the ARMA-GARCH specification by using our novel sup-Lagrange Multiplier test to take into account the conditional heteroskedasticity. We use the consistent Hannan-Rissanen criterion to select the order of the ARMA model to be tested Han82. The maximum order tested is the ARMA(12,12) and the procedure used selects the ARMA(1,1) model. As for the order of the GARCH model specification we select the GARCH(1,2) on the basis of previous investigations. Hence we test the ARMA(1,1)-GARCH(1,2) against the TARMA(1,1)-GARCH(1,2) specification with $d=1$. We obtained a value of the supLM statistic equal to 25.668 (threshold = 6.273). Given that the critical value at 1% level results 17.65 (from Table I of And03 with $\pi_0=.20$), our asymptotic sLMg test rejects and points to a significant threshold effect either in the intercept and/or at lag 1, so the results of the test corroborate the hypothesis that the dynamic of strikes is governed by a regime-switching dynamics with conditional heteroskedasticity. We propose the following two-stage TARMA-GARCH model:

align[align omitted — 505 chars of source]

We adopt a two-stage estimation approach since the existing results for TARMA models include least squares estimators for which consistency and asymptotic normality have been established Gia21, Li11b. In this way, we can adopt a maximum likelihood approach for estimating the GARCH part and exploit mature and reliable existing implementations such as those of the R package rugarch Gha20. Equation (ref) reports the estimated coefficients of the model, with standard errors in parentheses below each coefficient. The threshold is estimated to be 6.23, equivalent to roughly 1.7 million non-worked hours due to strikes in a month.

align[align omitted — 737 chars of source]

The estimated TARMA model fulfils the sufficient conditions of ergodicity and invertibility as discussed in Cha19. Figure (ref) shows the series (light blue line) together with the fitted values from the TARMA model (blue line), the estimated conditional standard deviation $\hat h^{1/2}_t$ from the GARCH(1,1) fit (green line), and the estimated threshold ($\hat r = 6.23$, dashed red line). The TARMA specification for the conditional mean has a lower regime which is characterized by a persistent seasonality with a clear threshold effect at lag 12. The MA(3) parameters are not significant but play a role in providing a solution that mimics the observed dynamics. Note that the upper regime contains no seasonal MA lags and this is a strong indication of the presence of a threshold effect in the moving average model, something that conventional TAR models are not able to reproduce or, at best, can only reproduce partially by incorporating higher order lags. Clearly, the dynamics reacts differently to shocks depending upon the regime. The threshold identifies the time corresponding to the structural break in the variance of the series and splits it into two periods. Indeed, from the mid-eighties onwards the series belongs to the lower regime where the dynamics is much more characterized by seasonality. The series of the conditional standard deviation shows both clusters of volatility and asymmetric peaks, and appears to change starting from late 1990s. The parameter estimates for the GARCH(1,1) part of the model are shown in Equation (ref). The estimates indicate a non-negligible heteroskedastic structure in the innovations. Also in this case, the estimated model results stationary and ergodic and its adequacy is assessed in the Supplementary materials (ref), which also reports model diagnostics.

figure[figure omitted — 342 chars of source]

The fit is able to reproduce the strong asymmetric behaviour of the series: the histogram of the data in reported in Figure (ref) (left), where we superimposed a kernel density estimate over a simulated trajectory of 100k observations from the fitted model (blue line). This is also witnessed by the Kolmogorov-Smirnov two-sample test, computed between the distribution of the observed data and of that of the simulated series ($p$-value = 0.08). Moreover, we test the residuals of the model for the presence of non-linear serial dependence, up to lag 12, with the entropy metric $S_{\rho}$, see Figure (ref) (right). Since no lags exceed the bootstrap rejection band we can be confident that the model managed to capture the underlying non-linear features of the series. Finally, in Figure (ref) of the Supplementary Material we show that there is no unaccounted cross-dependence between the residuals of the fitted model and the covariates salary and price, introduced at the beginning of the section.

figure[figure omitted — 609 chars of source]

The existence of a threshold-type dynamics is consistent with the underlying theory, see e.g. granovetter1978, where individuals make choices that are influenced by the behaviour of other individuals. Also, bursztyn2021 stress the importance of the activity of people within their social network when choosing to take part in a protest. The existence of strong ties between individuals is important for the persistence of the protests and the clustering effect. As for the structural change in the variance, we note that the strengthening of the government and the exclusion of unions from the process of reforming policies has downplayed collective strikes as a political and social bargaining tool hamann2013. godard2011 argues that the change in the dynamics of strikes observed since the 1980s can be due to four different reasons: $i)$ they have been diverted into alternative forms of conflict; $ii)$ the capitalist system managed to reduce the number of citizens who disapprove of the economic and political system itself, or at least it managed to reduce their will to act upon it; $iii)$ the conflict underwent a transformation and has become more deeply embedded and linked to general behaviour within and outside the workplace, such as cynicism and escapism, but also depression; $iv)$ conflict has become dormant. While some of these reasons appear more convincing than others, their multifactorial interaction could have played a role in explaining the observable shift. In addition, other possible reasons, partially overlapping the aforementioned, are the reduction of the union density, the birth of autonomous and non-political trade unions, and the creation of local agreements, three events that characterized Italy after the 1970s and that are particularly relevant in the 1990s and 2000s, see also giangrande2021.

Conclusions

Our proposal, that compares the ARMA-GARCH versus the TARMA-GARCH specifications, allows us to test for nonlinearity in the conditional mean, without being affected by heteroskedasticity, making it possible to assess the two types of nonlinearity separately. Since threshold ARMA models can accommodate a wide range of complex dynamics, we expect the test to have power against most types of nonlinearity in the conditional mean, besides threshold models. For this reason, it can be used as an omnibus explorative nonlinearity test with good properties of robustness against heteroskedasticity. This aspect will be further explored in future investigations. The results point to a significant threshold effect in the series of Italian strikes, which appears to be governed by a regime-switching dynamics with conditional heteroskedasticity. The analysis of the strikes offers some evidence that TARMA models provide a flexible and parsimonious specification that describes some of its empirical non-linear stylized facts. Our findings are coherent with both features of the history of the Italian labour market, such as the decrease in the social turmoil after the first half of the 1980s godard2011, and the social aspects of unrest and labour markets in general, based on the threshold model by granovetter1978 and also in line with more recent studies of the impact of social ties and the existence of social networks able to trigger political engagement and participation to protest bursztyn2021.