Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
67,681 characters · 3 sections · 124 citation commands
COVID-19: Tail Risk and Predictive Regressions
Keywords: COVID-19, pandemic, tail risk, predictive regressions, forecasting, robust inference JEL Classification: C13, C51
\@ifstar{\starsection}{\nostarsection}{Introduction} Several recent papers have focused on econometric and statistical analysis and forecasting of key time series and variables associated with the on-going COVID-19 pandemics, including infection and death rates, and their effects on economic and financial markets (see Stock1, Stock1, Toda1, Toda1, Miles, Miles, Villaverde, Villaverde, Paul, Paul, Pesaran, Pesaran, Stock, Stock, Toda, Toda, Linton, Linton, Manski, Manski, among others).
This paper contributes to the above literature by focusing on the robust analysis of the effects of the pandemics on financial markets across the World. We provide the results of robust estimation and inference on predictive regressions for returns on major stock indexes in 23 developed and emerging economies in North and South America, Europe, and Asia incorporating the time series of reported infections and deaths from COVID-19.
We also present a detailed study of persistence, heavy-tailedness and tail risk properties of COVID-19 infections and deaths time series that emphasize the necessity in applications of robust inference methods in the analysis and forecasting of the COVID-19 pandemic and its impact on economic and financial markets and the society.
Econometrically justified and robust analysis in the paper is based on heteroskedasticity and autocorrelation consistent (HAC) inference methods, recently developed robust $t$-statistic inference procedures and robust tail index estimation approaches.
The results of the analysis, in particular, point to potential non-stationarity in the form of unit roots in the time series of daily infections and deaths from COVID-19 that are commonly used in research on modelling and forecasting of the COVID-19 pandemic and its effects. The results emphasize the necessity in basing the analysis of models incorporating the COVID-19-related time series such as daily infection and death rates on their (stationary) differences. The analysis using robust tail index inference methods further indicates potential heavy-tailedness with possibly infinite variances and first moments in the time series of daily infections and deaths from COVID-19 and their differences in countries across the World.
In order to account for the problems of potential non-stationarity in the daily COVID-19 infections and deaths time series, the paper provides the analysis of predictive regressions for financial returns incorporating both the lagged daily infections/deaths from COVID-19 and their differences. Further, the properties of autocorrelation, heavy-tailedness and heterogeneity in the COVID-19 infections/deaths time series are accounted for by the use in the predictive regression analysis of both the widely applied standard HAC inference methods as well as the recently proposed $t-$statistic approaches to robust inference under the above problems in the data (see IM, IM, IM1, Section 3.3 in ibragimov2015heavy, ibragimov2015heavy and Section (ref) in this paper).
The standard HAC inference methods indicate (apparently spurious) statistical significance of the (potentially non-stationary) lagged daily infections and deaths from COVID-19 in predictive regressions for returns on the major stock indices in some countries. HAC methods also point to statistical significance of (stationary) lagged daily changes in the number of COVID-19 infections and deaths for some countries. For daily changes in COVID-19 infections, high statistical significance, with the expected negative signs of the predictive regression coefficients, is observed in the case of Italy, India, Brazil and Argentina.
Motivated by the results on high persistence and heavy-tailedness in daily COVID-19 infections/ deaths time series obtained in the paper and also by poor finite sample properties of HAC inference methods (see Section (ref) and references therein), we further provide the assessment of statistical significance of the coefficients in the predictive regressions using $t-$statistic approaches to robust inference based on group estimates. According to the statistically justified analysis using the robust $t-$statistic approaches, the lagged daily COVID-19 infection and death rates and their (stationary) differences appear not to be statistically significant in predictive regressions for stock index returns in all the countries considered in the analysis.
Overall, one of the main conclusions from the results in the paper is that statistical and econometric analyses and forecasts of the on-going COVID-19 pandemic and its impacts on economic and financial markets and the society should be based on theoretically justified robust inference methods. The methods used in the analysis and the forecasting of the pandemic and its effects should account, in particular, for the problems of potential non-stationarity, autocorrelation, heavy-tailedness and heterogeneity in the key time series and variables related to COVID-19, including the infections and deaths time series.
The paper is organised as follows. Section (ref) provides the results of (non-)stationarity and unit root tests for the time series of COVID-19 infection and death rates in the countries across the World. Section (ref) describes the data used in the analysis. Section (ref) presents the analysis of heavy-tailedness and tail risk properties of daily COVID-19 infection and death rates. Section (ref) provides the results of theoretically justified and robust statistical analysis of predictive regressions for the returns on major stock indices in the countries considered incorporating the time series on COVID-19 infections and deaths. Section (ref) makes some concluding remarks and discusses directions for further research. Appendices A and B provide the diagrams and tables on the results of the statistical analysis in the paper.
\@ifstar{\starsection}{\nostarsection}{Data}
The analysis in the paper uses the data on daily COVID-19 infections and deaths in different countries across the World (the UK, Germany, France, Italy, Spain, Russia, the Netherlands, Sweden, India, Austria, Finland, Ireland, the US, Lithuania, Canada, Brazil, Mexico, Argentina, Japan, China, South Korea, Indonesia and Australia) for the period from 22 January 2020 to 22 March 2021. The data is obtained from the Data Repository maintained by the Center for Systems Science and Engineering (CSSE) at Johns Hopkins University.\footnote{https://github.com/CSSEGISandData/COVID-19} The data on daily prices of major stock indices for the countries considered is obtained from Yahoo Finance and the data on interest rates is from the Global Rates database.\footnote{https://www.global-rates.com/interest-rates/central-banks/central-banks.aspx} We consider the following stock indices: FTSE 100 (UK), DAX (Germany), CAC 40 (France), FTSE MIB (Italy), IBEX 35 (Spain), MOEX (Russia), AEX (Netherlands), OMXS 30 (Sweden), SENSEX (India), ATX (Austria), OMX Helsinki 25 (Finland), ISEQ (Ireland), Dow Jones, S&P 500 (USA), OMX Vilnius (Lithuania), TSX (Canada), iBovespa (Brazil), IPC Mexico (Mexico), Merval (Argentina), NIKKEI 225 (Japan), SHANGHAI (China), KOSPI (South Korea), JCI (Indonesia), ASX 50, ASX 200 and Australian All Ordinaries (Australia). The analysis uses central bank rates for the countries considered; European interest rate is used for the country members of European monetary union.
Throughout the paper, $I_t$ and $D_t$ denote the (cumulative) number of COVID-19 infections and deaths from the beginning of the period on 22 January 2020 to day $t$ in the countries considered. Further, $\Delta I_t$ and $\Delta D_t$ denote the differences of the above time series, that is the number of reported infections and deaths in day $t.$ By $\Delta^2 I_t$ and $\Delta^2 D_t$ we denote the cumulative infections/deaths time series' second differences, that is, the daily changes in the number of COVID-19 infections and deaths in the countries dealt with. The estimation and testing in the paper is based on the periods with positive values of the number of COVID-19 infections and deaths $I_t$ and $D_t$ in the countries considered.
\@ifstar{\starsection}{\nostarsection}{Empirical results}
We begin the analysis by the study of the degree of integration in the time series $I_t,$ $D_t,$ $\Delta I_t,$ $\Delta D_t,$ $\Delta^2 I_t$ and $\Delta^2 D_t$ of COVID-19 infections and deaths and their differences in the countries considered. Tables (ref) and (ref) present the results of several unit root tests for the time series of daily infections and deaths $\Delta D_t, \Delta D_t$ and the time series of daily changes in their number $\Delta^2 I_t, \Delta^2 D_t$ . The results are provided for the (right tailed) likelihood ratio unit root test proposed by JanssonNielsen2012 (with the test statistic $LR$ in Tables (ref) and (ref); see also skrobotov2018bootstrap, skrobotov2018bootstrap), the GLS-based modified Phillips-Perron type tests (with the corresponding test statistics $MZ_\alpha$, $MSB$, $MZ_t$), the modified point optimal test (with the test statistic $MP_t;$ see NP2001, NP2001) and the GLS-based Augmented Dickey-Fuller test (with the test-statistic denoted by $ADF$ in Tables (ref) and (ref); see ERS1996, ERS1996). To address the issue of possible heavy tails and infinite variance of the series (see the next section), for calculation of the $p-$value of the unit root tests, we use recently justified sieve wild bootstrap algorithm with a Rademacher distribution employed in the wild bootstrap re-sampling scheme (see cavaliere2016unit, cavaliere2016unit).
An important tuning parameter in the above tests is related to the choice of lag length used in the analysis. We use the modified Akaike information criterion (MAIC) lag choice approach based on standard ADF regressions as suggested by PQ2007. According to the results (the wild bootstrap $p-$values are given in brackets), the unit root hypothesis in the time series $\Delta I_t$ of daily COVID-19 infections is not rejected at reasonable significance levels, e.g., 5% and 10%, by all the employed tests for all the countries considered except Spain, Sweden, Ireland, China and South Korea. For the time series $\Delta D_t$ of daily COVID-19 related deaths, the unit root hypothesis in is not rejected for all the countries considered except Spain, Sweden, Finland, Argentina and China. For the daily infections and deaths time series $\Delta I_t$ and $\Delta D_t$ in China, the rejection of the unit root hypothesis is on every reasonable significance level (even at 1%).
On the other hand, according to the results in Tables (ref) and (ref), the unit root hypothesis is rejected at all reasonable significance levels by all the tests for the time series $\Delta^2 I_t$ and $\Delta^2 D_t$ of daily changes in the number of COVID-19 infections and deaths in all the countries considered.
The above results of unit root tests have several important implications for statistical analysis of models and key time series related to the COVID-19 pandemic and its effects. According to the results, in the countries across the World, the time series of daily COVID-19 infections and deaths and thus the time series of total (cumulative) infections/deaths from the disease up to a certain date that are typically employed in the analysis and forecasting of the pandemic and its impact appear to exhibit non-stationarity. The daily COVID-19 infections/deaths time series $\Delta I_t$ and $\Delta D_t$ appears to exhibit unit root process persistence for most of the countries considered. This, in turn, implies very high persistence in the time series $I_t$ and $D_t$ of total infections/deaths up to a certain date that appear to be integrated of order 2.\footnote{The conclusions on persistence properties of the time series $I_t, D_t$ and $\Delta I_t, \Delta D_t$ are somewhat similar to those for the CPI and the inflation rate (the change in the logarithm of the CPI) time series, where often unit root hypothesis is not rejected for the inflation rate and thus the (logarithm) of the CPI levels appears to be integrated of order 2 (see the analysis of non-stationarity in Section 14.6 in SW, SW, for the inflation rate and its changes in the US). These conclusions imply the necessity of the use of differences of the inflation rate in time series modeling of inflation and its relationship to other key economic variables such as the unemployment level in the Phillips curve (see Chs. 14 and 16 in SW, SW).}
Non-stationarity of daily COVID-19 infections and deaths time series implies that, due to the spurious regression problem, the statistical analysis of models incorporating these and other nonstationary variables related to the pandemic should be based on their stationary differences as in the case of predictive regressions for financial returns in Section (ref).
In this section, we provide the analysis of heavy-tailedness and tail risk properties of daily COVID-19 infection and death rates in the countries considered. The estimates point to pronounced heavy-tailedness in the infections/deaths time series. This further motivates the necessity in applications of robust methods in modelling and forecasting the dynamics of infection and death rates and other variables related to the pandemic and inference on their effects on the world financial and economic markets.
As indicated in many empirical and theoretical works in the literature (see, among others, the analysis and the reviews in embrechts1997modelling, embrechts1997modelling, cont2001empirical, cont2001empirical, gabaix2009power, gabaix2009power, ibragimov2015heavy, ibragimov2015heavy, and references therein), distributions of many variables related to or affected by crises and natural disasters and characterised by the presence of extreme values and outliers, such as financial returns, catastrophe risks or economic losses from natural catastrophes, exhibit deviations from Gaussianity in the form of heavy power law tails. For a positive heavy-tailed variable (e.g., representing a risk, the absolute value of a financial return, or a loss from a natural disaster $X$) one has
for large $x>0,$ with a constant $C>0$ and the parameter $\zeta>0$ that is referred to as the tail index (or the tail exponent) of $X.$ The value of the tail index parameter $\zeta$ is important as it characterises the probability mass (heaviness and the rate of decay) in the tails of power law distribution ((ref)). Heavy-tailedness (i.e., the tail index $\zeta$) of the variable $X$ governs the likelihood of observing extremes and outliers in the variables. The smaller values of the tail index $\zeta$ correspond to a higher degree of heavy-tailedness in $X$ and, thus, to a higher likelihood of observing extremely large values in realisations of the variable. In addition, importantly, the value of the tail index $\zeta$ governs finiteness of moments of $X,$ with the moment $EX^p$ of order $p>0$ of the variable being finite: $EX^p<\infty$ if and only if $\zeta>p.$ In particular, the variance of $X$ is defined and is finite if and only if $\zeta>2,$ and the first moment of the variable is finite if and only if $\zeta>1.$
The degree of heavy-tailedness and finiteness of variances for variables is crucial for applicability of standard statistical and econometric approaches, including regression and least squares methods. Similarly, the problem of potentially infinite fourth moments of (economic and financial) time series dealt with needs to be taken into account in applications of autocorrelation-based methods and related inference procedures in their analysis (see the discussion in cont2001empirical, cont2001empirical, Ch. 1 in ibragimov2015heavy, ibragimov2015heavy, and references therein).
Many recent studies argue that the tail indices $\zeta$ in heavy-tailed models ((ref)) typically lie in the interval $\zeta\in (2, 4)$ implying finite variances and infinite fourth moments for financial returns in developed economies and maybe smaller than 2 implying possibly infinite variances for financial returns in emerging markets (see, among others, LP, LP, gabaix2009power, gabaix2009power, ibragimov2015heavy, ibragimov2015heavy, and references therein).\footnote{Heavy-tailed power law behavior is also exhibited by many other such important economic and financial variables as income and wealth (with $\zeta\in (1.5, 3)$ and $\zeta\approx 3,$ respectively; see, among others, gabaix2009power, gabaix2009power, and the references therein); financial returns from technological innovations, losses from operational risks and those from earthquakes and other natural disasters (with tail indices that can be considerably less than one, see ibragimov2015heavy, ibragimov2015heavy, and references therein).}
The recent study by Taleb provides (Hill's, see below) tail index estimates supporting extreme heavy-tailedness with $\zeta$ smaller than 1 and infinite first moments in the number of deaths from 72 major epidemic and pandemic diseases from 429 BC until the present. Toda1 report (Hill's) estimates of the tail index close to 1 implying infinite variances and first moments in the distribution of COVID-19 infections across the US counties at the beginning of the pandemic.
Several approaches to the inference about the tail index $\zeta$ of heavy-tailed distributions are available in the literature (see, among others, the reviews in embrechts1997modelling, embrechts1997modelling, gabaix2011rank, gabaix2011rank, Ch. 3 in ibragimov2015heavy, ibragimov2015heavy, and references therein). The two most commonly used ones are Hill’s estimates and the OLS approach using the log-log rank-size regression.
It was reported in a number of studies that inference on the tail index using widely applied Hill’s estimates suffers from several problems, including sensitivity to dependence and small sample sizes (see, among others, Ch. 6 in embrechts1997modelling, embrechts1997modelling). Motivated by these problems, several studies have focused on alternative approaches to the tail index estimation. For instance, huisman2001tail propose a weighted analogue of Hill’s estimator that is reported to correct its small sample bias for sample sizes less than 1,000. Using extreme value theory, MullerWang focus on inference on the quantiles and tail probabilities of heavy-tailed variables with a fixed number $k$ of their extreme observations (order statistics) employed in estimation as is typical in relatively small samples of fat-tailed data. embrechts1997modelling, among others, advocate sophisticated nonlinear procedures for tail index estimation.
gabaix2011rank focus on econometrically justified inference on the tail index $\zeta$ in heavy-tailed power law models ((ref)) using the popular and widely applied approach based on log-log rank-size regressions $\log(Rank)=a-b \log(Size),$ with $b$ taken as an estimate of $\zeta.$ The reason for popularity of the approach is its simplicity and robustness. gabaix2011rank provide a simple remedy for the inherent small sample bias in log-log rank-size approaches to inference on tail indices, and propose using the (optimal) shifts of 1/2 in ranks, with the tail index estimated by the parameter $b$ in (small sample bias-corrected) regressions $\log(Rank-1/2)=a-b \log(Size).$ gabaix2011rank further derive the correct standard errors on the tail exponent $\zeta$ in the log-log rank-size regression approaches. The standard error on $\zeta$ in the above log-log rank-size regressions is not the OLS standard error but is asymptotically $(2/k)^{1/2}\zeta,$ where $k$ is the number of extreme (the largest) observations on the heavy-tailed variable $X$ used in tail index estimation (see also Ch. 3 in ibragimov2015heavy, ibragimov2015heavy). The numerical results in gabaix2011rank point to advantages of the proposed approaches to inference on tail indices, including their robustness to dependence in the data and deviations from exact power laws in the form of slowly varying functions.
Naturally, it is important that inference on the tail index in heavy-tailed power law distributions ((ref)) is based on i.i.d. or stationary observations (op. cit.). The results in the previous section point to potential non-stationarity of the time series $\Delta I_t, \Delta D_t$ of daily COVID-19 infections and deaths and stationarity of their differences $\Delta^2 I_t, \Delta^2 D_t$ (the daily changes in the number of daily COVID-19 infections and deaths) in most of the countries considered. We, therefore, focus on estimation of the tail indices $\zeta$ in heavy-tailed power law models for the differences $\Delta^2 I_t$ and $\Delta^2 D_t.$ The implied tail indices for daily COVID-19 infections and deaths $\Delta I_t, \Delta D_t$ equal to the same values as in the case of $\Delta^2 I_t$ and $\Delta^2 D_t,$ as the former time series are cumulations of the latter ones.
Figures (ref) and (ref) provide the plots of Hill's estimates of the tail indices $\zeta$ in power law distributions for the time series $\Delta^2 I_t$ and $\Delta^2 D_t$ of daily changes in the number of COVID-19 infections and deaths in several countries considered (Australia, China, France, India, Italy, Russia, the UK and the US) with different number $k$ of extreme (largest) observations used in tail index estimation (the so-called Hill's plots, see Ch. 6 in embrechts1997modelling, embrechts1997modelling, and also Taleb, Taleb, for similar plots employed in the analysis of the inverse $\theta=1/\zeta$ of the tail index $\zeta$ in power law models ((ref)) for the number of deaths from major epidemic and pandemic diseases from ancient times until the present).\footnote{In the diagrams, $k$ plotted on the OX axis denotes the tail truncation level - the number of extreme observations - used in tail index estimation, in % of the total sample size $N,$ that is, $k$ equals 2.5-15% of the total sample size $N.$} Similarly, Figures (ref) and (ref) provide the log-log rank-size regression estimates of the tail indices $\zeta$ with optimal shifts $1/2$ in ranks proposed in gabaix2011rank for the time series $\Delta^2 I_t$ and $\Delta^2 D_t$ in the above countries that use different truncation levels $k$ for the largest observations used in inference (see IIK, IIK, for the analysis of such log-log rank-size plots for foreign exchange rates in emerging economies). The plots in Figures (ref)-(ref) also provide the corresponding 95% confidence intervals for tail indices $\zeta$ in power law models ((ref)) for the changes in the number of daily COVID-19 infections and deaths.
The analysis of Figures (ref)-(ref) indicates that Hill's and log-log rank-size regression tail index estimates for the time series $\Delta^2 I_t$ and $\Delta^2 D_t$ of daily changes in the number of COVID-19 infections and deaths tend to stabilize as a sufficient number $k$ of extreme (largest) observations (order statistics) on $\Delta^2 I_t$ and $\Delta^2 D_t$ is used in inference. As expected, log-log rank-size regression estimates tend to be less sensitive to the choice of $k$ compared to Hill's estimates.
Importantly, the left-end points of the confidence intervals for tail indices $\zeta$ in power law models for daily changes in COVID-19 infections and deaths calculated using different tail truncation levels $k$ in the countries in the diagrams tend to be less than two indicating possibly infinite second moments and variances. Further, from the analysis of Figures (ref)-(ref) it follows that the tail indices may be even less than one for some of the countries indicating extreme heavy-tailedness with possibly infinite first moments.
The conclusions from heavy-tailedness analysis for other countries considered in the paper are similar to those above.
The conclusions on heavy-tailedness in the time series of daily COVID-19 infections and deaths are important as, according to the above discussion, they point to high likelihood of observing their large values. They further emphasise the necessity in the use of robust methods in statistical analysis and forecasting of the dynamics of the COVID-19 pandemic and its impacts, including the approaches robust to the problems of heavy-tailedness and heterogeneity in the data.
This section presents the main results of the paper on statistically justified and robust evaluation of the effects of the COVID-19 pandemic on financial markets in different countries across the World. We focus on the analysis of predictive regressions for returns on major stock indices in the countries considered (see the introduction) incorporating the time series characterizing the infection and death rates from COVID-19 in the countries. Importantly, due to the problems of nonstationarity and the unit root dynamics in the time series $\Delta I_t, \Delta D_t$ of daily COVID-19 infections and COVID-19 related deaths in most of the countries discussed in Section (ref), estimation of the predictive regressions is provided for regression models for stock index returns $R_t$ with both the lagged daily infections/deaths $\Delta I_{t-1}, \Delta D_{t-1}$ and their (stationary) lagged differences $\Delta^2 I_{t-1}, \Delta^2 D_{t-1}$ (the daily changes in the number of COVID-19 infections and deaths) used as regressors.
More precisely, the estimation results are provided for predictive regressions in the form
where $R_{t}$ are the excess returns on major stock indices in the countries considered at the end of the day $t$ given by the difference between the end of the day-$t$ stock index returns and the countries' interest rates (see Section (ref)), and the regressors $X_{t-1}$ are either the number $\Delta I_{t-1}, \Delta D_{t-1}$ of COVID-19 infections/deaths on day $t-1$ in the countries dealt with or the daily changes $\Delta^2 I_{t-1}=\Delta I_{t-1}-\Delta I_{t-2},$ $\Delta^2 D_{t-1}=\Delta D_{t-1}-\Delta D_{t-2}$ in the number of infections/deaths.\footnote{The well-known stylized fact of absence of linear autocorrelations in daily financial returns (see cont2001empirical and references therein) implies exogeneity of regressors in the predictive regressions considered.}\footnote{We also conduct the analysis similar to that in the paper for the first and second differences of logarithms of daily infections and deaths. It points to unit root non-stationarity in the first log differences and the implied stationarity in the second log differences, similar to the differences of daily infections/deaths in Section (ref). We further obtain estimates in analogues of predictive regressions ((ref)) with the lagged first and the (stationary) second differences of logarithms of daily infections and deaths used as regressors. The conclusions from the estimates of such regressions are mostly similar to those for regressions with the above differences $\Delta I_t, \Delta D_t$ $\Delta^2 I_t, \Delta^2 D_t$ of daily infections/deaths used as regressors. The estimation results are available on request.}
In order to account for autocorrelation and heteroskedasticity in the regressors and the error terms in predictive regressions ((ref)) we use the widely applied HAC based methods (the standard errors and $t-$statistics with the quadratic spectral - QS - kernel and automatic choice of bandwidth as in andrews1991heteroskedasticity, andrews1991heteroskedasticity, and the corresponding $p-$values based on standard normal approximations) in the analysis of statistical significance of the regression coefficients.
It is well known, however, that commonly used HAC inference methods and related approaches based on consistent standard errors often have poor finite sample properties, especially in the case of pronounced dependence, heterogeneity and heavy-tailedness in the data (see the discussion and the analysis in IM, IM, IM1, Section 3.3 in ibragimov2015heavy, ibragimov2015heavy, and references therein). To account for these problems, we also provide the analysis of statistical significance of predictive regression coefficients using the $t-$statistic approaches to robust inference under dependence, heterogeneity and heavy-tailedness of largely unknown form recently developed in IM (IM, IM1). Following the approaches, robust large sample inference on a parameter of interest (e.g., a predictive regression coefficient $\beta$) is conducted as follows: The data is partitioned into a fixed number $q\ge 2$ (e.g., $q=2, 4, 8$) of groups, the model is estimated for each group, and inference is based on a standard $t-$test with the resulting $q$ parameter estimates.
In the context of inference on the coefficient $\beta$ in time series predictive regressions ((ref)), the regression is estimated for $q$ groups of consecutive time series observations with $(j-1)T/q<t\le jT/q,$ $j=1, ..., q,$ resulting in $q$ group estimates $\hat{\beta}_j,$ $j=1, ..., q.$ The robust test of a hypothesis on the parameter $\beta$ is based on the $t-$statistic in the group OLS regression estimates $\hat{\beta}_j,$ $j=1, ..., q.$ E.g., the robust test of the null hypothesis $H_0: \beta=0$ against alternative $H_a: \beta\neq 0$ is based on the $t-$statistic $t_{\beta}=\sqrt{q}\frac{\overline{\hat{\beta}}}{s_{\hat{\beta}}},$ where $\overline{\hat{\beta}}=q^{-1}\sum_{j=1}^q \hat{\beta}_j$ and $s_{\hat{\beta}}^2=(q-1)^{-1}\sum_{j=1}^ q (\hat{\beta}_j-\overline{\hat{\beta}})^2.$ The above null hypothesis $H_0$ is rejected in favor of the alternative $H_a$ at level $\alpha\le 8.3\%$ (e.g., at the usual significance level $\alpha=5\%$) if the absolute value $|t_{\hat{\beta}}|$ of the $t-$statistic in group estimates $\hat{\beta}_j$ exceeds the $(1-\alpha/2)-$quantile of the standard Student-$t$ distribution with $q-1$ degrees of freedom.\footnote{One-sided tests are conducted in a similar way, and the approaches further provide robust confidence intervals for the unknown parameters $\beta$ (IM1, IM1, ibragimov2015heavy, ibragimov2015heavy).}
The $t-$statistic based approaches do not require at all estimation of limiting variances of estimators of interest. As discussed in IM1, ibragimov2015heavy, IM, they result in asymptotically valid inference under the assumptions that the group estimators of a parameter of interest are asymptotically independent, unbiased and Gaussian of possibly different variances.\footnote{Justification of asymptotic validity of the robust $t-$statistic inference approaches in IM, IM1 is based on a small sample result in bakirov2006student that implies validity of the standard $t-$test under independent heterogeneous Gaussian observations and its analogues for two-sample $t-$tests obtained in IM1.} The assumptions are satisfied in a wide range of econometric models and dependence, heterogeneity and heavy-tailedness settings of a largely unknown type. The numerical analysis in IM, IM1, ibragimov2015heavy indicates favorable finite sample performance of the $t-$statistic based robust inference approaches in inference on models with time series, panel, clustered and spatially correlated data.\footnote{See also esarey for a detailed numerical analysis of finite sample performance of different inference procedures, including $t-$statistic approaches, under small number of clusters of dependent data and their software (STATA and R) implementation.}\footnote{The t-statistic robust inference approach proposed in IM provides a formal justification for the widespread Fama–MacBeth method for inference in panel regressions with heteroskedasticity (see FM, FM). Following the method, one estimates the regression separately for each year, and then tests hypotheses about the coefficient of interest using the t-statistic of the resulting yearly coefficient estimates. The Fama–MacBeth approach is a special case of the t-statistic based approach to inference, with observations of the same year collected in a group.}\footnote{See, among others, Bloom, Krueger, BW, verner, ChenRI and Tim for empirical applications of the robust $t-$statistic inference approaches in IM, IM1.} Importantly, the $t-$statistic based approaches to robust inference may also be used under convergence of group estimators of a parameter interest to scale mixtures of normal distributions as in the case of models under heavy-tailedness with infinite variances and in regressions with non-stationary exogenous regressors. \footnote{See Section 3.3.3 in ibragimov2015heavy for applications of the robust $t-$statistic approaches in inference in infinite variance heavy-tailed models. The recent works by Anat, Pedersen and IKP provide further applications of the approaches in robust inference on general classes of GARCH and AR-GARCH-type models exhibiting heavy-tailedness and volatility clustering properties typical for real-world financial returns, foreign exchange rates and other important economic and financial time series. The recent paper by IKS focuses on applications of the $t-$statistic approaches in inference on predictive regressions with persistent and/or fat-tailed regressors and errors. }
Tables (ref) and (ref) provide the results of the assessment of statistical significance of the coefficients $\beta$ on the lagged time series $\Delta I_{t-1}, \Delta D_{t-1}$ of daily COVID-19 infections/deaths and their differences $\Delta^2 I_{t-1}, \Delta^2 D_{t-1}$ - the daily changes in the number of infections and deaths from the disease - in predictive regressions ((ref)) for the countries considered. More precisely, the tables provide the values of HAC $t-$statistic with the QS kernel and the automatic choice of bandwidth discussed above as well as the values of the $t-$statistic $t_{\beta}$ in estimates $\hat{\beta}_j,$ $j=1, ..., q,$ of the slope parameter $\beta$ obtained using $q=4, 8, 12$ and 16 groups of consecutive time series observations. The asterisks in the tables indicate statistical significance of the slope coefficient (*** for significance at 1%, ** for significance at 5% and * for significance at 10%) implied by formal comparisons of the HAC $t-$statistics with the quantiles of a standard normal distribution. As described above, following the $t-$statistic approaches to robust inference in IM, IM1, (the absence of) statistical significance of the slope coefficient $\beta$ is assessed using the comparisons of the $t-$statistic in group estimates of the regression coefficient in the table with the quantiles of Student-$t$ distributions with $q-1$ degrees of freedom.
The values of HAC $t-$statistics in Tables (ref) and (ref) indicate an apparently spurious statistical significance of the (potentially non-stationary) lagged daily infections and deaths $\Delta I_{t-1}, \Delta D_{t-1}$ from COVID-19 in predictive regressions for returns on the main stock indices in some countries. In the case of daily COVID-19 infections, this is observed for the UK, Italy, Russia, the US, Brazil, China, South Korea and Australia, and, in the case of daily COVID-19 related deaths, for India, Canada, Brazil, Mexico, China, South Korea, Indonesia and Australia. The coefficients $\Delta I_{t-1}, \Delta D_{t-1}$ whose significance is indicated by the HAC $t-$statistics have the expected negative sign.
According to the HAC $t-$statistics in the tables, in the case of (econometrically justified) predictive regressions for excess returns on major stock indices considered involving (stationary) differences $\Delta I^2_{t-1},$ $\Delta D^2_{t-1}$ of the lagged daily COVID-19 infections and deaths (daily changes in the number of infections and deaths from the disease), the coefficients at the regressors $\Delta I^2_{t-1},$ $\Delta D^2_{t-1}$ are not statistically significant at conventional levels for most of the countries. Exceptions are, in the case of daily changes $\Delta I^2_{t-1}$ in the number of COVID-19 infections, are Italy, India, Brazil and Argentina, where the coefficient at the regressor $\Delta I^2_{t-1}$ is highly significant (at 1% level) according to the HAC $t-$statistics, with the expected negative sign, and also the UK where some significance (at 10% level) is observed, albeit with the positive sign at the coefficient. In the case of predictive regressions incorporating the daily changes $\Delta D^2_{t-1}$ in the number of COVID-19 related deaths, the HAC $t-$statistics indicate high statistical significance (at 1%) of the coefficient at the regressor $\Delta D^2_{t-1}$ only for the UK, India, Brazil, Mexico, South Korea and Australia; among these, the predictive regression coefficient has the expected negative sign only in the case of Australia.
However, the lagged daily COVID-19 infection and death rates $\Delta I_{t-1}$ and $\Delta D_{t-1}$ and their (stationary) differences $\Delta^2 I_{t-1}$ and $\Delta^2 D_{t-1}$ appear not to be statistically significant in predictive regressions for stock index returns in all countries considered according to the (econometrically justified) robust $t-$statistic approaches with different choices of the number of groups $q.$
\@ifstar{\starsection}{\nostarsection}{Conclusion} This paper presented the results of theoretically justified and robust statistical analysis of the effects of the COVID-19 pandemic on financial markets in different countries across the World. The analysis is based on robust inference in predictive regressions for the returns on the countries' major stock indices incorporating the time series characterizing the dynamics in the COVID-19 related deaths rates.
The paper further presented the results of the statistical analysis of (non-)stationarity, heavy-tailedness and tail risk in the time series on infections/death rates from COVID-19 in the countries considered. The obtained results point to non-stationary unit root dynamics and pronounced heavy-tailedness with possibly infinite variances and fist moments in the time series of daily COVID-19 infections and deaths in most of the countries dealt with.
According to the results in the paper, the standard HAC inference methods indicate apparently spurious statistical significance of the (potentially non-stationary) lagged daily infections/deaths from COVID-19 in predictive regressions for returns on the major stock indices in some countries. HAC methods also point to statistical significance of (stationary) lagged daily changes in the number of COVID-19 infections and deaths for some countries. For daily changes in COVID-19 infections, high statistical significance, with the expected negative signs of the predictive regression coefficients, is observed In the case of Italy, India, Brazil and Argentina.
Motivated by the results on high persistence and heavy-tailedness in daily COVID-19 infections/ deaths time series obtained in the paper and also by poor finite sample properties of HAC inference methods, we further provide the analysis of statistical significance of the coefficients in the predictive regressions using the recently developed $t-$statistic approaches to robust inference. The analysis using the robust $t-$statistic inference approaches indicates that the lagged daily COVID-19 death rates and their (stationary) differences appear to be statistically insignificant in predictive regressions for stock index returns in essentially all countries considered in the analysis.
The analysis and conclusions in the paper emphasize the necessity in the use of robust inference methods accounting for autocorrelation, heterogeneity and heavy-tailedness in statistical and econometric analysis and forecasting of key time series and variables related to the COVID-19 pandemic and its effects on economic and financial markets and society. They further emphasize the importance of the use of correctly specified models of the COVID-19 pandemic and its effects incorporating stationary time series and variables such as the daily changes in COVID-19 related deaths used in predictive regressions in this work.
Further research may focus on robust tests for structural breaks in models of the dynamics of the pandemic and its effects on financial and economic markets using the robust inference approaches based on two-sample $t-$statistics in IM1, and applications of inference methods such as sign- and rank-based tests that are robust to relatively small sample sizes of observations in statistical analysis of key models related to the spread of COVID-19. It would also be of interest to apply further estimation approaches for heavy-tailed models for time series associated with the pandemic that are robust to small samples, including the recently developed fixed-$k$ inference approaches for power-law models ((ref)) in MullerWang. The analysis in these directions is currently under way by the authors and co-authors.
\restoregeometry \makeatletter \let\origsection\@ifstar{\starsection}{\nostarsection}
\newcommand\nostarsection[1] { \origsection{#1} }
\newcommand\starsection[1] { \origsection*{#1} }
\makeatother
\setcounter{section}{0} \setcounter{subsection}{0} \setcounter{figure}{0} \setcounter{table}{0} \renewcommandAppendix \Alph{section}{Appendix \Alph{section}} \renewcommand\Alph{section}\arabic{figure}{\Alph{section}\arabic{figure}} \renewcommand\Alph{section}\arabic{table}{\Alph{section}\arabic{table}}
\@ifstar{\starsection}{\nostarsection}{Tables}
\@ifstar{\starsection}{\nostarsection}{Figures} \newgeometry{margin=2cm}