EconBase
← Back to paper

A Peek into the Unobservable: Hidden States and Bayesian Inference for the Bitcoin and Ether Price Series

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

49,808 characters · 14 sections · 74 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Peek into the Unobservable: Hidden States and Bayesian Inference for the Bitcoin and Ether Price Series

abstractConventional financial models fail to explain the economic and monetary properties of cryptocurrencies due to the latter's dual nature: their usage as financial assets on the one side and their tight connection to the underlying blockchain structure on the other. In an effort to examine both components via a unified approach, we apply a recently developed Non-Homogeneous Hidden Markov (NHHM) model with an extended set of financial and blockchain specific covariates on the Bitcoin (BTC) and Ether (ETH) price data. Based on the observable series, the NHHM model offers a novel perspective on the underlying microstructure of the cryptocurrency market and provides insight on unobservable parameters such as the behavior of investors, traders and miners. The algorithm identifies two alternating periods (hidden states) of inherently different activity -- fundamental versus uninformed or noise traders -- in the Bitcoin ecosystem and unveils differences in both the short/long run dynamics and in the financial characteristics of the two states, such as significant explanatory variables, extreme events and varying series autocorrelation. In a somewhat unexpected result, the Bitcoin and Ether markets are found to be influenced by markedly distinct indicators despite their perceived correlation. The current approach backs earlier findings that cryptocurrencies are unlike any conventional financial asset and makes a first step towards understanding cryptocurrency markets via a more comprehensive lens.

Introduction

Motivation, Methodology and Main Results

The present study is motivated by the still limited understanding of the economic and financial properties of cryptocurrencies. Sheding light on such properties constitutes a necessary step for their wider public adoption and is fundamental for blockchain stakeholders, investors, interested authorities and regulators (Co19,Li19,Pre20). More importantly, it may provide hints about market manipulation and fraud detection. Unfortunately, existing financial models that are used to study fiat currency exchange rates fail to capture the convoluted nature of cryptocurrencies (Ca19). The additional challenge that they face is the tight connection between cryptocurrency prices and the underlying blockchain technology which drives the dynamics of the observable market. To some extent, this is expressed via the particular market microstructure of cryptocurrencies: the market depth which depends on the exchange and the market maker, the functionality of exchanges as custodians (unique property among financial assets) and the absence of stocks, equities or other financial investment instruments (with the exception of Bitcoin futures, Ka19) which render acquiring and/or trading the cryptocurrency the main way of investing in this new technology (Ko18). The miners and/or stakers emerge as the main actors who drive the creation and distribution of the currency whereas the cheap and immediate transactions essentially obviate the need for conventional brokers. All these features (among many others), starkly distinguish cryptocurrencies from conventional financial assets or fiat money. However, a precise understanding of their defining financial and economic properties is still elusive (Br18,Co18,Ur18). With this in mind, the concrete research questions that we set out to understand are the following:

itemize[leftmargin=*] • How do cryptocurrencies compare -- in terms of their economic and financial properties -- to well understood financial assets like commodities, precious metals, equities and fiat currencies (Li17,Ba18,Kan19)? How do they relate to traditional financial markets and global macroeconomic indicators? • What are the defining microstructure characteristics of the cryptocurrency market and which are the distinguishing features (if any) between different coins (Kat19)?

To address these questions, we use a recently developed instance of Non-Homoge\-neous Hidden Markov (NHHM) modeling, namely the Non-Homogeneous P\'olya Gamma Hidden Markov model (NHPG) of Ko19,Kok20b, which has been shown to outperform similar models in conventional financial data (Me11). Using financial and blockchain specific covariates on the Bitcoin (Na08) and Ether (Bu14,But20) log-return series (henceforth BTC and ETH, respectively), the NHHM methodology aims not only to capture dynamic patterns and statistical properties of the observable data but more importantly, to shed some light on the unobservable financial characteristics of the series, such as the activity of investors, traders and miners. The present model falls into the Markov-switching or regime-switching literature with two possible states that is the benchmark for predicting exchange rates (En94,Le06,Fr05,Be16,Gr13,Wr09). This linear model was first introduced by Ha89 as an alternative approach to model non-linear and non-stationary data. It involves switches between multiple structures (equations) that can characterize the time series behavior in different regimes (states). The switching mechanism is governed by an unobservable state variable that follows a first-order Markov chain\footnote{For example, in the seminal paper of Ha89, the author used the underlying hidden process to define the business cycles (recession periods). More recent examples and a comprehensive theory about NHHM in finance can be found in Ma14.}. Therefore, the NHMM is suitable for describing correlated and heteroskedastic data with distinct dynamic patterns during different time periods, as are precisely cryptocurrency prices (Agg19,Bo19a,Kat19). Although standard in financial applications (Ma14), Hidden Markov models have only been applied in the cryptocurrency context by Po19 as state space models, by Ko18 to capture the liquitity uncertainty and Ph17 in the context of price bubbles. Yet, their more extensive use is supported by the specific characteristics of cryptocurrency data that have been identified by earlier research. Ba17,Ja18 and De18 demonstrate the non-stationarity of BTC prices and volume and underline the importance of modeling non-linearities in Bitcoin prediction models. This is further elaborated by Be16,Pi17,Ph18 and Yu19 who suggest that model selection and the use of averaging criteria are necessary to avoid poor forecasting results in view of the cryptocurrencies' extreme and non-constant volatility. Along these lines, Ci16 show that the Bitcoin price series exhibits structural breaks and suggest that significant price predictors may vary over time. Additional motivation for the analysis of cryptocurrency data with regime-switching models as the one employed here, is provided by Ka17 who demonstrate the heteroskedasticity of BTC prices and Bau18 who identify periods of different trading activity. Our main findings can be summarized as follows

itemize[leftmargin=*] • The NHPG algorithm identifies two hidden states with frequent alternations for the BTC log-return series, cf. (ref). State 1 corresponds to periods with higher volatility and returns and accounts for roughly one third of the sample period (2014-2019). By contrast, state 2 marks periods with lower volatility, series autocorrelation (long memory), trend stationarity and random walk properties, cf. (ref). At the more variable state 1, the BTC data series is influenced by miners' activity and more volatile covariates (stock indices) in comparison to more stable indicators (exchange rates) in state 2, cf. (ref). • The results for the hidden process are the same for both the long run (2014-2019) and the short run (2017-2019) BTC data, cf. (ref). However, differences in the significant predictors indicate more speculative activity in the short run compared to more fundamental investor behavior in the long run, cf. (ref). In sum, speculative activity (noise traders) is identified in the less frequent state 1 and in the short run whereas increased activity of fundamental investors is seen in state 2 and in the long run. • The algorithm does not mark a well defined hidden process with clear transitions for the ETH series, cf. (ref). This is further supported by the low number and the small values of significant predictors from the current set, cf. (ref). These results imply that ETH prices are still driven by variables beyond the currenlty selected set of predictors, showing characteristics of an emerging market that is more isolated than BTC from global financial and macroeconomic indicators.

More details are presented in (ref). Overall, the outcome of the NHPG model can be useful for investors and blockchain stakeholders by providing hints on periods of differentiating activities and effects in the cryptocurrency markets. From a theoretical perspective, it backs earlier findings that cryptocurrencies are unlike any other financial asset and suggests that their understanding requires not only the integration of existing financial tools but also a more refined framework to account for their bundled technological and financial features (Co18).

Related Literature

The literature on the financial properties of cryptocurrencies is expanding at an exponential rate and an exhaustive review is not possible (see Cor19 and references therein for a more comprehensive reference list). More relevant to the current context is the scarcity (to the best of our knowledge) of papers that address the bundled nature of cryptocurrencies as both blockchain applications and financial assets. Existing studies focus either on the underlying blockchain technology/consensus mechanism or on the observable financial market but not on both. By contrast, the current NHPG model parses the observable financial information to recover the underlying structure of cryptocurrency markets and hence makes a first step towards a unified approach to fill this gap. Its limitations are discussed in (ref). In the remaining part of this section, we provide a (non-exhaustive) list of studies that focus on the financial part. Early research, mainly focusing on BTC has provided mixed insights on the properties of cryptocurrencies. Kl18 claim that BTC is fundamentally different from valuable metals like gold due to its shortage in stable hedging capabilities. Along with Ch15, Ci16 also argue that standard economic theories cannot explain BTC price formation and using data up to 2015, they provide evidence that BTC lacks the necessary qualities to be qualified as money. However, Dy16 demonstrate that BTC has similarities to both gold and the US dollar (USD) and somewhat surprisingly, that it may be ideal for risk-averse investors. Bo17,Bo17a and Bo17b also explore BTC's characteristics as a financial asset and find that while BTC is useful to diversify financial portfolios -- due to its negative correlation to the US implied volatility index (VIX) -- it otherwise has limited safe haven properties. Using data from a longer period (between 2010 and 2017), De18 conclude the opposite, namely that BTC may indeed serve as a hedging tool, due to its relationship to the Economic Policy Uncertainty Index (EUI). In comparative studies, Fr18,Corb18 provide empirical evidence of bubbles in both BTC and ETH and Gk18 suggest that BTC is less risky than ETH, i.e., that it exhibits less fat tailed behavior. Phi18 confirm that Bitcoin exhibits long memory and heteroskedasticity and argue that cryptocurrencies display mild leverage effects, predictable patterns with mostly oscillating persistence, varied kurtosis and volatility clustering. Comparing BTC with ETH, they argue that kurtosis is lower for ETH being easier to transact than BTC. Along this line, the findings of Me19 and Kat19 further motivate the use of non-homogeneous and regime-switching modeling for both the BTC and ETH log-returns series. The differences between cryptocurrencies and conventional financial markets are further elaborated by Ka17,Ha17,Ph18. High volatility, speculative forces and large dependence on social sentiment at least during its earlier stages are shown by some as the main determinants of BTC prices (Ga15,Ge15,Co15,Yi18). Yet, a large amount of price variability remains unaccounted for (Ho18,Mc18,Ja18). Moreover, the proliferation of cryptocurrencies on different blockchain technologies suggests that their current correlation may be discontinued in the near future and calls for comparative studies as the one conducted here (Bo19).

Outline

The rest of the paper is structured as follows. In (ref), we describe the NHPG model and simulation scheme and present the set of variables that have been used (some preliminary descriptive statistics and tests about this data are relegated to (ref)). (ref) contains the main results and their analysis. In the first part ((ref)), we present the outcome of the algorithm and discuss the statistical findings for the hidden states and the generated subseries. In the second part, (ref), we focus on the significant explanatory variables for the BTC data series in both the short and long run and the ETH data series. We conclude the paper with a discussion of the limitations of the present model and directions for future work in (ref).

Methodology & Data

Given a time horizon $T\ge0$ and discrete observation times $t=1,2,\dots,T$, we consider an observed random process $\left\{Y_{t}\right\}_{t\le T}$ and a hidden underlying process $\left\{Z_{t}\right\}_{t\le T}$. The hidden process $\left\{Z_t\right\}$ is assumed to be a two-state non-homogeneous discrete-time Markov chain that determines the states ($s$) of the observed process. In our setting, the observed process is either the BTC or the ETH log-return series. Importantly, the description of the hidden states is not pre-determined and is subject to the outcome of the algorithm and interpretation of the results. Let $y_{t}$ and $z_{t}$ be the realizations of the random processes $\left\{Y_{t}\right\}$ and $\{Z_t\}$, respectively. We assume that at time $t,\ t=1,\dots,T$, $y_{t}$ depends on the current state $z_{t}$ and not on the previous states. Consider also a set of $r-1$ available predictors $\left\{X_{t}\right\}$ with realization $x_{t}=(1,x_{1t},\dots ,x_{r-1t})$ at time $t$. The explanatory variables (covariates) $\left\{X_t\right\}$ that are used in the present analysis are described in (ref). The effect of the covariates on the cryptocurrency price series $\{Y_t\}$ is twofold: first, linear, on the mean equation and second non-linear, on the dynamics of the time-varying transition probabilities, i.e., the probabilities of moving from hidden state $s=1$ to the hidden state $s=2$ and vice versa. Given the above, the cryptocurrency price series $\{Y_{t}\}$ can be modeled as \[Y_t\mid Z_t=s \sim \mathcal{N}(x_{t-1}B_s,\sigma^2_{s}), \;s=1,2,\] where $B_{s}=(b_{0s},b_{1s},\dots ,b_{r-1 s})'$ are the regression coefficients and $\mathcal{N}(\mu,\sigma^2)$ denotes the normal distribution with mean $\mu$ and variance $\sigma^2$. The dynamics of the unobserved process $\left\{Z_{t}\right\}$ can be described by the time-varying (non-homogeneous) transition probabilities, which depend on the predictors and are given by the following relationship $$P(Z_{t+1}=j\mid Z_t=i)=p^{(t)}_{ij}=\frac{\exp(x_{t}\beta_{ij})}{\sum^{2}_{j=1}\exp(x_t\beta_{ij})}, \; i,j=1,2,$$ where $\beta _{ij}=(\beta _{0,ij},\beta _{1,ij},\dots ,\beta_{r-1,ij})^{\prime }$ is the vector of the logistic regression coefficients to be estimated. Note that for identifiability reasons, we adopt the convention of setting, for each row of the transition matrix, one of the $\beta_{ij}$ to be a vector of zeros. Without loss of generality, we set $\beta_{ij}=\beta_{ji}=\mathbf{0}$ for $i,j=1,2, i\neq j$. Hence, for $\beta_i:=\beta_{ii},\;i=1,2$, the probabilities can be written in a simpler form $$p^{(t)}_{ii}=\frac{\exp(x_{t}\beta_{i})}{1+\exp(x_t\beta_{i})} \ \text{and}\ p^{(t)}_{ij}=1-p^{(t)}_{ii} ,\ i,j=1,2,\ i\neq j.$$ To make inference on the hidden process, we use the smoothed marginal probabilities $P(Z_t=i\mid Y_{1:T},z_{t+1},\theta)$ which are the probabilities of the hidden state conditional on the full observed process as derived from the Forward-Backward algorithm (Ha89). In the rest of the paper, we use the notation $P\(Z_t=i\right)$ for convenience.

algorithm[algorithm omitted — 1,932 chars of source]

Simulation Scheme

The unknown quantities of the NHPG are $\left\{\theta_s=\(B_{s},\sigma _{s}^{2}\right),\beta_s, s=1,2 \right\}$, i.e., the parameters in the mean predictive regression equation and the parameters in the logistic regression equation for the transition probabilities. We follow the methodology of Ko19. In brief, the authors propose the following MCMC sampling scheme for joint inference on model specification and model parameters.

enumerate[itemsep=0cm] • Given the model's parameters, the hidden states are simulated using the Scaled Forward-Backward of algorithm of Sc02. • The posterior mean regression parameters are simulated using the standard conjugate analysis, via a Gibbs sampler method. • The logistic regression coefficients are simulated using the P\'{o}lya-Gamma data augmentation scheme Po13, as a better and more accurate sampling methodology compared to the existing schemes.

The steps 1-3 of the MCMC algorithm are detailed in (ref).

Data

We assess the ability of 11 financial--macroeconomic and 3 cryptocurrency specific variables, outlined in (ref), in explaining and forecasting the prices of BTC and ETH via the NHPG model. In the related cryptocurrency literature these indices are commonly studied under various settings (Wi13,Ye15,Bo17b,Es17,Pi17,Ho18, Ja18 and Po19). The findings of the descriptive statistics and preliminary stationarity tests, cf. (ref), indicate that the logarithmic return (log-return), i.e., the change in log price, $r_t=\log\(y_t\right)-\log\(y_{t-1}\right)$, series of BTC and ETH exhibit trend non-stationarity, non-linearities, rich (i.e., non-random) underlying information structure and non-normalities. Based on these properties, the NHPG model seems appropriate for the study of the log-return data series. Accordingly, we apply the NHPG algorithm on daily log-returns of BTC and ETH, with normalized explanatory variables. We perform two experiments over two different time frames: in the first, we study the BTC series between 1/2014 and 8/2019 and in the second, we study both the BTC and ETH series between 1/2017 and 8/2019. The second time frame has been selected to allow reasonable comparisons between the BTC and ETH prices after eliminating an initial period following the launch of the ETH currency. It is further motivated by the outcome of a test-run of the NHPG model on BTC prices, cf. (ref), which indicates a transition point to a different period for the BTC price series in January 2017.

figure[figure omitted — 463 chars of source]
table[table omitted — 2,211 chars of source]

Results & Analysis

In this section, we discuss the findings from the NHPG model on the BTC and ETH log-return series. We first present the graphics with the output of the algorithm for the whole 2014-2019 period on BTC log-returns ((ref)) and the shorter 2017-2019 period on both BTC and ETH log-returns ((ref)). Then, we interpret the results and compare the statistical properties and the significant covariates between the two hidden states of both the BTC and ETH series and between the short and long run BTC series ((ref)).

Hidden States: Bitcoin 2014--2019

(ref) displays the BTC log-return series (blue line) along with the smoothed marginal probabilities (gray bars) of the hidden process being at state 1. Using as a threshold the probability $P\(Z_t=1\right)>0.5$, we estimate the hidden states for each time period. The NHPG model identifies a subseries of 667 observations in state 1 and a subseries of 1388 observations in state 2. The description of the hidden states is not predetermined by the model and is done a posteriori, by comparison of the statistical properties of the two subseries that have been generated. As it is obvious from (ref), state 1 corresponds to periods of larger log-returns and increased volatility in comparison to state 2. The frequent changes are in line with previous studies on the heteroskedasticity and on the regime switches (structural breaks) of the Bitcoin time series (Ph17,Ka17 and Ko18,Ar19, respectively). Yet, the refined outcome of the NHPG model, which determines the time periods that the series spends in each state, allows for a more granular approach. Specifically, it adds information about the significant covariates that affect both the observable and the unobservable process and on the financial properties of each state. This is done in (ref) below.

figure[figure omitted — 320 chars of source]

Hidden States: Bitcoin and Ether 2017--2019

(ref) shows the results of the NHPG model for both the BTC (left panel) and ETH (right panel) log-return series over the shorter 1/2017-8/2019 period. The algorithm has again identified two states in the BTC series, (ref), as indicated by the clear distinction between high-low marginal probabilities of state 1, i.e., $P\(Z_t=1\right)$, that are given by the gray bars. Moreover, a comparison with the same period in (ref) demonstrates that the NHPG has produced the same result (zoom in) -- in terms of statistical quality -- even over this smaller period, i.e., the algorithm has converged and returns essentially the same probabilities for the underlying process. However, as we will see below, cf. (ref), the statistical analysis unveils differences in the significant predictors and financial properties between the short and long run.

figure[figure omitted — 689 chars of source]

The picture is different for the ETH series, cf. (ref). Here, the hidden process is not well defined since the probabilities of state 1 at each time period are mostly close to $0.5$. This indicates high degree of randomness in the transitions of the algorithm and along with the low number of significant covariates that have been identified for ETH (cf. (ref) below), it suggests that ETH prices are still influenced by forces which are beyond the current set of financial and blockchain indicators (Ka18,Ph18,Phi18). This implies that ETH -- when viewed as a financial asset -- shows characteristics of an evolving, non-static and still emerging market. However, the relative isolation of ETH from other financial assets agrees with earlier findings in the literature (Ph18,Co18). Our next task is to provide additional insight on the structural financial and economic attributes that differentiate these two states for all experiments. Based on the similarities between the short and long run BTC time frames and the poor convergence of the algorithm for the ETH series, we focus on the long-run BTC series.

Hidden States: Financial Properties (BTC 2014-2019)

The results of both the descriptive statistics and the relevant statistical tests are summarized in (ref). Each entry -- BTC price, log-price and log-return series -- consists of two rows that correspond to the subseries of state 1 (upper row) and state 2 (lower row), respectively. The first two columns of (ref) verify that the estimated hidden process segments the series into two subseries with high/low mean and variance values for all the examined data series. Log-returns exhibit increased kurtosis in comparison to the initial estimates, cf. (ref), for both subseries (in particular for state 2). Similarly, the skewness of both subseries has increased and has turned positive with the skewness of the second subseries being again much higher than that of the first (cf. Tak18). These distributional properties lead to rejection of normality for either subseries and suggest the presence of heavy-tailed data (phenomena in which exreme events are likely, Zh18)\footnote{Inclusion of a third hidden state could potentially lead to smoothing of these measurements, cf. (ref).}. The identification of two subchains with different kurtosis and skewness can be a useful tool to investors (Jo03,Ko93,Di02). As risk measures, kurtosis and skewness cause major changes to the construction of the optimal portofolio (Ch97,Con13), especially in emerging and highly volatile markets (Ca07).

table[table omitted — 1,424 chars of source]

The asymmetry on the distributions and the difference of volatility between the two subchains can be related to the activity of informed or fundamental vs uninformed, noise or non-fundamental investors (or traders). Intuitively, the activity of uninformed investors leads to periods with higher volatility (cf. Bau18 and references therein). This is true for state 1 and refines the findings of Zar19b,Bau18 who attribute the informational inefficiency of BTC not only to its endogenous factors of an emerging, non-mature market but also to the non-existence of fundamental traders. The differences between the two states are further explained by the statistical tests. While the p-values of the Dickey-Fuller (DF) and Jarque-Bera (JB) remain the same as for the combined data series, cf. (ref), the results for the Ljung-Box-Q (LBQ), KPSS and Variance Ratio (VR) tests unveil different characteristics of the two subseries. In state 2 of the log-return series, the zero hypothesis is rejected for the LBQ test but not for the KPSS and VR tests. This suggests that the subseries defined by state 2 is a random walk with trend stationarity and long memory. These findings are related to (and to some extent refine) the results of Ji18,La18,Khu18,Me19,Zar19 by determining periods with (state 2) and without (state 1) permanent effects (long memory). The subchain of state 1 stills exhibits richer structure which can be potentially attributed to the combined activity and herding behavior of the non-fundamental traders (Bo19b,Sil19,St19).

Significant Explanatory Variables: Bitcoin and Ether

The second functionality of the NHPG model is to identify the significant explanatory variables from the set of available predictors that affect the underlying series both linearly, i.e., in the mean equation (observable process), and non-linearly, i.e., in the non-stationary transition probabilities (unobservable process). The algorithm also distinguishes between the variables that are significant in each state. The corresponding results for the BTC log-return series over both the 2014-2019 and 2017-2019 time periods are given in (ref) and the results for the ETH log-return series over the 2017-2019 time period are given in (ref). We use $B_i$ to denote the posterior mean equation coefficients and $\beta_i$ the posterior mean logistic regression coefficients for states $i=1,2$, as described in (ref).

table[table omitted — 4,349 chars of source]

The predictors that have been found significant at the 0.05 level are marked with bold font and $\ast$. The main findings are the following:

description[leftmargin=*] • The significant predictors (covariates) that dominate both the observable and the unobservable processes in the more volatile state 1 (cf. (ref)), correspond to more volatile financial instruments such as stock markets (S&P500 and NASDAQ). By contrast, state 2 is mostly influenced by the more stable exchange rates, cf. (ref). These findings suggest increased speculative activity in state 1 in comparison to fundamental investors in state 2. • While the algorithm has identified essentially the same hidden process for both the short and long run windows, cf. (ref), the significant predictors that affect both the observable and unobservable processes are remarkably different: more volatile for the short run versus more fundamental (monetary) for the long run. In line with Hor19, these findings provide evidence for increased speculative behavior in the short run. They also extend BTC's financial and safen haven properties to more recent windows (Po19,Ba17,Bo17a). Additionally, they refine the results of Cor18 and Cha19 who argue about the differences in the short and long run BTC markets and the hedging properties of BTC against volatile stock indices in time varying periods, respectively. • The lower number of significant predictors in the ETH log-return series reflects the inability of the NHPG model to parse the underlying process, cf. (ref). This differentiates the ETH from the BTC market and provides evidence that ETH is still at its infancy, evolving independently from established economic indicators and fundamentals. Yet, the main -- and somewhat unexpected -- conclusion is that, despite the evident correlation between the prices of BTC and ETH (Pearsons serial correlation 0.62), the two cryptocurrencies are affected by different fundamental financial and macroeconomic indicators over the same time period.
figure[figure omitted — 378 chars of source]

Finally, an observation that applies to all series is that the current set of predictors cannot fully explain the data volatility. Excluding the miners' activity (as expressed by the Hash Rate) which appears significant in state 1 for all series (both for the observable and the unobservable processes), this observation follows from the small values of the predictors in the mean equation of state 1 (cf. columns $B_1$ in (ref)) and the absence of predictors in the mean equation (observable process) of state 2 (cf. columns $B_2$ in (ref)).

table[table omitted — 2,533 chars of source]

Discussion: Limitations and Future Work

The application of NHHM modeling in cryptocurrency markets comes with its own limitations. From a methodogical perspective, the main concerns stem from the decision rule for each state which is probabilistic and the exogenously given number of hidden states. In the present study, we used the threshold of $0.5$ to decide transitions from state 1 to state 2 and vice versa. However, in the related financial literature, there are many different approaches even with lower thresholds. Moreover, while two hidden states are generally considered the norm in most financial applications, the current results suggest that it may be worth exploring the possibility of a third hidden state. Alternatively, one may define a gray zone for time periods in which the algorithm returns probabilities around 0.5 for both states. This will allow for the identification of periods with high uncertainty about the underlying process and hence, will lead to more scarce, yet more uniform (in terms of distributional properties) subseries. From a contextual perspective, the present approach does not account for qualitative attributes of the predictive variables. For instance, it does not measure centralization of the transactions or alleged fake volumes among different exchanges (Ga18,Bo19a). Coupling the present approach with transaction graph analysis, and user metrics to capture potential market manipulation and the purpose of usage such as speculative trading or exchange of goods and services (Ch15,Bl17,Ba18) will lead to improved results. Lastly, as more blockchains transition to alternative consensus mechanisms such as Proof of Stake, further iterations of the current model should also include the underlying technology (e.g., staking versus mining) as a determining factor Vot20. At the current stage, such a comparative study is not possible from a statistical perspective, since the market capitalization and trading volume of \enquote{conventional} Proof of Work cryptocurrencies is still not comparable to that of coins with alternative consensus mechanisms Oce20. The long-anticipated transition of the Ethereum blockchain to Proof of Stake consensus may define such an opportunity in the near future But20. Along these lines, extensions of the current model may enrich the set of covariates (explanatory variables) to capture technological features and/or advancements of various cyrptocurrencies, refine the NHPG model with potentially three hidden states and finally, couple the statistical/economic findings with transaction graph analysis. The expected outcome is a more detailed understanding of the financial properties of cryptocurrencies and the assembly of a model with improved explanatory and predictive ability for cryptocurrency markets.

Acknowledgements

Stefanos Leonardos and Georgios Piliouras were supported in part by the National Research Foundation (NRF), Prime Minister's Office, Singapore, under its National Cybersecurity R&D Programme (Award No. NRF2016NCR-NCR002-028) and administered by the National Cybersecurity R&D Directorate. Georgios Piliouras acknowledges SUTD grant SRG ESD 2015 097, MOE AcRF Tier 2 Grant 2016-T2-1-170 and NRF 2018 Fellowship NRF-NRFF2018-07.