Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
63,269 characters · 12 sections · 39 citation commands
Structural Estimation of Behavioral Heterogeneity
\thispagestyle{empty}
We develop a behavioral asset pricing model in which agents trade in a market with information friction. Profit-maximizing agents switch between trading strategies in response to dynamic market conditions. Due to noisy private information about the fundamental value, the agents form different evaluations about heterogeneous strategies. We exploit a thin set\textemdash a small sub-population\textemdash to pointly identify this nonlinear model, and estimate the structural parameters using extended method of moments. Based on the estimated parameters, the model produces return time series that emulate the moments of the real data. These results are robust across different sample periods and estimation methods.
Key words: asset pricing, behavioral finance, extended method of moments, identification, structural model
JEL code: C13, C58, G12, G17
Zhentao Shi (corresponding author): [email removed], Department of Economics, 912 Esther Lee Building, the Chinese University of Hong Kong, Shatin, New Territories, Hong Kong SAR, China. Tel: (852) 3943-1432. Fax (852) 2603-5805. Huanhuan Zheng: [email removed], Lee Kuan Yew School of Public Policy, National University of Singapore, 469C Bukit Timah Road, Singapore 259772. We benefit from in-depth discussion with Taisuke Otsu. We thank Zhenyu Gao, Oliver Linton, Peter Phillips, Michael Zheng Song and Jun Yu for helpful comments. All remaining errors are ours.
Financial markets undergo cycles of booms and busts. Price fluctuations generate profit opportunities for different investment strategies. No single investment strategy can always triumph\textemdash they also experience cycles of gain and loss in response to shifting market environment. It is essential for profit-seeking investors to choose their strategies according to the dynamic market conditions. We try to understand, theoretically and empirically, the impact of information friction on strategy switching. Our model follows the common approach in heterogeneous agent models (HAM), in which an agent selects from multiple investment principles such as the fundamental and technical trading strategies, while we introduce information friction to generate endogenous switching between different strategies.
In this model, every agent receives a private signal\textemdash an unbiased forecast about the fundamental value of the risky asset. Given the presence of information dispersion embodied in the realization of the private signal, the agents conceive different evaluations for the same investment strategy. As a result, each agent chooses, from a set of investment strategies, the one that maximizes the expected profit. The agents' actions reshape the asset price, and the evolutionary environment forces the agents to revise their subsequent choices in the next period. Such dynamic interaction between the agents and the asset price induces behavioral heterogeneity among agents and boom-bust cycles in the financial market.
We formally identify the proposed behavioral model via a thin set\textemdash a small subset of the population that reflects some special cases of the model khan2010irregular. We combine the unconditional moments, which involve the whole sample, with those conditional moments motivated from the thin sets. The two kinds of moments differ in the rates of convergence, so that the standard asymptotic theory for generalized method of moments (GMM) is not directly applicable. We employ extended method of moments (XMM) gagliardini2011efficient for estimation and statistical inference.
Applying XMM to historical observations of the Standard and Poor 500 index (S&P 500), we estimate and test our structural model in several sample periods. The predicted returns from the model closely match the real data in terms of the mean, standard deviation, skewness and kurtosis. Furthermore, we find empirical evidence that supports the evolutionary trading heterogeneity driven by information dispersion. When the price is relatively close to the fundamental value, an investment strategy based on the historical price trend is popular in the market. When the asset is excessively mispriced, however, the agents tend to switch to a fundamental strategy to pursue higher expected profits; their collective actions gradually drive the price toward the fundamental value, which corrects the market.
Our paper makes several contributions to the literature. In terms of modeling, the dynamics in trading heterogeneity has been modeled by latent boom-burst market states chiarella2012estimating, real business cycles lof2012heterogeneity, and switching stochastic processes brock1998heterogeneous. In particular, Markov transition of discrete regimes is popular in modeling the switching processes, and finds many empirical applications in the stock market, commodity market and derivative market frijns2010behavioral,jongen2012explaining,ter2013dynamic,eichholtz2015fundamentals. While these empirical papers directly model the aggregate time series, they leave unexplained why some agents switch their strategies but the others do not. Our new model combines he2016trading's microeconomic mechanism that endogenizes the Markov switching process and hirshleifer1992managerial's prioritization of profit instead of utility for the agency problem in asset management. Built on a microeconomic foundation of individual behavior, our theoretical model provides implication of the aggregate time series.
Econometric identification is the bridge that links the economic structural model and the data. Well-known is the difficulty to check identification in nonlinear models rothenberg1971identification,newey1994large,komunjer2012global. Formal identification is largely missing in the literature of HAM, where nonlinearity is the rule rather than the exception. While following the convention of HAM, we construct our model with identification in mind. The switching between the fundamental and technical strategies opens the opportunity for us to scrutinize in “slow motion” the instant, or the thin set, when the market is overwhelmed by one strategy. When a single strategy dominates, identification can be easily verified. We explore the thin-set identification and manage to recover all structural parameters in our model. To the best of our knowledge, this is the first paper that formally analyzes and establishes identification in the literature of structural modeling of heterogeneous behavior in the financial market.
In terms of estimation methods, existing empirical works of HAM mostly use nonlinear least squares boswijk2007behavioral,chiarella2012estimating,frijns2010behavioral, except that franke2012structural utilize the simulated method of moments (SMM). We derive an explicit formula of the pricing mechanism that implies moment restrictions in closed-form, which simplifies and speeds up the estimation. XMM is exactly the right bottle opener for a champagne brewed by the thin-set identification, thanks to the econometricians who crafted it.
The rest of the paper is organized as follows. Section (ref) develops the information-driven structural asset pricing model of behavioral heterogeneity. Section (ref) discusses identification, data handling, and estimation. Section (ref) reports the empirical findings, and compares them with those based on alternative approaches. Section (ref) concludes the paper. Moreover, we have prepared an Online Supplement with additional empirical results, extension, implementation, and examples.
In this section, we summarize the key building blocks of the information-based structural model. Step-by-step derivation of the model is given in Appendix Section (ref). A continuum (of measure one) of agents trade on one risky asset and one risk-free asset. The logarithm of the fundamental value of the risky asset at period $t$, denoted as $\mu_{t}$, is an exogenous random variable that market participants cannot interfere. It follows a random walk $\mu_{t}=\mu_{t-1}+\sigma_{\mu}\varepsilon_{t}^{\mu}$ for some $\sigma_{\mu}>0$, where $\varepsilon_{t}^{\mu}$ is independently and identically distributed across $t$ with mean and variance standardized as 0 and 1, respectively. Let $\boldsymbol{\mu}^{t}=\left(\mu_{t},\mu_{t-1},\mu_{t-2},\ldots,\mu_{0}\right)$ be the history of the fundamental value. At the beginning of period $t$, each agent receives a private signal $x_{it}=\mu_{t}+\sigma_{x}\varepsilon_{it}$, an unbiased forecast of the fundamental value $\mu_{t}$. The noise $\varepsilon_{it}|\boldsymbol{\mu}^{t}\sim\mathrm{i.i.d.}\Lambda$, where $\Lambda$ is a strictly increasing distribution function with the support of the real line, its density symmetric around 0, and the variance standardized as 1.
Let $\mathbf{p}^{t-1}=\left(p_{t-1},p_{t-2},\ldots,p_{0}\right)$ be the logarithm of the past price. Both $\mathbf{p}^{t-1}$ and $\boldsymbol{\mu}^{t-1}$ are public information for all investors at the beginning of time $t$. Each agent consults two financial advisors who conduct fundamental analysis and chartist analysis independently, to which we refer as $f$-advisor and $c$-advisor, respectively. The advisors make forecast according to their own perception of price movement, which may not be consistent with the true price formation mechanism. The $f$-advisor expects the price to respond to the fundamental value. Once she learns the private information $x_{it}$ from her client, she updates the expected $\mu_{t}$ to be $\frac{\mu_{t-1}+\alpha x_{it}}{1+\alpha}$, which is an average of $\mu_{t-1}$ and $x_{it}$ weighted by the precision (the inverse of variance), where $\alpha=\sigma_{\mu}^{2}/\sigma_{x}^{2}$ measures the precision of private information relative to public information. Believing in the efficient market hypothesis, she expects the period-$t$ return to be $\frac{\alpha\sigma_{x}}{1+\alpha}\left(\varepsilon_{it}-\delta_{t}\right),$ where $\delta_{t}=\left(\left(1+\alpha\right)p_{t-1}-\mu_{t-1}-\alpha\mu_{t}\right)/\left(\alpha\sigma_{x}\right)$. The $f$-advisor maximizes the constant absolute risk aversion (CARA) utility function and recommends the optimal investment flow $q_{it}^{f*}=\eta\frac{\alpha\sigma_{x}}{1+\alpha}\left(\varepsilon_{it}-\delta_{t}\right)$ into the risky asset, where $\eta$ is the trading intensity of the fundamental strategy with respect to asset mispricing.
In the meantime, the $c$-advisor utilizes technical analysis to forecast price movement. Her strategy is based only on the historical price trend, rather than $x_{it}$ or $\boldsymbol{\mu}^{t-1}$. Her expected period-$t$ return is $\Delta_{t-1}=p_{t-1}-p_{t-1}^{\mathrm{ref}}$, where $p_{t-1}^{\mathrm{ref}}$ is the reference price derived from certain technical rules. Under the same utility function, the $c$-advisor recommends the optimal investment flow $q_{t}^{c*}=\tau\Delta_{t-1}$ into the risky asset, where $\tau$ is the trading intensity of the chartist strategy. Unlike $q_{it}^{f*}$ that varies with $i$, for each individual $q_{t}^{c*}$ is the same.
We focus on the fundamental and technical strategies of bounded rationality out of many alternatives for the following reasons. (i) The two strategies are used commonly in practice allen1990charts. (ii) Models accounting for such two strategies are powerful in explaining financial market phenomena such as bubbles and crashes lux1995herd,huang2010financial and providing empirical specifications that outperform random walk chiarella2012estimating. (iii) Due to resource constraints, it is reasonable to prioritize investment strategies with good tracking records, supported by theoretical or empirical foundations; it is costly to hire a large number of financial advisors to conduct various analysis. (iv) No evidence suggests that other types of analysis consistently outperform fundamental and technical analysis in terms of profitability or utility.
Neither strategy is rational in that they ignore how agents' trading behavior affects the price. Forming rational expectation is difficult in the current setup due to the uncertainty about the convergence of the price to the fundamental value. It deviates from the rational expectation model, in which the price must return to its value at the terminal period. The fundamental strategy that utilizes private information does not always dominate the chartist strategy because the price\textemdash determined by the aggregate action of market participants\textemdash may not necessarily reflect the information.
Next, we discuss how the agents select trading strategies. Unlike the financial advisors who care about utility, the agents seek to maximize their investment profit in excess to the risk-free asset hirshleifer1992managerial.\footnote{Allowing the agents to have a different target function from financial advisors highlights the contrast between practitioners, who adopt straightforward criteria to swiftly respond to market, and researchers, who focus on sophisticated measures of utility. Our main results still hold when agents maximize CARA utility as their financial advisors do.} Let $\pi_{it}^{f}$ be the expected profit of the fundamental strategy based on the information available at the beginning of period $t$, and $\pi_{t}^{c}$ be that of the chartist strategy. In our model, we have
An agent chooses the strategy that yields higher expected profit. Due to the constraints on risk exposure and resources, we assume that every agent adopts one and only one strategy. Investors are not confident to select strategies that they are unfamiliar with, especially those insufficiently corroborated by studies or experience. It is therefore reasonable to presume that agents ignore strategies that are not scrutinized by their financial advisors.
Now we set equal $\pi_{it}^{f}$ and $\pi_{t}^{c}$ to solve the threshold signals that make agents indifferent between the fundamental and the chartist strategy. The quadratic form in ((ref)) yields the lower bound $\overline{\varepsilon}_{t}^{m}=\delta_{t}-\zeta_{t-1}$ and upper bound $\bar{\varepsilon}_{t}^{M}=\delta_{t}+\zeta_{t-1}$, where $\zeta_{t-1}=\frac{1+\alpha}{\alpha\sigma_{x}}\sqrt{\frac{\tau}{\eta}}\left|\Delta_{t-1}\right|$. The individual choice of the strategy hinges on the private signal. We assume that the agent will choose the fundamental strategy when she is indifferent between the two options. When $\varepsilon_{it}\in(-\infty,\overline{\varepsilon}_{t}^{m}]\cup[\overline{\varepsilon}_{t}^{M},\infty),$ the agent will adopt the fundamental strategy and we call her a fundamentalist. When $\varepsilon_{it}\in\left(\overline{\varepsilon}_{t}^{m},\overline{\varepsilon}_{t}^{M}\right),$ she will take the chartist strategy, and we call her a chartist. Given the distribution of the private signal, the fraction of chartists is \[ m_{t}=\Lambda\left(\overline{\varepsilon}_{t}^{M}\right)-\Lambda\left(\overline{\varepsilon}_{t}^{m}\right). \] The fraction of fundamentalists is $1-m_{t}$. If in addition $\Lambda$ is unimodal, an application of the Leibniz integral rule to $m_{t}$ shows that it strictly decreases in $\left|\delta_{t}\right|\in\left(0,\infty\right)$. Since $\left|\delta_{t}\right|$ captures the degree of mispricing, the fraction of chartists is relatively large (small) when the market is moderately (excessively) mispriced.
After selecting their preferred strategies at the beginning of period $t$, all agents place their trading orders simultaneously to a market maker. Following lux1995herd, we assume that the market maker adjusts the price according to \[ p_{t}\left(\theta\right)=p_{t-1}+\rho D_{t}\left(\theta\right), \] where $\rho>0$ is the marginal impact of aggregate demand on the asset price, $\theta=\left(\eta,\tau,\alpha,\sigma_{\mu}\right)$ is the set of the other structural parameters,\footnote{Since $\sigma_{x}=\sigma_{\mu}/\sqrt{\alpha}$, we do not need to include $\sigma_{x}$ into $\theta$ given the presence of $\sigma_{\mu}$ and $\alpha$.} and
is the aggregate demand in the market, where $\varphi\left(a\right)=\int_{-\infty}^{a}zd\Lambda\left(z\right)$ is the upper-truncated mean. According to the model, the stock market return follows
The above equation characterizes the asset price movements. It will be the key equation for the empirical estimation.
We apply the market-maker framework, instead of the market-clearing mechanism, because the former enables the nonlinear model to be analytically tractable over multiple horizons while the latter does not necessarily yield a solution for the equilibrium price. 10.2307/27647318 find that the market-maker mechanism performs as well as, if not outperforms, the market-clearing mechanism in terms of generating the efficient price.
We conclude this section by comparing our model with the Markov regime-switching regression. The Markov switching model is originated from hamilton1989new, and has been extended over the decades kim1994dynamic,kim1999has, with the latest development endogenizing the latent state variable kim2008estimation,chang2016new. Regime-switching models are featured by the transition probability among discrete states. In contrast, the microeconomic mechanism in our model dictates the variation of the fraction of agents who adopt either strategy in the dynamic market environment. On the one hand, our approach preserves the Markov property since $R_{t}\left(\theta\right)$ depends only on $\left(p_{t-1},p_{t-1}^{\mathrm{ref}},\mu_{t},\mu_{t-1}\right)$, which the econometrician directly observes when analyzing the data, but no other past observations. On the other hand, our approach differs from the Markov regime-switching regression as we do not directly model the aggregate time series. Instead, the aggregate market demand is generated by summing up the individual demand. Furthermore, an agent's switching between heterogeneous strategies is endogenous, because the threshold of the strategy choice is implied by the profit maximization problem. In other words, we attempt to provide a microeconomic foundation for the association between the latent states and the aggregate time series.
A model is judged not only by its microeconomic foundation, but also by its empirical fitness. We push the model to encounter data in this section. We verify that the structural parameters can be identified from the distribution of the observable random variables, and then propose an estimation procedure.
The structural model is a description of the data generating process, while the analysis of identification bridges the gap between the theoretical model and the observed data. The unobservable noises in the structural model stem from $\left(\varepsilon_{t}^{\mu},\varepsilon_{it}\right)$, which are independently and identically distributed across time. As a result, $\left(R_{t}\left(\theta\right)=\rho D_{t}\left(\theta\right)\right)_{t=1}^{T}$ is strictly stationary according to the model.
In reality, the econometrician observes two time series $\mathbf{p}^{T}$ and $\boldsymbol{\mu}^{T}$. If the observable random variables are truly generated from the theoretical model, can we uniquely determine the value of the “deep parameters” $\left(\sigma_{\mu},\eta,\tau,\alpha,\rho\right)$ from the joint distribution of $\left(\mathbf{p}^{T},\boldsymbol{\mu}^{T}\right)$? Obviously, $\sigma_{\mu}$ can be directly identified from $\boldsymbol{\mu}^{T}$. We narrow down the question to recovering the parameters $\left(\eta,\tau,\alpha,\rho\right)$ by matching the distribution of $\left(R_{t}\left(\theta\right)\right)_{t=1}^{T}$, which comes from the theory, with the distribution of the observable $\left(R_{t}^{\text{r}}=p_{t}-p_{t-1}\right)_{t=1}^{T}$, where the superscript “r” stands for “real”. Nevertheless, $\left(\eta,\tau,\rho\right)$ cannot be identified jointly. In view of ((ref)) and ((ref)), if we multiply $\rho$ by a non-zero constant and divide $\eta$ and $\tau$ by the same constant, the resulting $R_{t}\left(\theta\right)$ in ((ref)) remains. Hence we have to normalize $\rho=1$ and discuss the identification of the other three parameters $\left(\eta,\tau,\alpha\right)$.
It is well-known that global identification is often difficult in nonlinear models rothenberg1971identification,newey1994large,komunjer2012global. In the literature of HAM, identification of structural parameters is largely ignored. In this paper, we formally establish point identification for this highly nonlinear structural model. We take the thin-set identification approach khan2010irregular,lewbel2016identification, conditioning on some events that occur on a set of measure zero if the random variables are continuously distributed.\footnote{This thin-set identification strategy is not peculiar to our model. In Supplement Section S5, we provide examples in which thin-set identification can be invoked to establish point identification for other HAM models. } The key insight for the point identification is that when the event \[ G_{1}=\left\{ \Delta_{t-1}=0\right\} \] occurs, the expected return of the chartist strategy becomes zero, and all investors thereby turn to the fundamental strategy. Conditional on $G_{1}$, we have $\bar{\varepsilon}_{t}^{m}=\bar{\varepsilon}_{t}^{M}$ and $m_{t}=0$, and can simplify ((ref)) as
where $\tilde{z}_{1t}=\mu_{t-1}-p_{t-1}$ and $\tilde{z}_{2t}=\mu_{t}-p_{t-1}$ are observable, and $\theta_{1}=\eta/\left(1+\alpha\right)$ and $\theta_{2}=\eta\alpha/\left(1+\alpha\right)$ are explicit functions of the deep parameters. As long as the conditional distribution $\left(\tilde{z}_{1t},\tilde{z}_{2t}\right)\big|G_{1}$ is not perfectly collinear, we can identify $\theta_{1}$ and $\theta_{2}$, and then recover $\alpha=\theta_{2}/\theta_{1}$ and $\eta=\theta_{1}+\theta_{2}$. The occurrence of $G_{1}$ highlights the particular instant when the market is overwhelmed by the fundamental strategy, and the identification of $\alpha$ and $\eta$ follows.
Once $\alpha$ is identified, we can further condition on another event \[ G_{2}=\left\{ \tilde{\delta}_{t}\left(\alpha\right)=0\right\} \] where $\tilde{\delta}_{t}\left(\alpha\right)=\left(1+\alpha\right)p_{t-1}-\mu_{t-1}-\alpha\mu_{t}$. Under the event $G_{2}$, we verify in Appendix Section (ref) that ((ref)) becomes \[ R_{t}\left(\theta\right)=\psi\left(\varsigma_{t-1}\right)\tau\Delta_{t-1}=\psi\left(\sqrt{\tau}\frac{1+\alpha}{\alpha\sigma_{x}\sqrt{\eta}}\left|\Delta_{t-1}\right|\right)\tau\Delta_{t-1}, \] where $\psi\left(a\right)=2\Lambda\left(a\right)-1$ is strictly increasing, and non-negative when $a\geq0$.
Taking the expectation operator $E\left[\left|\cdot\right||G_{2}\right]$ on both sides of the above equation, we have \[ E\left[\left|R_{t}\left(\theta\right)\right|\big|G_{2}\right]=\tau E\left[\psi\left(\sqrt{\tau}\frac{1+\alpha}{\alpha\sigma_{x}\sqrt{\eta}}\left|\Delta_{t-1}\right|\right)\left|\Delta_{t-1}\right|\bigg|G_{2}\right]. \] Since $\left(\alpha,\sigma_{x},\eta\right)$ are already recovered, in the above equation $\tau$ is the only known parameter. Because the right-hand side is monotonically increasing in $\tau$ for any $\tau\geq0$ as long as $\Delta_{t-1}\neq0$, the parameter $\tau$ is identified.
The discussion of identification ensures that we can pin down the deep parameters from the observable time series given sufficiently many observations. We proceed to the estimation strategy.
Recall that $R_{t}^{\mathrm{r}}$ is the real return and $R_{t}\left(\theta\right)$ is the return according to the model. If the real data is truly generated from the structural model, the distribution of $\left(R_{t}^{\mathrm{r}}\right)_{t=1}^{T}$ must be the same as the that of $\left(R_{t}\left(\theta\right)\right)_{t=1}^{T}$. In reality, the structural model is at best a simplification of the real world.
Moment matching is one of the most popular econometric methods to estimate structural models. We estimate the structural parameter $\theta$ by matching moments of the marginal distribution of returns. First, as $\sigma_{\mu}$ is identified from the standard deviation of $\varepsilon_{t}^{\mu}$, we specify the first moment function \[ g_{1t}\left(\theta\right)=\left(\varepsilon_{t}^{\mu}\right)^{2}-\sigma_{\mu}^{2}, \] since $E\left[T^{-1}\sum_{t=1}^{T}g_{1t}\left(\theta\right)\right]=E\left[T^{-1}\sum_{t=1}^{T}\left(\varepsilon_{t}^{\mu}\right)^{2}\right]-\sigma_{\mu}^{2}=0$. Next, as the two parameters $\eta$ and $\alpha$ can be identified given $G_{1}$, we match the conditional mean and variance. Notice that these two moments are implied by the thin-set identification, and conditioning on $G_{1}$ literally means selecting only the observations such that $\Delta_{t-1}=0$. Since $\Delta_{t-1}$ is continuously distributed, the event $G_{1}$ happens with probability zero. To avoid the problem of too few local observations, we use a kernel function to assign weights to each observation, as in smith2007efficient and gospodinov2012local. We assign large weights on observations with small $\left|\Delta_{t-1}\right|$ and small weights on those with large $\left|\Delta_{t-1}\right|$. Given an appropriate bandwidth $h_{T}$, we would have enough observations to guarantee the estimation consistency at $\Delta_{t-1}=0$ asymptotically as $T\to\infty$. Let $w_{t}^{G_{1}}\left(h_{T}\right)=\phi\left(\Delta_{t-1}/h_{T}\right)$ be the weight of the $t$-th observation, where $h_{T}$ is the bandwidth and $\phi\left(a\right)=\left(2\pi\right)^{-1/2}\exp\left(-0.5a^{2}\right)$ is the density function of the standard normal. We construct two Gaussian-kernel-weighted moment functions
where $\tilde{R}_{t}^{\mathrm{r}}=R_{t}^{\mathrm{r}}-T^{-1}\sum_{t=1}^{T}R_{t}^{\mathrm{r}}$ is the demeaned $R_{t}^{\mathrm{r}}$, and $\tilde{R}_{t}^{2}\left(\theta\right)$ is defined similarly. On the other hand, the chartist parameter $\tau$ is identified conditional on $G_{2}$. The argument for identification of $\tau$ conditional on $G_{2}$ motivates another kernel-weighted moment function
where $w_{t}^{G_{2}}\left(\alpha,h_{T}\right)=\phi\left(\tilde{\delta}_{t}\left(\alpha\right)/h_{T}\right)$. We use the same bandwidth $h_{T}$ in $w_{t}^{G_{1}}\left(h_{T}\right)$ and $w_{t}^{G_{2}}\left(\alpha,h_{T}\right)$ for simplicity.
Under the assumption that the model is correctly specified, the moments \[ E\left[\left\{ g_{jt}\left(\theta\right)\right\} _{j=1,\ldots,4}\right]=0_{4\times1} \] pointly identify $\theta$. However, since $\left\{ g_{jt}\left(\theta\right)\right\} _{j=2,3,4}$ are constructed from a sub-population, they only use a small fraction of the data. As a consequence, the rates of convergence of the kernel-weighted sample moments are slower than the usual rate of $\sqrt{T}$, so are the rates of the estimated parameters. It is desirable to improve the rate of convergence of these parameter estimates by local identification information.
Following gagliardini2011efficient and antoine2012efficient, we assume local identification in the sense of rothenberg1971identification. That is, $\theta_{0}$ is locally identified if there exists an open neighborhood of $\theta_{0}$ containing no other $\theta$ that can generate the same distribution. Local identification does not contradict the thin-set identification. Local identification is based on the unconditional information of the population. The thin-set point identification here, however, relies on the conditioning of two special events that form the sub-population.
Assuming local identification, we further construct four unconditional moments with the whole sample. Specifically, we match the mean, variance, skewness and kurtosis of the returns:
We focus on these moments, thanks to the well-documented stylized facts about financial time series, i.e., excessive volatility, negative skewness, and fat tail in returns cont2001. Under local identification, these unconditional moment functions $\left\{ g_{jt}\left(\theta\right)\right\} _{j=5,\ldots,8}$ improve asymptotic efficiency of the estimator.
The construction of the moments gives a clear interpretation of indirect inference gourieroux1993indirect. While $\theta$ is the deep parameter from the structural model, those eight conditional and unconditional moments consist of a set of reduced-form parameters. The principle of indirect inference matches the reduced-form parameters from the observable data and the counterparts from the structural model. Model misspecification can be accommodated by indirect inference, in which the estimated structural parameter $\theta$ is the one that minimizes some distance between the reduced-form parameter from the real world and that from the economic theoretical model. Even though our stylized fundamentalist-chartist model is certainly a simplistic narrative, the estimation will tune the model to its best approximation to the features of the observed return time series.
The standard theory of GMM requires that all moments converge at rate $\sqrt{T}$. Such a premise is violated if we combine the eight moments $E\left[g_{jt}\left(\theta\right)\right]$, $j=1,\ldots,8$. The unconditional moments and conditional ones converge to their population means at different rates. Let $\mathbf{g}_{t}\left(\theta\right)=\left(g_{jt}\left(\theta\right)\right)_{j=1,\ldots,8}$ be the vector of the moment functions. Evaluated at a neighborhood of the true value, the (scaled) sample unconditional moments $T^{-1/2}\sum_{t=1}^{T}g_{jt}\left(\theta\right)=O_{p}\left(1\right)$ for $j\in\left\{ 1,5,\ldots,8\right\} $, while the (scaled) sample conditional moments $\left(Th_{T}\right)^{-1/2}\sum_{t=1}^{T}g_{jt}\left(\theta\right)/\sum_{t=1}^{T}w_{t}^{G_{1}}\left(h_{T}\right)=O_{p}\left(1\right)$ for $j\in\left\{ 2,3\right\} $ and $\left(Th_{T}\right)^{-1/2}\sum_{t=1}^{T}g_{4t}\left(\theta\right)/\sum_{t=1}^{T}w_{t}^{G_{2}}\left(\alpha,h_{T}\right)=O_{p}\left(1\right)$. With such a mixture of sample moments converging at various rates, the standard asymptotic theory of GMM is inapplicable. Fortunately, gagliardini2011efficient and antoine2012efficient have developed XMM, an extension of GMM, to explicitly incorporate moments with different rates of convergence. This latest methodological advancement makes the following empirical estimation possible.
We implement XMM by the continuous updating estimator (CUE) hansen1996finite. Let $\overline{g}_{j}\left(\theta\right)=T^{-1}\sum_{t=1}^{T}g_{jt}\left(\theta\right)$ be the simple sample average of $\left(g_{jt}\left(\theta\right)\right)_{t=1}^{T}$. Define the CUE criterion function as \[ J\left(\theta\right)=T\overline{\mathbf{g}}'\left(\theta\right)\widehat{\Omega}^{-1}\left(\theta\right)\overline{\mathbf{g}}\left(\theta\right), \] where $\overline{\mathbf{g}}\left(\theta\right)=\left(\overline{g}_{j}\left(\theta\right)\right)_{j=1}^{8}$, and $\widehat{\Omega}\left(\theta\right)$ is the sample long-run variance of $\left(\mathbf{g}_{t}\left(\theta\right)\right)_{t=1}^{T}$. CUE automates the choice of the weighting matrix so that we do not have to track the rate of each sample moment, and the scaling factors in the unconditional moments, $1/\sum_{t=1}^{T}w_{t}^{G_{1}}\left(h_{T}\right)$ and $1/\sum_{t=1}^{T}w_{t}^{G_{2}}\left(\alpha,h_{T}\right)$, are also canceled out in $\widehat{\Omega}\left(\theta\right)$.
We denote the XMM estimator as\footnote{ While standard kernel weight only depends on $h_{T}$, here $w_{t}^{G_{2}}\left(\alpha,h_{T}\right)$ also depends on $\alpha$. In Appendix Section (ref) we explain that it does not affect the asymptotic distribution of $\widehat{\theta}_{\mathrm{XMM}}$.}
Under regularity assumptions (see gagliardini2011efficient or antoine2012efficient), if $h_{T}\to0$, $h_{T}\sqrt{T}\to\infty$ as $T\to\infty$, we have
where $\Sigma$ is the asymptotic variance and it can be consistently estimated by \[ \widehat{\Sigma}=\left[\left(\frac{1}{T}\sum_{t=1}^{T}\frac{\partial}{\partial\theta}\mathbf{g}'_{t}\left(\theta\right)\right) \widehat{\Omega}^{-1}\left(\theta\right)\left(\frac{1}{T}\sum_{t=1}^{T}\frac{\partial}{\partial\theta'}\mathbf{g}_{t}\left(\theta\right)\right)\right]^{-1}\big|_{\theta=\widehat{\theta}_{\mathrm{XMM}}}. \] Regarding the model specification test, antoine2012efficient prove that this $J$-statistic still follows the usual $\chi^{2}$ distribution. With eight moments and four unknown parameters, the degrees of freedom of the $\chi^{2}$ distribution is 4.
We use Robert Shiller's S&P 500 dataset to construct the price and the fundamental (Downloadable at \url{http://www.econ.yale.edu/ shiller/data.htm}). The raw time series $\mathbf{p}^{T}$ is taken as the monthly average of the daily closing prices, and $\boldsymbol{\mu}^{T}$ is calculated as the present value of all monthly dividend flows according to the Gordon growth model gordon1959dividends.
Our discussion of the econometric procedure leaves open several choices in the implementation. We discuss these issues one by one. The observed return and the fundamental time series both exhibit upward trends. We have to filter the trends so that we can focus on the fluctuation of the stationary time series. Detrending does not change the behavior of the investors as the growth trend is incorporated in their decision of the quantity they purchase and the strategy they take. For simplicity, we fit a linear trend for each time series and then detrend. We find that the difference in the two trends is very small, which supports the implication of the efficient market hypothesis that the growth rate of $\boldsymbol{\mu}^{T}$ and $\mathbf{p}^{T}$ converge in the long run. We observe that detrending in our data preserves the pattern of over- and under-pricing periods as the crossing points of the two raw time series are proximate before and after detrending. Moreover, for the chartist strategy we need a reference price $p_{t-1}^{\mathrm{ref}}$. We use a simple 12-month moving average rule $p_{t-1}^{\mathrm{ref}}=\frac{1}{12}\sum_{s=t-12}^{t-1}p_{s}$.
We have assumed that the density function of $\Lambda$ to be symmetric and its support is the real line. Many distributions satisfy these conditions, for example the standard normal, the hyperbolic secant distribution, the Logistic distribution, the Laplace distribution, and the $t$-distributions of degrees of freedom at least 3 (with their variance standardized as 1). While $\varepsilon_{it}$ is unobservable, data provides no guidance about the choice of $\Lambda$. We select $\Lambda$ as the standard normal for its theoretical and practical attractiveness. Firstly, under normality the truncated mean function $\varphi\left(a\right)=-\left(2\pi\right)^{-1/2}\exp\left(-a^{2}/2\right)$ is the (minus) density function of $N\left(0,1\right)$, which is a built-in function in all modern statistical programming languages. Secondly, the normal distribution is favorable in justifying the conditional expectation of $\mu_{t}$ in the fundamental strategy. Given $\mu_{t-1}$ and $x_{it}$, if the fundamentalist takes a prior distribution $\varepsilon_{t}^{\mu}\sim N\left(0,1\right)$, she will attain the posterior distribution $\mu_{t}|\left(\mu_{t-1},x_{it}\right)\sim N\left(\frac{\mu_{t-1}+\alpha x_{it}}{1+\alpha},\frac{\sigma_{\mu}^{2}}{1+\alpha}\right),$ which delivers exactly the weighted average rule for the fundamentalist's expectation of $\mu_{t}$.
Throughout this paper, we use the same set of tuning parameters for all estimation procedures and sample periods. The bandwidth $h_{T}$ in the kernel-weighted sample moments is set as $1.06\widehat{\sigma}_{\Delta}T^{-1/5}$ according to silverman1986density's rule of thumb, where $\widehat{\sigma}_{\Delta}$ is the sample standard deviation of $\left(\Delta_{t}\right)_{t=1}^{T}$. The long-run variance is estimated using the Bartlett kernel newey1987simple; the number of lags in the kernel is chosen as $1.14\left\lfloor T^{1/3}\right\rfloor $ where $\left\lfloor \cdot\right\rfloor =\max_{b\in\mathbb{N}}\left\{ b\leq\cdot\right\} $, with the constant and the rate recommended in andrews1991heteroskedasticity. The rates of these tuning parameters satisfy the requirement for the asymptotic normality, and the estimates are stable in a reasonable range.
When applying XMM to the data, we set the compact parameter space $\Theta$ as $\left[0.001,3\right]^{3}\times\left[0.001,6\right]$, which is sufficiently wide for $\theta$. We must deal with the local optimizers in general nonlinear programming. We try many initial values to enhance the probability of capturing the global minimizer. The initial value for $\sigma_{\mu}$ is always the sample mean of $\left(\varepsilon_{t}^{\mu}=\mu_{t}-\mu_{t-1}\right)_{t=1}^{T}$. This sample mean is a consistent estimator, although in theory it is not as efficient as the XMM estimator since it does not incorporate the information provided by the other moments. For the other three parameters $\left(\eta,\tau,\alpha\right)$, the initial value is independently drawn from the uniform distribution over their parameter space. Given a randomly generated initial value, we carry out the nonlinear optimization. We repeat such optimization for 100 times, save each local minimum, and take the smallest one as the global minimum.
In this section, we report the empirical results and compare them with alternative specifications. We first estimate the parameters with a recent time span from January, 1991 to December, 2013, to which we refer as Period 1. We then repeat the estimation procedure for two alternative time spans: January, 1961\textemdash December, 1990 (Period 2), and January, 1911\textemdash December, 1960 (Period 3) for robustness check.
In Section (ref), all the eight moments are incorporated in the estimation, to which we refer as the full model. Furthermore, we evaluate the effect of the kernel-weighted moments in Section (ref), and the mixture of the two strategies in Section (ref).
The time series of the linearly-detrended price $\mathbf{p}^{T}$ and fundamental value $\boldsymbol{\mu}^{T}$ in Period 1 are shown in the upper panel of Figure (ref). It is apparent that the price is more volatile than the fundamental. The price sometimes deviates significantly away from the fundamental value, which corresponds to boom-bust episodes in the financial history. In the long run, the price tracks the fundamental value in general, which supports the market efficiency theory in a long-term perspective.
We take XMM as our benchmark. We report the XMM estimates of $\theta=\left(\sigma_{\mu},\eta,\tau,\alpha\right)$ and the two-sided 95% asymptotic confidence intervals in Table (ref) for each sample period. All estimates are positive and none of the confidence intervals contains 0, which is consistent with the economic interpretation of these parameters.
The parameters $\eta$ and $\tau$ represent the trading intensity of the fundamental strategy and the chartist strategy, respectively. Based on the estimation results of Period 1, the estimate of $\eta$ means that raising the expected return of the fundamental strategy by 1% increases the investment flow by 0.10% on average. On the other hand, the estimate of $\tau$ implies that 1% change in the expected return of the chartist strategy leads to a 0.61% hike in the investment flow. The estimate of $\alpha$ is $1.71$ suggests that investors update their expected fundamental value aggressively by overweighing the private information relative to the common prior on the historical fundamental value, as the private information is more precise than the public information. In terms of the model specification test, the $J$-statistic is 2.22 with the $p$-value 0.70. It does not reject the model, indicating that our model can be a reasonable description of the data generating process for this sample period.
Besides the values of the structural parameters, we are also interested in the endogenous switching of the financial agents between the chartist and fundamental strategies. It is illustrated in the middle panel of Figure (ref). Consistent with the model's prediction, the market is dominated by fundamentalists when the asset is excessively mispriced, and by chartists when the price moves more closely around the fundamental value. How agents switched between the heterogeneous strategies during the recent global financial crisis is of particular interest. When the market was booming during 2005\textendash 2007, many agents clustered to be chartists, who traded on price trends. When the trend was reversed in late 2007, the market fraction of chartists declined sharply. Fundamentalists prevailed the market in 2008\textendash 2009, the most volatile years during the global financial crisis. In that episode, financial assets were overwhelmingly underpriced (as illustrated in the upper panel of Figure (ref)), and fundamentalists had accumulated strong buying force that drove the price up toward its fundamental value. However, the price may not converge to the fundamental value immediately after the fundamentalists occupy the market. The presence of information friction produces such inertia in our model. No similar patterns of switching was found during the dot-com crisis. In the early period of dot-com bubble formation, chartists dominated the market. As the asset became more and more overpriced, agents switched to fundamentalists. The bubble continued to grow even after fundamentalists fully occupied the market. In a highly noisy environment, some fundamentalists might wrongly extrapolate the asset to be underpriced even if it was actually overpriced.
To examine the performance of moment matching, we plug in the estimated parameters into the model to predict the return. In the lower panel of Figure (ref), the solid line is the empirical cumulative distribution function (ECDF) of the real data $\left(R_{t}^{\mathrm{r}}\right)_{t=1}^{T}$, and the dashed line is the ECDF of the fitted return series $\left(R_{t}\left(\widehat{\theta}_{\mathrm{XMM}}\right)\right)_{t=1}^{T}$. The two ECDF curves closely track each other. As shown in the upper panel of Table (ref), the predicted returns generated from XMM have a mean return close to zero, a variance around 0.04, a negative skewness, and a kurtosis that is larger than 3. These sample moments are similar to those of the real return series.
Next, we repeat the same exercises for other sample periods to check the robustness of the empirical results. Figure (ref) and the second column in Table (ref) display the results of Period 2. Again, the $J$-statistic does not reject the model specification. The estimated coefficients are comparable with those in Period 1. The fitted returns match well with the real data in terms of ECDF and the four moments, as shown in the middle panel of Table (ref). Moreover, consistent with the previous results, we observe from the middle panel of Figure (ref) that chartists prevailed when the asset was moderately priced, for example in 1976, while fundamentalists dominated the market when the price deviated significantly away from the fundamental, for example in 1978\textendash 1982.
Figure (ref) and the third column of Table (ref) report the results for Period 3, a half century that witnessed the Great Depression. The high volatility in this era is manifest as shown in Table (ref), with a kurtosis of 15.30 in the real return, the largest among the three sample periods. In terms of the point estimates, the scale of the estimated coefficients $\tau$ and $\eta$ are larger than those reported in the other two periods, showing that both fundamentalists and chartists responded more sensitively to the expected returns. The estimated coefficient $\alpha$ in Period 3 is about twice as large as that in Period 1 or 2, suggesting fundamentalists updated information more aggressively in response to the volatile market. In Figure (ref) we again observe the switching from chartists to fundamentalists when the asset was excessively mispriced and from fundamentalists to chartists when the asset was moderately mispriced.
In the following sections, we estimate some simple alternative models and compare the empirical results with those discussed in this section.
The standard GMM utilizes only the unconditional moments in estimation. For comparison, we implement GMM (CUE) with the moment functions $\left\{ g_{jt}\left(\theta\right)\right\} _{j=1,5,6,7,8}$, and the results are reported in Table (ref). Ignoring $\left\{ g_{jt}\left(\theta\right)\right\} _{j=2,3,4}$, which contains information from the theoretical model, weakens the asymptotic efficiency of parameter estimation. In our context, such efficiency loss is reflected in the confidence intervals\textemdash in most cases the confidence intervals of the GMM estimator are wider than their XMM counterparts. In particular, the confidence interval of $\eta$ includes 0, which is highly undesirable since the identification of the parameters relies on a positive $\eta$. In contrast, when conditional moments are accounted for, the confidence intervals of $\eta$ are clearly deviated away from 0 (see Table (ref)). In the meantime, with fewer restrictions GMM improves the in-sample fitting. The model emulates the data more closely in terms of moment matching, as shown in Table (ref) with the kurtosis of the predicted return closer to that of the real data.
A common feature of the strategy switching in Figure (ref)\textendash (ref) is that fundamentalists dominate the market more frequently than chartists. This observation raises the question of the necessity of introducing the two investment strategies to characterize the price movement. This section explores whether a solo-strategy model is sufficient to capture the price dynamics.
The fundamentalist-only model is a sub-model of the two-strategy benchmark model. When $\tau=0$ and $\eta>0$, the chartist strategy generates zero profit so that no investor will adopt it. The predicted return of the fundamentalist-only model follows ((ref)). Since the kernel-weighted moment functions $g_{jt}\left(\theta\right),$ $j\in\left\{ 2,3,4\right\} $, remain valid in the sub-model, we estimate the parameters $\left(\sigma_{\mu},\eta,\alpha\right)$ by XMM with the same eight moments as in ((ref)) but setting $\tau=0$. With the restriction $\tau=0$, the $J$-statistic follows $\chi^{2}\left(5\right)$ asymptotically under the null.
We report the empirical results in the upper panel of Table (ref). The estimates stay positive and statistically significant as the 95% confidence intervals are all above 0. This is consistent with the results from the full model and provides evidence of the presence of the fundamentalist trading in the market. Under the null hypothesis that the fundamentalist-only model is correctly specified, the $J$-statistics are 15.74, 16.44 and 32.15 in Period 1\textendash 3, respectively, which are associated with $p$-value less than 1%. The strong rejection means that the fundamental strategy solely is incapable of mimicking the observed price movements. Moreover, in Table (ref) the moments of fitted returns are far away from the real ones. In particular, the fitted kurtosis is less than 3 throughout the three sample periods, which contradicts the fat-tail phenomenon observed in the real data.
If the fundamentalist-only model is insufficient to capture the real return, how about the chartist-only model? The lower panel of Table (ref) displays the XMM estimation results across the three sample periods.\footnote{A formal test of the chartist-only model is more complicated than the fundamentalist-only model because the full model precludes $\eta=0$ due to its presence in the denominator in $\zeta_{t-1}$. The implementation is detailed in Supplement Section S2.} The $J$-statistics clearly reject the chartist-only model at any commonly used test size, and the moment matching in Table (ref) is poorer than the full model.
In view of the empirical results, neither the fundamentalist strategy nor the chartist strategy alone reasonably matches the data. The mixture of the two trading strategies is effective in improving the model fitting.
In this paper, we develop a structural asset pricing model with information-driven behavioral heterogeneity. For this highly nonlinear model, we formally identify the structural parameters via thin-set identification. The thin-set identification and the follow-up estimation techniques are applicable to other heterogeneous agent models involving a mixture of investment strategies.
We estimate the parameters by XMM, and conduct inference for the model specification. The empirical results show that the structural model emulates the S&P 500 index. Investors switch between the fundamental and chartist strategies evolutionarily in response to the dynamic market conditions. Agents tend to cluster toward the chartist strategy when the market environment waxes and wanes, and their collective trading actions cause substantial asset mispricing that sometimes turn into bubbles and crashes. However, when the asset is significantly overpriced or underpriced, agents tend to revert to the fundamental strategy, which corrects the mispricing and restores the market efficiency. The switching is found to be crucial for the empirical fitness of the structural model. Models with only one strategy significantly underperform the structural model in terms of matching the real price movement.
In this Appendix of this paper, we present the step-by-step development of the structural model as well as the derivation of some technical claims in the main text. Moreover, we provide an Online Supplement for additional empirical results, extension, and implementation.