Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
139,841 characters · 31 sections · 66 citation commands
Uniform Inference for Characteristic Effects of Large Continuous-Time Linear Models
\singlespacing
Key words: Large dimensions, high-frequency data, cross-sectional bootstrap, GMM
\eject
\onehalfspacing
Conditional factor models have been playing an important role in capturing the time-varying sensitivities of individual outcomes to the factors, in which the factor betas are varying over time. In this paper, we study a continuous-time conditional factor model:
where $\bfm Y_0$ is the starting value of the outcome process at time 0; the model contains a factor, idiosyncratic, and a drift process $\{(\bfm F_t, \bfm U_t, \bfsym \alpha_t): t\geq 0\}$. Here $(\bfm Y_t, \bfsym \alpha_t, \bfm U_t)$ are all high-dimensional (whose dimension $p\to\infty$), and the dimension of the factor process $\bfm F_t$ is fixed. In addition, the process $\{\bfsym \beta_t: t\geq 0\}$ represents the factor loadings (or “betas"), which is assumed to be stochastic in this paper. Model ((ref)) covers many uses of factor models in asset pricing. In financial economics, extensive empirical studies have shown that assets' individual betas can be largely explained by their “characteristics" (or called “instruments"). These include lagged characteristics that are common to all stocks, characteristics specific to individual stocks, as well as observations of other firm instruments gagliardini2016time. Estimated betas as functions of the conditioning characteristics represent the effects of characteristics on firm specific sensitivities to the risk factors. They pick up long-run patterns and fluctuations in the betas. Therefore, estimating the characteristic effects on the individual betas is one of the central econometric tasks in financial economics. Given the importance of the topic, the econometric problem, however, is challenging for the reasons we shall elucidate below.
Let $\bfsym \beta_{lt}$ denote the $K$-dimensional factor betas of the $l$-th individual at time $t$. Let $\bfm X_{lt} $ be a vector of observed characteristics that may be varying across individuals and times, and $\bfm X_t=\{\bfm X_{lt}: l\leq p\}$. We model: for each $l=1,...,p$,
where $\bfm g_t(\cdot)$ is an unknown time-varying nonparametric function that is assumed to be well approximable by a sieve representation. Here, each component $\gamma_{lt,r}$ of $\bfsym \gamma_{lt}$ ($r\leq K$) satisfies, for some constant $C>0$,
The main message from ((ref))-((ref)) is that in the identification condition $\bfm g_{t}(\bfm X_{lt})= \mathbb E( \bfsym \beta_{lt}| \bfm X_{lt} ) $, the variance of the “error components" $\bfsym \gamma_{lt}:= \bfsym \beta_{lt} - \bfm g_{t}(\bfm X_{lt}) $ is defined on a compact set that includes zero, and can be either exactly on, arbitrarily close to, or bounded away from its “boundary". Thus we allow arbitrarily unknown signal strengths of $\bfsym \gamma_{lt}$. So the economic meaning of model ((ref)) is that the factor beta is decomposed as the sum of two components: (i) a nonparametric function of the observed instruments, $\bfm g_{lt}:=\bfm g_{t}(\bfm X_{lt})$, which we call “characteristic betas", and (ii) a time-varying and individual specific component, $\bfsym \gamma_{lt}$, which we call “idiosyncratic betas". The characteristic beta $\bfm g_{lt}$ picks up long-run beta patterns and fluctuations, and possesses less volatile, while $\bfsym \gamma_{lt}$ captures high frequency movements in beta, and represents the remaining time-varying individual factor sensitivities after conditioning on the observed characteristics. In addition, the strength of the idiosyncratic betas is allowed to be arbitrary, reflecting the nature that the explainability from the characteristics is unknown.
The goal of this paper is to provide a uniformly valid inference of $\bfm g_{lt}$, the characteristic effects on factor betas. By “uniformly valid", we mean the coverage probability is asymptotically correct uniformly over a broad class of data generating processes (DGPs) that allows various possible signal strengths of $\bfsym \gamma_{lt}$, measured by a weighted cross-sectional variance: $$ \bfm V_{\gamma, t} := \operatorname{Var}\left(\frac{1}{\sqrt{p}}\sum_{m=1}^ph_{t,ml}\bfsym \gamma_{mt}\bigg{|}\bfm X_t\right). $$ We assume $\bfsym \gamma_{lt}$ to be conditionally cross-sectionally independent given $\bfm X_t$, and find that the strength of $ \bfm V_{\gamma, t}$ plays a crucial role in the asymptotic behavior of estimated characteristic effect, and affects both the rate of convergence and limiting distributions. Here $h_{t,ml}$ is a function of $\bfm X_t$ whose definition will be clear in the paper. In particular, the asymptotic distribution of $\bfm g_{lt}$ has a “discontinuity" when the strength of $\bfsym \gamma_{lt}$, the eigenvalues of $ \bfm V_{\gamma, t}$, are near zero. As a consequence, the usual pointwise inference procedures under a fixed DGP only produce confidence intervals that are valid for specific DGPs, therefore potentially produce misleading inferences. Specifically, benchmark methods in the literature, which ignore the high-frequency beta dynamics in $\bfsym \gamma_{lt}$, would produce under-coveraging confidence intervals of the characteristic effects. On the other hand, we show that even if $\bfsym \gamma_{lt}$ is allowed, standard “plug-in" procedures using the estimated asymptotic variances do not produce uniformly correct coverage probabilities, because they require very strong signal strengths of $\bfm V_{\gamma, t}$, leading to over-coveraging confidence intervals when the signal strength is weak. The discontinuity issue here is similar to the problem of estimating parameters on a boundary. As is shown by, e.g., andrews1999estimation and ketz2017testing, when a test statistic has a discontinuity in its limiting distribution, as occurs in estimating parameters on a boundary and in random coefficients models, pointwise asymptotics can be very misleading.
We reply on a cross-sectional bootstrap to achieve the uniform inference. It is important to note that the employed bootstrap is cross-sectional, which resamples the cross-sectional individuals while keeping all the serial observations for each resampled individual. This procedure is essential because the discontinuity arises when the cross-sectional variance $ \bfm V_{\gamma, t}$ is near the boundary, and the cross-sectional bootstrap avoids estimating $ \bfm V_{\gamma, t}$ in the current context. We show that the bootstrap procedure leads to a correct asymptotic coverage probability and is uniformly valid over a large class of DGPs, and explain the reasons in detail.
Allowing $ \bfm V_{\gamma, t} $ to be unrestricted on a compact set that includes zero as the boundary point is the main motivation of this paper, which arises from the following practical consideration. While characteristics may fully explain the factor betas at times when they are just updated and made publicly available ($\bfsym \gamma_{lt}\approx0$), there are also times when factor betas contain either unmeasurable or high-frequency components that are more volatile and cannot be captured by the characteristics ($\bfsym \gamma_{lt} $ very different from zero). In those occasions, modeling the beta as fully specified functions of observed characteristics can be very restrictive. This is particularly true for high-dimensional and high-frequency factor models in empirical asset pricing, where individual factor betas demonstrate large heterogeneity when the number of assets is large, and assets' returns are available at a very high frequency. On the contrary, characteristics such as the firm sizes and book-market values, often vary more smoothly and are measured at a much lower frequency, often (but not always) leaving large portions of stock betas' dynamics unexplained. As we show in this paper, without taking into account the high-frequency movements of betas after conditioning on the characteristics, the inference procedures of characteristics' effects are not asymptotically valid. Unfortunately, this is often the case in the financial economic literature, which has been dominated by modeling betas as fully specified functions of the observed characteristics, including both parametric (e.g., shanken1990intertemporal, cochrane1996cross, ferson1999conditioning, avramov2006asset, gagliardini2016time) and nonparametric models, e.g., CL07 and CMO. As is shown by ghysels1998stable, misspecifying beta risk may result in serious pricing errors that might even be larger than those produced by an unconditional asset pricing model.
The model can be generalized to the high-dimensional continuous-time linear moment conditions framework. For each $l\leq p,$ the parameter $ \bfsym \beta_{lt} $ is identified by the following moment condition:
where $\bfm c_{z,lt}=d[\bfm Z_l,\bfm Z_l]_t/dt$ is the instantaneous quadratic variation process of $\bfm Z_l=\{\bfm Z_{lt}\}_{t\geq 0}$, and $ \Psi_l(\bfsym \beta,\bfm c)$ is a known function linear with respect to $\bfsym \beta$. In addition, we observe a set of time-varying characteristics $\bfm X_{lt}$, so that we have the following decomposition:
where $\mathbb E(\bfsym \gamma_{lt}|\bfm X_{t})=0$. The model admits the linear factor model as a special case by setting $\bfm Z_{lt}=(Y_{lt}, \bfm F_t)$ and $$\Psi_l(\bfsym \beta_{lt}, \bfm c_{z, lt})=\bfm c_{FF,t}\bfsym \beta_{lt} -\bfm c_{YF,lt},\quad \bfm c_{z, mt}=(\bfm c_{FF,t}, \bfm c_{YF,mt}). $$ Here $\bfm c_{FF,t} $ and $\bfm c_{YF,lt} $ are the quadratic covariances for the processes $\bfm F_t$ and $Y_{lt} $. Moreover, the above model also admits a continuous-time linear regression model with individual-specific regressors: e.g., barndorff2004econometric,mykland2006anova,kalnina2012nonparametric, as well as the idiosyncratic volatility model applied by ang2009high,herskovic2016common, li2016inference.
We provide a general two-step estimation procedure to make uniform inferences about the characteristic effect $\bfm g_{lt}= \bfm g_{t}(\bfm X_{lt}) $ for each fixed $l\leq p$. In step (i), we estimate $\bfsym \beta_{lt}$ by either directly solving ((ref)) or using generalized method of moments (GMM), with the sample quadratic variation in place of $ \bfm c_{z, lt}$, and in step (ii), we estimate $\bfm g_{lt}$ by a standard nonparametric sieve regression on ((ref)) with the estimated $\bfsym \beta_{lt}$. We aim to construct a confidence interval $CI_{\tau}$ for $\bfm g_{lt} $ for any specific $l\leq p$ and at a specific time $t$, so that at the nominal level $1-\tau$, $$ \lim_{p, T\to\infty} \sup_{\mathbb P\in\mathcal P}\left|\mathbb P\left(\bfm g_{lt}\in CI_{\tau} \right)-(1-\tau)\right|=0 $$ Here the probability measure $\mathbb P$ is taken uniformly over a broad DGP class $\mathcal P$, which admits various cross-sectional variations in $(\bfsym \gamma_{lt}$, $\bfm g_{lt})$ and dynamics if they are also time-varying. While our framework belongs to a more general class of two-step GMM estimators, we encounter a new feature as in the linear factor model we described earlier: the estimated $\bfm g_{lt}$ possesses a discontinuity on its limiting distribution because the strengths of $ \bfm V_{\gamma, t} $ is unknown and can be near the boundary. This brings new challenges to achieving the uniformity in the inference. Uniformity in the above sense is essential in this context, because it makes the inference valid and robust to the unknown degrees of dynamics in $\bfsym \beta_{lt}$, especially, the strengths of the cross-sectional variations in $\bfsym \gamma_{lt}$.
In the presence of “boundary parameters", the asymptotic inference becomes nonstandard, and modified inference procedures have been proposed in both time series and cross-sectional models, e.g., ketz2017testing, ketz2018subvector,pedersen2017inference. The main difference between our model and those considered in the literature is that we do not model $ \bfm V_{\gamma, t} $ explicitly as an unknown “parameter", so it does not appear in the GMM objective function. In fact, the mapping from $\bfsym \gamma_{lt}$ to the limiting distribution of the estimated $\bfm g_{lt} $ is still continuous, which makes the cross-sectional bootstrap first-order valid. We provide more detailed explanations in Section 2.
The linear factor model covers many useful models in the arbitrage pricing theory and models of linkages between international stock markets, as well as macroeconomics. Earlier literature include, e.g., CR, CK93,king1994volatility, SW02,BN02. With the massive use of the newly available datasets of intraday asset prices and large number of cross-sectional data, the continuous-time factor model of high dimensions have received extensive attentions in the recent high-frequency literature, such as Pelger:2016, ait2017using,fan2016incorporating, li2018jump.
The study of the effects of characteristics on betas is an essential subject in financial economics. For instance, it is commonly known that firm sensitivities to risk factors depend on the firm specific raw size and value characteristics. As is noted by daniel1997evidence, “It is the firms' characteristics (size and ratios) rather than the covariance structure of returns that appear to explain the cross sectional variation in stock returns." ang2012testing also found that the market risk premium is less correlated with value stocks' beta (stocks with high book-to-market ratio) than with growth stocks' beta. Firms' momentum is also one of the commonly used characteristics, whose effect on the factor sensitivities has been found to be linearly growing with the momentum, indicating a constant effect. In addition, ferson1999conditioning found that the lagged characteristics track variations in expected returns that is not captured by the Fama-French FF three-factor model, and that these characteristics have explanatory power on the factor loadings because they pick up betas' time-variation. In addition, the effect of common characteristics such as the term spread and default spread demonstrate significantly different volatiles among betas of individual stocks and portfolios, explaining the larger heterogeneity of the factor loadings for the former. Other empirical evidence that systematic risk is related to firm characteristics and business cycle variables is provided by jagannathan1996conditional, lettau2001resurrecting, among many others.
While most of the aforementioned works assume that the betas are fully explained by observed characteristics, a similar decomposition to ((ref)) was given by kelly2017instrumented, where betas are decomposed into a linear function of lagged characteristics as well as an unobservable loading component. They specifically require $\bfsym \gamma_{lt}$ to be “strong", with cross-sectional variances that are bounded away from zero. fan2016projected and kim2018arbitrage respectively studied a model whose betas have a similar decomposition. They did not study the inference problem. Our paper is also related to the continuous-time GMM framework of li2016generalized, but is different on the inference aspect, where we study the continuous-GMM estimation in the presence of high-dimensional linear moment conditions, and when combined with the nonparametric regression, there is a discontinuity issue on the limiting distribution for the characteristic effect. Other literature on continuous-time regression models can be found from barndorff2004econometric,mykland2006anova, li2017adaptive, among others. Also note that we focus on the continuous components of the factor models, and factor model for the jump components was studied by li2018jump.
The rest of this paper is organized as follows. Section (ref) informally discusses the issue of uniformity and the intuitive solutions using the cross-sectional bootstrap. Section (ref) describes the continuous-time conditional factor model driven by stochastic processes. We separately study the known and the unknown factor cases, and present the asymptotic results of the estimators. Section (ref) extends the model to the more general continuous-time GMM framework, with many moment conditions. Section (ref) presents real data applications on the high-frequency stock return data of firms from S&P500. Finally, in the supplement, we give simulated example to examine the uniformity of the proposed inference in finite sample, additional empirical findings, as well as all the technical proofs.
Notation: We observe data every $\Delta_n$ unit of time and let $\Delta_n$ go to zero in the limit. For any process $Z$, let $ \Delta_i^n Z=Z_{i\Delta_n}- Z_{(i-1)\Delta_n} = \int_{(i-1)\Delta_n}^{i\Delta_n} dZ_t$. For simplicity, we will denote $Z_{i\Delta_n}$ by $Z_i$. We use the symbol $\overset{\mathcal{L}\text{-}s}\longrightarrow$ to denote stable convergence in law. We say a constant a absolute constant if it does not depend on any pointwise DGP. Let $\bfm I_d$ be the $d\times d$ dimensional identity matrix. For a matrix $\bfm A$, we use $\lambda_{\min}(\bfm A)$ and $\lambda_{\max}(\bfm A)$ to respectively denote its smallest and largest eigenvalues. In addition, let $\|\bfm A\|:=\lambda_{\max}^{1/2}(\bfm A'\bfm A)$, and $\|\bfm A\|_\infty=\max_{ij}|(\bfm A)_{ij}|$. In addition, we shall achieve inferences uniformly valid over a large class of data generating process $\mathcal P$. For a random sequence $X_n$, we write $X_n\asymp O_P(a_n)$ if $X_n=O_P(a_n)$ and $a_n/X_n=O_P(1)$.
While this paper studies continuous-time models, in this section we heuristically discuss the issue we encounter on the uniform inference using a discrete-time, unconditional factor model with observed factors. Consider the following discrete-time one-factor model:
Here $(y_{mt}, \bfm f_t,\bfm x_m)$ are observable and of finite dimensions, in particular for ease of presentation, $\dim(\bfm f_t)=\dim(\bfsym \gamma_m)=1$. The goal is to make inference about $\ensuremath{\boldsymbol{\theta}}$, which is a $d\times 1$ vector. We assume that $\mathbb E(\bfsym \gamma_m|\bfm X)=0$, $\mathbb E(u_{mt}|\bfm X, \bfm F, \bfsym \Gamma)=0$, where $(\bfm X,\bfsym \Gamma)=\{(\bfm x_m,\bfsym \gamma_m): m\leq p\}$ and $\bfm F=\{\bfm f_t: t\leq T\}$, and that $\{u_{mt},\bfsym \gamma_m\}$ are cross-sectionally independent across $m\leq p.$
A natural estimator for $\ensuremath{\boldsymbol{\theta}}$ is based on a combination of cross-sectional and time-series regression: $$ \widehat \ensuremath{\boldsymbol{\theta}}= \bfm s_f^{-1} \bfm s_x^{-1} \frac{1}{pT}\sum_{t=1}^T \sum_{m=1}^p\bfm x_m\bfm f_t y_{mt} , $$ where $ \bfm s_{f}=\frac{1}{T}\sum_{t=1}^T\bfm f_t^2$, and $\bfm s_x=\frac{1}{p}\sum_{m=1}^p\bfm x_m\bfm x_m'$. Then $\widehat\ensuremath{\boldsymbol{\theta}}$ has the following expansion
Thus the asymptotic distribution depends on the interplay of two leading terms. Term $(\bfm b)$ arises from the cross-sectional estimation, which has a rate $O_P(p^{-1/2}\|\bfm V_{\gamma}\|^{1/2})$, with
where $\operatorname{Var}(\cdot | \bfm X)$ denotes the conditional variance given $\bfm X$. Because $\bfsym \gamma_m$'s are cross-sectionally independent so this term admits a cross-sectional central limit theorem (CLT). In addition, term $(\bfm a)$ has a rate $O_P((Tp)^{-1/2})$ because $u_{mt}$'s are conditionally independent across $(i,t)$.
The first question to address is, what are the final rate of convergence and the limiting distribution? The key to this question is that the strengths of the eigenvalues of $\bfm V_{\gamma}$ are unknown, and are arbitrarily supported on a compact set $[0, C]$ for $C>0.$ If $\bfm V_\gamma$ is weak and near boundary (zero), whose eigenvalues, treated as sequences, decay at rate faster than $O_P(T^{-1})$, then $(\bfm a)$ is the dominating term, leading to, for some covariance $\bfm v_1$, $$ \sqrt{Tp} (\widehat \ensuremath{\boldsymbol{\theta}}-\ensuremath{\boldsymbol{\theta}})\to^d \mathcal N(0, \bfm v_1), $$ whose asymptotic distribution and $\bfm v_1$ are determined by $(\bfm a)$. Intuitively, this occurs when the observed characteristics capture almost all the beta fluctuations, leading to a fast rate of convergence. On the other hand, if $\bfm V_\gamma$ is strong with all eigenvalues bounded away from zero, $(\bfm b)$ becomes the dominating term, and we simply have, for $\bfm v_2:=\text{plim}_{p\to\infty}\bfm V_{\gamma}$, $$ \sqrt{p}(\widehat \ensuremath{\boldsymbol{\theta}}-\ensuremath{\boldsymbol{\theta}}) \to^d \mathcal N(0, \bfm v_2). $$ In this case, the limiting distribution is determined by the cross-sectional CLT of $(\bfm b)$. Intuitively, this means when the idiosyncratic betas have strong cross-sectional variations, time series regression is not helpful to remove their effects on estimating $\ensuremath{\boldsymbol{\theta}}$, and the cross-sectional projection dominates. This leads to a slower rate of convergence.
Consequently, there is a discontinuity on the limiting distribution of $\widehat \ensuremath{\boldsymbol{\theta}}-\ensuremath{\boldsymbol{\theta}}$ when $\bfm V_\gamma$ is near the boundary. In practice, anything in between the above two extreme cases might also happen, leading to an unknown rate of convergence $O_P(a_{pT})$, where $a_{pT}\in[(pT)^{-1/2}, p^{-1/2}]$. This issue is similar to the problems in estimating parameters that are possibly on the boundary of the parameter space andrews1999estimation, andrews2010inference. The problem arises as we do not pretest or know how strong $\bfsym \gamma$'s cross-sectional variation is, which can vary in a large class of data generating process. Most of the financial econometric studies take the “weak" case as the default assumption (e.g., ferson1999conditioning, gagliardini2016time, CMO), while more recent studies (e.g., fan2016projected, kelly2017instrumented, kim2018arbitrage) provide evidence of the presence of the “strong" case in some sampling periods. Above all, to our best knowledge, all the existing inferences are pointwise, and is not robust to the strength of $\bfsym \gamma$'s variations. Pointwise inferences, therefore, can be misleading.
The second question to address is, what is the impact of the unknown rate for $\bfm V_\gamma$ on the inference about $\ensuremath{\boldsymbol{\theta}}$? The “standard" inference procedure is to plug-in the estimated asymptotic covariances for $\bfm V_{\gamma}$ and $\bfm v_1$, using their sample analogues. This procedure, however, works only pointwise, and does not provide a uniformly valid confidence interval. To understand the issue, consider the estimation of $\bfm V_{\gamma}.$ If $\bfsym \gamma_m$ were known, white1980heteroskedasticity's heteroskedastic covariance estimator can be applied:
Replacing $\bfsym \gamma_m$ with its consistent estimator $\widehat\bfsym \gamma_m$, we obtain $\widehat\bfm V_{\gamma}= \bfm s_x^{-1}\frac{1}{p}\sum_{m=1}^p \bfm x_m\bfm x_m' \widehat\bfsym \gamma_m ^2 \bfm s_x^{-1}$. Then $\widehat\bfm V_{\gamma}- \bfm V_{\gamma}$ has a decomposition
where “LLN error" refers to the error associated with the law of large number. The main issue is that the $\gamma$-estimation error cannot be uniformly controlled. Note that in the ideal case where $\ensuremath{\boldsymbol{\theta}}$ were known, one would estimate $\bfsym \gamma_m$ from ((ref)) by running time series regression of $y_{mt}-\bfm x_m'\ensuremath{\boldsymbol{\theta}}\bfm f_t$ on $\bfm f_t$ for each fixed $m\leq p$. Then $\widehat\bfsym \gamma_m-\bfsym \gamma_m=\bfm r_m$, where $\bfm r_m:=\bfm s_{f}^{-1}\frac{1}{T}\sum_{t\leq T}\bfm f_tu_{mt}.$ This leads to $$ \gamma\text{-estimation error} \geq \bfm s_x^{-1}\frac{1}{p}\sum_{m=1}^p\bfm x_m\bfm x_m'\bfm r_m^2 \bfm s_x^{-1}\asymp O_P( T^{-1}). $$ This results in an estimation error $\|\widehat\bfm V_{\gamma}-\bfm V_{\gamma} \|$ being {lower bounded} by an order $ O_P( T^{-1})$, which is not negligible whenever $\lambda_{\min}(\bfm V_{\gamma})=O_P(T^{-1})$ (corresponding to the case of weak $\gamma$-signal).
Consequently, the usual plug-in covariance estimator using $\widehat\bfm V_{\gamma}$ would lead to an asymptotically incorrect distribution, and over-coveraging probabilities. On the other hand, ignoring $\bfm V_{\gamma}$ would result in under-coveraging probabilities when it is present. Hence it is not uniformly valid.\footnote{More precisely, when $\bfm V_\gamma=O_P(T^{-1})$, plugging in its consistent estimator over-estimates the asymptotic variance, leading to valid but severely conservative inferences. }
To resolve the uniformity issue, we propose to use the cross-sectional bootstrap, which is intuitive and very easy to implement. We simply take random samples with replacement across cross-sectional individuals $\{1,..., p\}.$ Once an individual $l^*\in \{1,..., p\}$ is sampled, its associated entire time series $\{y_{l^*,t}\}_{t\leq T}$ is sampled. Then we obtain the estimator $\widehat\ensuremath{\boldsymbol{\theta}}^*$ using the bootstrap data. Finally, we calculate the critical value of $ \widehat \ensuremath{\boldsymbol{\theta}}^*-\widehat\ensuremath{\boldsymbol{\theta}}$ from a set of bootstrap estimators. This procedure is very simple, but perhaps surprisingly, leads to the desired uniform coverage for $\ensuremath{\boldsymbol{\theta}}$.
To prove the bootstrap validity, it is essential to show that this procedure directly mimics the cross-sectional variations in $\{\bfsym \gamma_m\} $. To see this intuitively, we note that we can expand the bootstrap estimator as: for $\bfm R$ as an asymptotically negligible term, $$ \widehat\ensuremath{\boldsymbol{\theta}}^*- \widehat\ensuremath{\boldsymbol{\theta}}= { \bfm s_x^{-1}\frac{1}{p} \sum_{m=1}^p (\bfm x_m ^*\bfsym \gamma_m ^*- \bfm x_m \bfsym \gamma_m ) }+ { \bfm s_f^{-1} \bfm s_x^{-1} \frac{1}{pT}\sum_{t=1}^T\bfm f_t \sum_{m=1}^p( \bfm x_m^* u^*_{it} -\bfm x_m u_{mt} )} +\bfm R, $$ where $\{(\bfm x_m^*, \bfsym \gamma_m^*): i\leq p\}$ is a simple random sample from $\{(\bfm x_m, \bfsym \gamma_m): i\leq p\}$ with replacement. Then the bootstrap asymptotic variance of $\widehat\ensuremath{\boldsymbol{\theta}}^*$ is analogously $ \frac{1}{Tp} \bfm v_1 +\frac{1}{p} \widetilde{ \bfm V}_{\gamma}$, where $\frac{1}{Tp}\bfm v_1$ is the asymptotic variance of term $(\bfm a)$ in ((ref)), and $\widetilde\bfm V_\gamma$ is defined in ((ref)). The only approximation error for $ \bfm V_{\gamma}$ is therefore: $$ \widetilde{ \bfm V}_{\gamma}- \bfm V_{\gamma}=\underbrace{\bfm s_x^{-1}\frac{1}{p}\sum_{m=1}^p\bfm x_m\bfm x_m' \left[\bfsym \gamma_m ^2-\operatorname{Var}(\bfsym \gamma_m|\bfm X)\right]\bfm s_x^{-1}}_{\text{LLN error}}. $$ Consequently, the $\gamma$-estimation error component in ((ref)) is avoided. The LLN error is of a higher order than $\bfm V_{\gamma}$, regardless of the signal strength of $\bfm V_\gamma$. For instance, suppose $\bfsym \gamma_m$ is generated from a rescaled sequence, that is, $\bfsym \gamma_m= b_{T}\bar\bfsym \gamma_m$, where $b_{T}\geq 0$ is a non-random arbitrary sequence, and $\bar\bfsym \gamma_m$ satisfies, for $ C, c_0>0$, almost surely, $$ \lambda_{\min}\left[\frac{1}{p}\sum_{m=1}^p\bfm x_m\bfm x_m'\operatorname{Var}(\bar\bfsym \gamma_m|\bfm X)\right]>c_0,\quad \mathbb E(\|\bar\bfsym \gamma_m\|^4|\bfm X)<C.$$Then the LLN-error $=o_P(1)\bfm V_{\gamma}$. Hence the approximation error for the asymptotic variance of $\widehat\ensuremath{\boldsymbol{\theta}}$ is negligible regardless of the strength $b_{T}$. The bootstrap validity can be achieved.
andrews2000inconsistency gave a generic counter-example showing that the usual bootstrap is inconsistent when the parameter is near the boundary of its space. So modified inference procedures have been proposed, e.g., see more recently, ketz2017testing, ketz2018subvector,pedersen2017inference. We note several important differences between our problem and that of the cited literature. The main difference is that the mapping from the underlying DGP to the asymptotic distribution of $\widehat \ensuremath{\boldsymbol{\theta}}$ is continuous in our setting. Such a mapping is essentially $ \bfsym \Gamma:\to \mathcal M(\bfsym \Gamma)$, $$ \mathcal M(\bfsym \Gamma):=\underbrace{ \bfm s_x^{-1}\frac{1}{p} \sum_{m=1}^p \bfm x_m \bfsym \gamma_m }_{(\bfm b)}, $$ where $\bfsym \Gamma=\{\bfsym \gamma_m\}$. The main reason of achieving a continuous mapping is that we do not specify $\operatorname{Var}(\bfsym \gamma_m|\bfm X)$ or $\bfm V_\gamma$ as unknown parameters. All the parameters in our model, $\{(\ensuremath{\boldsymbol{\theta}}, \bfsym \gamma_m): i\leq p\}$, are inside the interior of their parameter space $\Theta_\theta\times \otimes_{m=1}^p \Theta_\gamma$, where $\Theta_\theta$ and $\Theta_\gamma$ are compact subsets respectively in $\mathbb R^K$ and $\mathbb R$. Therefore when estimating $\ensuremath{\boldsymbol{\theta}}$, the loss function (e.g., least squares) does not explicitly depend on the unknown “boundary parameter" $\operatorname{Var}(\bfsym \gamma_m|\bfm X)$. In the absence of $\operatorname{Var}(\bfsym \gamma_m|\bfm X)$ in the loss function and with the continuous mapping $\mathcal M(\bfsym \Gamma)$, the cross-sectional bootstrap is asymptotically valid. In sharp contrast, in the model of andrews2000inconsistency and ketz2017testing, $\operatorname{Var}(\bfsym \gamma_m|\bfm X)$ is explicitly modeled as an unknown parameter and appears in the loss function. Therefore, the discontinuity in the asymptotic distribution explicitly arises from the presence of the boundary parameter in their loss functions, but is not due to the issue of interplay among multiple terms in the asymptotic expansion. Another reason why the cross-sectional bootstrap is valid in our context is that, as we illustrated in the above, the effect of estimating $\bfsym \gamma_m$ can be avoided by resampling the cross-sectional units, and the bootstrap variance directly estimates the asymptotic variance of term $(\bfm b)$. In contrast, the usual “plug-in" method for the estimated asymptotic variance is not uniformly valid because the estimation error for $\operatorname{Var}(\bfsym \gamma_m|\bfm X)$ dominates the estimand.\footnote{A possible alternative approach is to employ the thresholding: estimate $\bfm V_{\gamma }$ using $\widehat \bfm V_{\gamma}1_{\{\|\widehat \bfm V_{\gamma}\| <c_T \log T\}}$ for some sequence $c_T\asymp \min\{T, \sqrt{p}\}^{-1}$, so that $c_T\log T$ “just dominates" $\|\bfm V_{\gamma} - \widehat\bfm V_{\gamma}\|$. The similar approach has been employed to deal with the distribution discontinuity in the context of random coefficient models, and moment inequalities (e.g., andrews2010inference). But in the current context, it has a few drawbacks. One is that it is hard to cover the entire space of all possible sequences for the eigenvalues of $\bfm V_{\gamma}$. It also leaves a question of choosing the constant in $c_T$. So we do not pursue it in this paper.}
The discussions in this section are based on a very simple setting, assuming that: (1) the model is discrete-time; (2) the betas are time-invariant; (3) factors are directly observable; (4) the parameter of interest is finite dimensional; (5) there are no drifts. In Section (ref) we shall formally explore this idea in a continuous-time conditional factor model with drifts using high-frequency data, and separately consider observed and estimated factors. We also extend the regression model to more general high-dimensional moment condition based models in Section (ref).
Consider a large panel of $p$ time series $\bfm Y_t = (Y_{1t}, \cdots, Y_{pt})$, where $p$ is large. In financial asset pricing applications, $\bfm Y_t $ can be the vector of log-prices of $p$ stocks at time $t$. We assume $\bfm Y=\{\bfm Y_t\}_{t\geq 0}$ is a multivariate It\^{o} semimartingale on a filtered probability space $(\Omega, {\cal F}, \{{\cal F}_t\}_{t\geq0}, \mathbb{P})$. For simplicity, we begin with a model without jumps, as we are interested in the continuous components of log-prices and factors. The jump-robust estimators are given in Section (ref), where we employ a standard procedure to truncate jumps out.\footnote{In addition, we assume there is no micro-structure noises. In empirical studies we use data of five-min frequency. In the presence of micro-structure noises, other solutions include sub-sampling (Zhang&MA:2005), realized kernel (BN-HLN:2008) and pre-averaging (Jacod&LMPV:2009). Our main results remain valid when using those more complicated noise-robust estimators. }
We assume the following (continuous) factor structure:
where $\bfm Y_0$ is the starting value of the process $\bfm Y$ at time 0, the drift process $\bfsym \alpha = \{\bfsym \alpha_s\}_{s\geq0}$ is an optional $\mathbb{R}^p$-valued process, the factor loading process $\bfsym \beta=\{\bfsym \beta_t\}_{t\geq0}$ is an optional $p\times K$ matrix process. The $K$ dimensional continuous factor process $\bfm F_t$ and the idiosyncratic continuous risk $\bfm U_t$ can be represented as
where $\bfm W^U$ and $\bfm W^F$ are two multi-dimensional Brownian motions and are orthogonal to each other (that is, their quadratic covariation is zero), and $\bfsym \alpha^F = \{\bfsym \alpha_s^F\}_{s\geq0}$ is the drift process of the factors. At any time point $t$, we write $\bfsym \beta_t=(\bfsym \beta_{1t},...,\bfsym \beta_{pt})'$ and in general, each $\bfsym \beta_{lt}$ ($l=1,\cdots,p$) is a $ K\times 1$ vector of adapted stochastic processes. In the literature, this beta is referred to as the continuous beta (BLT:2016).
In addition, for each firm $l\leq p$, we observe a set of (possibly) time-varying characteristics:
We allow the characteristics $\bfm X_{lt}$ to consist of (1) common time-varying characteristics $\bfm x_{t}$ (such as term and default spread and macroeconomic variables); (2) individual specific characteristics $\bfm x_{l}$ that are time-invariant over the sampling period $[0,T]$ (such as size and value which change annually); and (3) characteristics $\bfm x_{l,t}$ that are both time-varying and individual specific. Here we present $(\bfm x_{l,t}', \bfm x_l', \bfm x_{t}')'$ with a bit abuse of notation.
In this paper, we consider the following decomposition of the continuous betas:
The overall effect of characteristics on the factor loadings is represented by $\bfm g_{lt}:=\bfm g_{t}(\bfm X_{lt}) $, and is called “characteristic beta". Here $\bfm g_{t}(\cdot) $ is a nonparametric function of macroeconomic and firm variables, possessing less volatile and picks up long run beta fluctuations. On the other hand, $\bfsym \gamma_{lt}$ represents the remaining time-varying individual factor risks after conditioning on the observed characteristics, and captures high frequency movements in betas. The two components capture different aspects of beta dynamics. For the identification purpose, we assume $\operatorname{\mathbb E}(\bfsym \gamma_{lt}| \bfm X_{t})=0$ for $\bfm X_t=\{\bfm X_{lt}: l\leq p\}$, which well separates the characteristic effects from the remaining effects. The goal is to make uniform inference about $\bfm g_{lt}$ over a broad DGP class, which admits various cross-sectional variations in $(\bfsym \gamma_{lt}$, $\bfm g_{lt})$ and dynamics.
We separately study two cases: known and unknown factor cases. In the “known factor case", the high-frequency return data of the common factors are observable, as in the case of ait2014idiosyncratic who constructed Fama-French factors using high-frequency returns. On the other hand, the “unknown factor case" refers to situations in which we do not observe the high-frequency factors, but can estimate them from a large continuous-time panel (up to a locally time-invariant rotation matrix). We shall show that the effect of estimating $\bfm F_t$ leads to an asymptotic bias that needs to be corrected.
The condition $\operatorname{\mathbb E}(\bfsym \gamma_{lt}|\bfm X_t)=0$ serves as a central condition to achieve the identification of the characteristic effects, under which both components in the beta decomposition are well separated. We now discuss the plausibility of this condition and possible approaches to relaxing it. In the presence of omitted characteristics, say $\widetilde\bfsym \gamma_{lt}$, its “explanable component" $\mathbb E(\widetilde\bfsym \gamma_{lt}|\bfm X_t)$ is “absorbed" in $\bfm g_{lt}$, and $\bfsym \gamma_{lt}$ only contains the orthogonal component $\widetilde\bfsym \gamma_{lt}- \mathbb E(\widetilde\bfsym \gamma_{lt}|\bfm X_t)$.
In the absence of this condition, identification is lost, and we need further exogenous variables to identify the effect of characteristics. Consider the ideal case that $\bfsym \beta_{lt}$ were known. Then in ((ref)), $\bfm X_{lt}$ is endogenous. To identify $\bfm g_t(\cdot)$, consider an instrumental variable approach: we need to find an exogenous instrumental variable $\bfm w_{lt}$ so that $\operatorname{\mathbb E}(\bfsym \gamma_{lt}|\bfm w_{lt})=0.$ Define the operator: $$ \mathcal T: \bfm g\to \operatorname{\mathbb E}(\bfm g(\bfm X_{lt})|\bfm w_{lt}). $$ We then have $\mathcal T(\bfm g_t)=\operatorname{\mathbb E}(\bfsym \beta_{lt}|\bfm w_{lt})$. The identification of $\bfm g_t$ depends on the invertibility of $\mathcal T$, and holds if and only if the conditional distribution of $\bfm X_{lt}|\bfm w_{lt}$ is complete, which is an untestable condition (see, e.g., newey2003instrumental). Suppose $\mathcal T$ is indeed invertible, it is well known that estimating $\bfm g_t$ becomes an ill-posed inverse problem, and regularizations are needed, with possibly a very slow rate of convergence. We refer to the literature for related estimation and identification issues: hall2005nonparametric, darolles2011nonparametric, chen2012estimation, etc. Therefore, while relaxing the condition $\operatorname{\mathbb E}(\bfsym \gamma_{lt}|\bfm X_t)=0$ is possible using the nonparametric instrumental variable approach, it requires a very different argument for the identification and estimation. We do not pursue it in this paper.
Suppose there are in total $n$ sample intervals on the interval $[0,T]$, with equal interval length, being $\Delta_n$. For any $t\in[0,T)$, define $I_t^n=\{ \lfloor t/\Delta_n \rfloor +1,\cdots, \lfloor t/\Delta_n \rfloor + k_n \}$, where $ \lfloor \cdot \rfloor$ is the floor (greatest integer) function, and $k_n$ is the number of high frequency observations within the window $I_t^n$. Ignoring the jumps, fix any $t$, and for all $i\in I_t^n$, by the Burkholder-Davis-Grundy inequality (cf. Chapter 2 of Jacod&protter:2011), we have the following approximation for $\Delta_i^n \bfm Y:= \bfm Y_{i\Delta_n} - \bfm Y_{(i-1)\Delta_n}$ (which is a $p\times 1$ vector):
where {$o_P(\Delta_n\sqrt{k_n})$ } holds for each fixed element of $\Delta_i^n\bfm Y$, and is uniform in $i\in I_t^n$. Let $\bfm G_t$ be the $p\times K$ matrix of $\{ \bfm g_{lt} \}_{l=1}^p$ and $\bfsym \Gamma_t$ be the $p\times K$ matrix of $\{\bfsym \gamma_{lt}\}_{l=1}^p$. We shall then use all observations on $I_t^n$ to estimate $\bfm G_{t}$. Let $\dim(\bfsym \beta_{lt})= K$, and $\dim(\bfm X_{lt})=K_x$. Throughout the paper, we shall assume $p, k_n\to\infty$, $\Delta_n\to0$, while $K, K_x$ are fixed constants.
To define the estimators of $\bfsym \Gamma_t$ and $\bfm G_t$, we introduce the following notation. Let $\bfsym \phi_{lt}=(\phi_1(\bfm X_{lt}),...,\phi_J(\bfm X_{lt}))'$ be a $J\times 1$ vector of sieve basis functions of $\bfm X_{lt}$, which can be taken as, e.g., Fourier basis, B-splines, and wavelets. Let $\bfsym \Phi_t=(\bfsym \phi_{1t },...,\bfsym \phi_{pt})'$ be the $p\times J$ basis matrix, and define the projection matrix: $$ \bfm P_t=\bfsym \Phi_t(\bfsym \Phi_t'\bfsym \Phi_t)^{-1}\bfsym \Phi_t',\quad p\times p. $$ We subsequently discuss the estimation procedures for the known and unknown factor cases.
In the known factor case, we also observe $\{\Delta_i^n\bfm F\}_{i\in I_t^n}$ in each interval. We use the following two-step estimation:
Step 1. Run time-series regression:
Write $\widehat\bfsym \beta_{t}=(\widehat\bfsym \beta_{1t},...,\widehat\bfsym \beta_{pt})'$ to be the $p\times K$ matrix.
Step 2. Run cross-sectional regression:
Putting together, the estimators can be expressed as:
When factors are unknown, we first estimate the latent factors and use them in place of $\Delta_i^n\bfm F$ in ((ref)). In the continuous-time factor model literature such as ait2017using and Pelger:2016, these factors are estimated using the regular principal component (PCA) method, extended from connor1986performance. But different from these works, we employ the PCA on the “projected returns", a method that was proposed by fan2016projected in the discrete unconditional factor model. Here we extend the method of this procedure to the continuous-time conditional factor model, and study its asymptotic effect for inference about the characteristic betas.
We use the simplified notation $\bfm P_{i-1}:=\bfm P_{(i-1)\Delta_n}$, $\bfm G_{i-1}:=\bfm G_{(i-1)\Delta_n}$, and $\bfsym \Gamma_{i-1}:=\bfsym \Gamma_{(i-1)\Delta_n}$. On each local window $I_t^n $, we can define the following $p\times k_n$ matrix:
Define the estimated factors, a $k_n\times K$ matrix $$ \widehat{\Delta^n\bfm F}=(\widehat{\Delta_i^n\bfm F}: i\in I_t^n)', \quad k_n\times K, $$ whose columns equal $\sqrt{k_n\Delta_n}$ times the eigenvectors of the $k_n\times k_n$ matrix $\frac{1}{pk_n\Delta_n}(\bfm P\Delta^n\bfm Y)_t'(\bfm P\Delta^n\bfm Y)_t$ corresponding to the first $K$ eigenvalues. We then use estimated factors in place of $\{\Delta_i^n\bfm F\}_{i\in I_t^n}$ in ((ref)):
and note that $\frac{1}{k_n\Delta_n}\sum_{i\in I_t^n} \widehat{\Delta_i^n\bfm F} \, \widehat{\Delta_i^n\bfm F}' =\bfm I_K$. The $l$-th components $\widehat\bfm g_{lt}^{\operatorname{latent}} $ and $\widehat\bfsym \gamma_{lt}^{\operatorname{latent}} $, respectively estimate the characteristic and idiosyncratic betas for the $l$-th individual. The superscript “latent" indicates that the estimators are defined for the case of latent factors.
We now give an intuitive explanation on the estimated factors. Note that $\bfm P_{i-1}$ is the cross-sectional projection matrix onto the space expanded by the sieve transformations of the characteristics at time $(i-1)\Delta_n$. Apply the projection to the discretized model:
where $\ensuremath{\boldsymbol{\psi}}_{i-1}$ denotes the higher-order drift and diffusion terms. By the identification conditions $\mathbb{E}(\bfsym \Gamma_t|\bfm X_t)=0$ and that $\mathbb{E}( \bfm U_{t+s} - \bfm U_t | \bfm X_t)=0$, the two components of the “projection errors" are projected off, whose rate of decay (after standardized by $\Delta_n^{-1/2}$) is of $O_P(p^{-1/2})$. Ignoring the higher-order term, we have
which is nearly “idiosyncratic-free", and leads to $$ (\bfm P\Delta^n\bfm Y)_t'(\bfm P\Delta^n\bfm Y)_t\approx\Delta^n\bfm F \bfm G_t'\bfm G_t\Delta^n\bfm F'. $$ Therefore the columns of $\Delta^n\bfm F$ are approximately the eigenvectors of the “idiosyncratic-free" matrix $(\bfm P\Delta^n\bfm Y)_t'(\bfm P\Delta^n\bfm Y)_t$, up to a rotation. Hence we can estimate them by applying PCA on $(\bfm P\Delta^n\bfm Y)_t'(\bfm P\Delta^n\bfm Y)_t$.
Furthermore, ((ref)) also yields:
It shows that $\bfsym \Gamma_{i-1}$ represents the loading on the risk factors of the remaining components of returns, after the characteristic effect is conditioned.
In the general case with jumps, we employ the truncation method to remove those jumps. For notation simplicity, we omit the details and simply assume the jumps are of finite variation. In the known factor case, we replace each $\Delta_i^n \bfm Y$ and $\Delta_i^n \bfm F$ (previously assumed to be continuous) with their truncated versions:
where $\Delta_i^n Z_{\psi_n^Z}:=\Delta_i^n Z_l \, 1_{\{ \| \Delta_i^n Z_l \| \leq \psi_n^{Z_l} \}}$ denotes the usual truncated process for the process $\Delta_i^nZ$, with some random sequence $\psi_n^Z$ that depends on certain property of $Z$ and converges in probability to zero as $\Delta_n\rightarrow0$ (e.g., Mancini:2001).\footnote{The common practice is the set $\psi_n^{Z_l} = \alpha_l \Delta_n^\varpi$, where $\varpi\in(0,1/2)$, $\alpha_l=C (\frac{1}{t} \text{IV}(Z_l)_t)^{1/2}$ with $C=3,4$ or $5$ and $\text{IV}(Z_l)_t$ is the integrated volatility of $Z_l$ over $[0,t]$.} In the unknown factor case, we only need to replace each $\Delta_i^n \bfm Y$ with its corresponding truncated versions.
We now present the technical assumptions for the asymptotic properties of the estimated characteristic effect $\widehat\bfm g_{lt}$ for a fixed $l\leq p$. This subsection presents the required conditions to prove the limiting distribution. In particular, we allow the idiosyncratic components $\{\bfm U_t, \bfsym \Gamma_t\} $ to be cross-sectionally weakly dependent. In Section (ref), we shall present the required conditions for the cross-sectional bootstrap, where we assume them to be cross-sectionally independent.
We assume that the following conditions hold uniformly over a class of DPG's: $\mathbb P\in\mathcal P$. We apply the standard assumptions to define the stochastic processes as follows (e.g., Protter:2005).
The above assumption ensures that $\bfm g_{t}(\bfm X_{lt})$ and $\bfsym \phi_{lt}$ are smooth transformations of $\bfm X_{lt}$. In particular, condition (i) is regarding the smoothness with respect to time, so $\bfm g_{t}(\bfm X_{lt})$ and $\bfsym \phi_{lt}$ are also semimartingales; Condition (ii) is regarding the smoothness with respect to cross-sections. It holds if $\max_{t, \bfm x}\|\bfm g_t(\bfm x)-\sum_{j=1}^J\bfm b_j\phi_j(\bfm x)\| \leq CJ^{-\eta}$ for some sieve coefficients $\{\bfm b_j: j\leq J\}$.
We now describe the asymptotic variance of $\widehat \bfm g_{lt}$, and introduce further notation. Let $\bfm c_{FF,t}$ ($K\times K$) and $\bfm c_{uu,t}$ ($p\times p$) be the instantaneous quadratic variation processes of $\bfm F=\{\bfm F_t\}_{t\geq0}$ and $\bfm U=\{\bfm U_t\}_{t\geq0}$. Let $U_{l,t}$ denote the $l$-th component of $\bfm U_t$, where $l\leq p$. Also, let $$ h_{t,lm} =\bfsym \phi_{lt}'(\frac{1}{p}\bfsym \Phi_t'\bfsym \Phi_t)^{-1}\bfsym \phi_{mt},\quad l,m\leq p. $$ The asymptotic variance depends on the following matrices.
The asymptotic distribution of $\widehat\bfm g_{lt}$ is jointly determined by two uncorrelated components:
where $\bfm P_{t,l} $ denotes the $l$-th column of $\bfm P_t$. Assumptions (ref) condition (ii) requires that $\bfsym \gamma_{lt}$ be cross-sectionally weakly dependent, and $ \bfsym \Gamma_{t}' \bfm P_{t,l} $ admits a cross-sectional CLT. In addition, condition (iii) requires $\{U_{mt}: m\leq p\}$ be cross-sectionally weakly dependent. These conditions ensure the asymptotic normality of ((ref)).
Assumption (ref) is similar to the pervasive condition in the approximate factor model's literature, which identifies the latent factors (up to a rotation). In particular, condition (i) requires that the betas should nontrivially depend on $\bfm X_{t}$.
We first present the estimated $\bfm g_{lt}$ for a fixed $(l,t)$ when factors are observable.
When the factors are latent and estimated, $\widehat \bfm g_{lt}$ consistently estimates a rotated $\bfm g_{lt}$. Up to the rotation, the asymptotic variance is identical to that of the known factor case. However, the effect of estimating the factors gives rise to a bias term. Let $\widehat \bfm V_t$ be a $k_n\times k_n$ diagonal matrix consisting of the first $K$ eigenvalues of $\frac{1}{pk_n\Delta_n}(\bfm P\Delta^n\bfm Y)_t'(\bfm P\Delta^n\bfm Y)_t$. Let
We have the following theorem.
We make several remarks.
Generally, while $\|\bfm V_{\gamma,t}\|=o_P(k_n^{-1}) $ and $\lambda_{\min}(\bfm V_{\gamma, t})\gg k_n^{-1}$ are two special cases, we do not know the actual strength of the eigenvalues of $\bfm V_{\gamma,t}$ in practice. In fact, its eigenvalues can be any sequences in a large range, resulting in an unknown rate of convergence for $ \left(\widehat \bfm g_{lt}-\bfm g_{lt}\right)$ between $O_P((k_np)^{-1/2})$ and $O_P(p^{-1/2})$. This calls for a need of uniform inference. We shall rely on the cross-sectional bootstrap, as we formally present in the next subsection.
We now derive a bias-corrected spot estimated $\bfm g_{lt}$ in the case of estimated factors. The bias correction is valid uniformly over various signal strengths. Note that in ${\rm BIAS}_g$, $\bfm M_t$ can be naturally estimated by $\widehat\bfm M_t=\frac{1}{ \sqrt{p}} \widehat\bfm V_t^{-1} \widehat \bfm G_t'.$ The major challenge arises in estimating the $p\times p$ quadratic variation $\bfm c_{uu, t}$, which is high-dimensional when $p$ is large. We consider two cases for the bias correction.
CASE I: cross-sectionally uncorrelated
When $\{\Delta_i^n U_1,\cdots,\Delta_i^n U_p\}$ are cross-sectionally uncorrelated, the ${\cal F}_{i-1}$ conditional quadratic variation $\bfm c_{uu,(i-1)\Delta_n}$ is a diagonal matrix. Let $ \widehat{\Delta_{i}^n\bfm U}=\Delta_{i}^n\bfm Y-(\widehat\bfm G_{i-1}+\widehat\bfsym \Gamma_{i-1})\Delta_{i}^n\bfm F$. Apply white1980heteroskedasticity's covariance estimator using the residuals: $$ \widehat{{\rm BIAS}}_g=\widehat \bfm M_t \frac{1}{k_n\Delta_n\sqrt{p}} \sum_{i\in I_t^n} \bfm P_{i-1} \operatorname{diag}\{\widehat{\Delta_{i}^n\bfm U} \widehat{\Delta_{i}^n\bfm U}'\} \bfm P_{i-1,l} $$
CASE II: cross-sectionally weakly correlated (sparse)
In this case $\bfm c_{uu, (i-1)\Delta_n}$ is no longer diagonal. We shall assume it is a sparse covariance matrix, in the sense that many of its off-diagonal entries are zero or nearly so. Then the thresholding estimator of POET can be applied, yielding a nearly $\min\{k_n, p\}^{1/2}$- consistent sparse covariance estimator $\widehat{\bfm c}_{uu,t}$. More specifically, let $s_{dl}$ be the $(d,l)$-th element of $\frac{1}{\Delta_nk_n}\sum_{i\in I_t^n}\widehat{\Delta_i^n\bfm U}\widehat{\Delta_i^n\bfm U}'$. Let the $(d,l)$-th entry of the estimated covariance be: $$ (\widehat \bfm c_{uu,t})_{dl}=
$$ where $\text{th}(\cdot)$ is a thresholding function, whose typical choices are the hard-thresholding and soft-thresholding. Here the threshold value $\varrho_{dl}= \bar C (s_{dd}s_{ll})^{1/2}\omega_{np}$ for some constant $\bar C>0$, with $\omega_{np}=\sqrt{\frac{\log p}{k_n}} +\frac{J}{p}\max_{j,d}\frac{1}{J}\|\bfsym \phi(\bfm x_{jd})\|^2\sqrt{\log J}$.\footnote{Hard-threholding takes $\text{th}(s_{dl})=s_{dl}$, while soft-thresholding takes $\text{th}(s_{dl})=\mbox{sgn}(s_{dl})( |s_{dl}|-\varrho_{dl} )$. We shall justify the choice of $\omega_{np}$ in Section \ref{sec:a.6}. In addition, the choice of the constant $\bar C$ can be either guided using cross-validation, or simply a constant near one. For returns of S\&P 500, the rule of thumb choice $\bar C=0.5$ empirically works very well. We refer to \cite{POET} for more discussions on thresholding.} Estimate the bias by: $$ \widehat{{\rm BIAS}}_g=\widehat \bfm M_t\frac{1}{k_n\sqrt{p}}\sum_{i\in I_t^n} \bfm P_{i-1} \, \widehat{\bfm c}_{uu, t} \, \bfm P_{i-1,l}. $$
Formally, we have the following theorem for the bias-corrected estimator.
As we explained in Section (ref), the major difficulty is to handle $\bfsym \Gamma_t'\bfm P_{t,l}=\frac{1}{p}\sum_{m=1}^p \bfsym \gamma_{mt}h_{t,ml}$ in the asymptotic expansion of $\widehat\bfm g_{lt}-\bfm g_{lt}$, which contributes to the asymptotic variance by $\frac{1}{p}\bfm V_{\gamma, t}$. If we use the standard “plug-in" method to estimate $\bfm V_{\gamma, t}$, we would have to introduce a $\gamma$-estimation error $$ \frac{1}{p}\sum_{m=1}^p h_{t,ml}^2\left[ \widehat \bfsym \gamma_{mt}\widehat \bfsym \gamma_{mt}' - \bfsym \gamma_{mt} \bfsym \gamma_{mt}' \right], $$ which is not negligible when $\bfm V_{\gamma, t}$ is near zero. Instead, we propose to use cross-sectional bootstrap to achieve uniform inference, and let the bootstrap distribution “mimic" cross-sectional variations.
Let $\mathcal M:=\{m_1,..., m_p\}$ be a simple random sample with replacement from $\{1,..., p\}$, and we always fix $m_l= l$ when we are interested $\bfm g_{lt}$ for the $l$-th specific individual. As we do not need to mimic the time series variations, so for each sampled index $m\in \mathcal M$, the entire time series $\{\Delta_i^nY_{m}, i\in I_t^n \}$ and $\{\bfm X_{m, i}, i\in I_t^n\}$ are kept. Therefore, we independently resample the time series and obtain the bootstrap data: $\{\Delta_i^nY_m, i\in I_t^n \}_{m\in \mathcal M}$ and $ \big\{\bfm X_{mi}: i\in I_t^n \big\}_{m\in\mathcal M}$. In addition, we also keep the entire time series $\{\Delta_i^n\bfm F \}$ in the case of known factors, and $\{\widehat{\Delta_i^n\bfm F} \}$ in the case of unknown factors. The effect of estimating $\Delta_i^n\bfm F$ does not play a role in the cross-sectional variations. Hence we do not re-estimate the factors in each bootstrapped sample.
Let $\Delta_i^n\bfm Y^*=(\Delta_i^nY_{m_1},...,\Delta_i^nY_{m_p})'$, $\bfsym \Phi_t^*=(\bfsym \phi_{m_1,t},...,\bfsym \phi_{m_p,t})'$ and $\bfm P_t^*=\bfsym \Phi_t^*(\bfsym \Phi_t^{*'}\bfsym \Phi_t^*)^{-1}\bfsym \Phi_t^{*'}$. Let $\bfm P_{t,l}^*$ be the $l$-th column of $\bfm P_t^*$. Define
When $\bfm g_{lt}$ is multidimensional, it is easier to present the confidence interval for a linear transformation $\bfm v'\bfm g_{lt}$. The following algorithm summarizes the steps for computing the confidence intervals.
We need the following conditions for the bootstrap validity.
Note that the bootstrap asymptotic variance for $\widehat \bfm g_{lt}^*$ is analogously $ \frac{1}{k_np} \bfm V_{u, t} +\frac{1}{p} \widetilde{ \bfm V}_{\gamma, t}$, where $ \widetilde{\bfm V}_{\gamma, t} =\frac{1}{p}\sum_{m=1}^ph_{i,ml}^2 \bfsym \gamma_{mt} \bfsym \gamma_{mt}'$. The only approximation error for the $ \bfm V_{\gamma, t}$ part is the “law of large numbers error": $$ \widetilde{ \bfm V}_{\gamma, t}- \bfm V_{\gamma, t}=\underbrace{\frac{1}{p}\sum_{m=1}^ph_{t ,ml}^2 \left[\bfsym \gamma_{mt} \bfsym \gamma_{mt} ' -\operatorname{Var}(\bfsym \gamma_{mt}|\bfm X_t)\right]}_{\text{LLN error}}. $$ Consequently, the $\gamma$-estimation error component is avoided. This forms the foundation of the bootstrap asymptotic validity. The moment bound in Assumption (ref) (i) on $\bfsym \Gamma_t$ is used to ensure that the LLN-error is negligible regardless of the strength of $\bfm V_{\gamma,t}$.
Theorem (ref) requires the idiosyncratic components $\{\{\Delta_i^nU_m\}_{t\in [0, T]},\{\bfsym \gamma_{mt}\}_{t\in [0, T]}\}_{m\leq p}$ be cross-sectionally independent. We can relax this assumption to allow for cross-sectionally block-dependent idiosyncratic components when these blocks are known, and rely on block cross-sectional bootstraps. More specifically, suppose the cross-sectional index set has a non-overlapping partition $\{1,..., p\}=B_1\cup...\cup B_H$, with $H\to\infty$, and the cardinality of each “block" $B_h$ is finite: $\max_{h\leq H}|B_h|_0=O(1).$ For a fixed $k\leq K$. We assume: \\ (i) Individuals' block mememberships are known; (that is, $\{B_h\}_{h\leq H}$ are known)\footnote{In the more complicated case where blocks are unknown, one could first apply a block-thresholding method to estimate the cross-sectional covariance matrix of $\Delta_t^n\bfm U$, to consistently recover the block structures first. See, e.g., cai2012adaptive.}, \\ (ii) $\mbox{Cov}(\Delta_t^nU_{l_1}, \Delta_t^nU_{l_2})=0$ and $\mbox{Cov}(\gamma_{mt, k_1}, \gamma_{mt, k_2})=0$ for any $t\in[0,T]$ if the two individuals $(l_1,l_2)$ belong to different blocks, for any $k_1, k_2\leq K$; $\gamma_{mt,k}$ denotes the $k$-th element of $\bfsym \gamma_{mt}$.
Therefore, conditionally on the factors, individuals are possibly correlated only within the same block. Empirically, the assumption of known blocks of finite size can be supported by setting blocks as industry sectors, which is motivated from the economic intuition that firms within similar industries are expected to have higher correlations conditioning on the factors, e.g., ait2017using. These blocks form a natural basis for the application of non-overlapping block bootstraps. Suppose we are interested in the inference for $\bfm g_{lt}$ for a specific firm $l\leq p$, and it is known that $l\in B_{h_0}$ for a particular $h_0\leq H$. Set $l$ as the first element of $B_{h_0}$. We employ the block-bootstrap on the cross-sectional units, and can proceed with the following algorithm.
The proof of the first-order bootstrap validity is very similar to that of Theorem (ref), building on the results of the validity of block-bootstrap andrews2004block, lahiri1999theoretical. We omit the formal proof for technical simplicity.
We can also estimate the long-run characteristic effect: $\int_0^T\bfm g_{lt}dt$ in the case of known factors.\footnote{Due to the rotation discrepancy, estimating the long-run effect with unknown factors subjects to the issue of time-varying rotations, and is a very challenging problem in the presence of time-varying betas, and we shall leave it for the future research.} Using the standard overlapping spot estimates (see, e.g. Jacod&Rosenbaum:2013), we estimate it by $$\widehat{\int_0^T\bfm g_{lt}dt}:=\sum_{t=1}^{[T/\Delta_n]-k_n}\widehat \bfm g_{lt}\Delta_n.$$ Asymptotically, it has the following decomposition:
As in the spot estimation case, the asymptotic expansion also admits two components: the effect of estimating integrated betas and the effect of cross-sectional estimation. The final limiting distribution is determined by the interplay of both terms. Due to the unknown signal strength of the idiosyncratic beta $ \sum_{t=1}^{[T/\Delta_n]-k_n}\Delta_n \bfsym \Gamma_t'\bfm P_{t,l}$, we still rely on the bootstrap to make uniform inference about $\bfm v'\int_0^T\bfm g_{lt}dt$ for any specific vector $\bfm v$ of interest.
Denote by $\widehat{\int_0^T\bfm g_{lt}dt}^{*b}=\sum_{t=1}^{[T/\Delta_n]-k_n}\widehat \bfm g_{lt}^{*b}\Delta_n$ as the bootstrap estimator in the $b$-th generated sample. Let $\widetilde q_{\tau}$ be the $1-\tau$ bootstrap quantile of $\{ | \bfm v'\widehat{\int_0^T\bfm g_{lt}dt}^{*b} - \bfm v'\widehat{\int_0^T\bfm g_{lt}dt} | \}_{b\leq B}$. The confidence interval for $\bfm v'\int_0^T\bfm g_{lt}dt$ is given by
The following condition plays a similar role as that of Assumption (ref), but for estimating the integrated characteristic effects.
We consider a more general continuous-time model with linear moment conditions. Suppose we observe data that are discretized realizations from a continuous-time stochastic process $\{\bfm Z_t=(\bfm Z_{1t},...,\bfm Z_{pt}): t\in[0,T]\}$. We assume $\bfm Z=\{\bfm Z_t\}_{t\geq 0}$ be a multivariate It\^{o} semimartingale on a filtered probability space $(\Omega, {\cal F}, \{{\cal F}_t\}_{t\geq0}, \mathbb{P})$. Igonoring jumps, we have $$ \bfm Z_{lt} =\int_0^t \bfsym \alpha_{ls}^Z ds + \int_0^t \bfsym \sigma_{ls}^Z d \bfm W^Z_{ls}, \quad l=1,..., p,\quad \forall t\in[0,T] $$ where $\{ \bfm W_{lt}^Z\}_{t\leq 0}$ is a Brownian motion, and $ \{ \bfsym \alpha_{lt}^Z\}_{t\geq0}$ is the drift process. For each $t\in[0, T],$ let the parameter $ \bfsym \beta_{lt} $ be identified by the following moment condition:
where $\bfm c_{z,lt}=d[\bfm Z_l,\bfm Z_l]_t/dt$ is the instantaneous quadratic variation process of $\bfm Z_l=\{\bfm Z_{lt}\}_{t\geq 0}$, and $ \Psi_l(\bfsym \beta,\bfm c)$ is a known function linear with respect to $\bfsym \beta$ and continuously differentiable with respect to $\bfm c$. In addition, we assume that $\bfsym \beta_{lt}$ depends on a set of characteristics $\bfm X_{lt}$, so has the following decomposition:
where $\mathbb E(\bfsym \gamma_{lt}|\bfm X_{t})=0$. Here $\bfm X_{t}$ a set of (possibly) time-varying characteristics. The effect of characteristics on $\bfsym \beta_{lt}$ is represented by $ \bfm g_{t}(\bfm X_{lt}) $. The goal is to make uniform inference about $ \bfm g_{t}(\bfm X_{lt}) $ that allows the cross-sectional variations of $\bfsym \gamma_{lt}$ to be possibly arbitrarily close to the “boundary".
Many applications in economics and finance give rise to the continuous-time many linear moment conditions as in ((ref)) and ((ref)), as we now illustrate with a few examples.
We apply the continuous-time generalized methods of moments (GMM) based on the following moment conditions:
We shall first assume the underlying $\bfm Z_{lt}$ is continuous. Let $\dim(\bfm Z_{lt})= K_z$ and $\dim(\Psi_l)=K_\psi$. We shall assume $ K_z $ and $K_\psi$ are fixed constants, and $K_\psi\geq K$, so $\bfsym \beta_{lt}$ is possibly over-identified.
Ignoring the jumps, over the $i$-th sampling interval, we have the following discrete time observation:
Let $\widehat \bfm c_{z, lt}$ be the sample quadratic variation of $\bfm Z_{lt}$ on window $I_t^n$: $$ \widehat \bfm c_{z, lt}=\frac{1}{k_n\Delta_n}\sum_{i\in I_t^n} \Delta_i^n\bfm Z_l \, \Delta_i^n\bfm Z_l'. $$ We employ a simple two-step estimation:
Step 1. Let $\widehat\bfsym \beta_{lt}$ be the solution that satisfies:
Let $\widehat\bfsym \beta_{t}=(\widehat\bfsym \beta_{1t},...,\widehat\bfsym \beta_{pt})'$ be the $p\times K$ matrix. Here $\ensuremath{\boldsymbol{\Omega}}_{lt}$ is a known positive definite weight matrix. In the general case with jumps, we solve $$ \widehat\bfsym \beta_{lt}:=\arg\min_{\bfsym \beta}\Psi_l(\bfsym \beta, \widehat\bfm c_{z, lt}^\psi)' \ensuremath{\boldsymbol{\Omega}}_{lt} \Psi_l(\bfsym \beta, \widehat\bfm c_{z, lt}^\psi) ,\quad l=1,..., p. $$ where we replace each $\Delta_i^n \bfm Z_l$ with their truncated versions: $ \widehat \bfm c_{z, lt}^{\psi}=\frac{1}{k_n}\sum_{i\in I_t^n} \Delta_i^n\bfm Z_l^{\psi_n} \, \Delta_i^n\bfm Z_l^{\psi_n\prime}. $ Here $\Delta_i^n \bfm Z_{l}^{\psi_n}:=\Delta_i^n \bfm Z_l \, 1_{\{ \| \Delta_i^n \bfm Z_l \| \leq \psi_n \}}$ denotes the usual truncated process for the process $\Delta_i^n \bfm Z_{l}$, with some random sequence $\psi_n$ that converges in probability to zero as $\Delta_n\rightarrow0$.
Step 2. Define
Therefore, the estimator ((ref)) in the linear factor model with known factors is a special case of the two-step estimator presented here, while the estimator of the unknown factor case can also be considered as a special case, which replaces $\Delta_i^n\bfm Z_l$ in the definition of $\widehat\bfm c_{z,lt}$ with the “estimated regressors".\footnote{ The general asymptotic theories presented in this section focus on the “known regressor case". In the case of “estimated regressors", corresponding to the latent factor case in the linear factor model, the estimated regressors should be estimated separately and plugged in, as we did for the unknown factors in Section 3. The effect of estimating the unknown regressors might introduce additional biases that need be corrected. For simplicity, we do not cover this case in the general GMM framework.}
Our estimator belongs to the general class of two-step GMM estimators, as previously discussed in newey1994asymptotic, chen2003estimation,chen2015sieve,chernozhukov2016double: a nuisance parameter is estimated in the first step, and substituted in the second step estimation. As is shown by these authors, when the second set of moment conditions is not “Neyman orthogonal" with respect to $\bfsym \beta_{lt}$ (roughly speaking, its “directional derivative" with respect to $\bfsym \beta_{lt}$ is nonzero), the first-step estimation error $\widehat\bfsym \beta_{lt}-\bfsym \beta_{lt}$ is not negligible, and plays a leading role in the asymptotic distribution for $\widehat\bfm g_{lt}$.
In the current large panel context with many moment conditions ($l=1,..., p$), indeed the second moment condition for $\bfm g_{lt}$, given by $\mathbb E(\bfsym \beta_{lt}|\bfm X_t)=\bfm g_{lt}$, is not Neyman orthogonal with respect to $\bfsym \beta_{lt}$. However, two new phenomena are present here. To see this, let $\nabla_\beta\Psi_m( \bfm c )$ be the $K\times K$ gradient matrix of $\Psi_m(\bfsym \beta, \bfm c)$ with respect to $\bfsym \beta$, which does not depend on $\bfsym \beta$ due to the linearity. Write
The first-order condition of ((ref)) leads to, the $p\times K$ matrix $\widehat\bfsym \beta_t$ satisfies: $$ \widehat\bfsym \beta_t-\bfsym \beta_t\approx-
$$ where $\bfsym \beta_{lt}$ and $\bfm c_{z, lt}$ are evaluated at the true values. Then applying Step 2 of the estimation, combined with the delta-method, yields, for each $l\leq p$, $$ \widehat\bfm g_{lt}-\bfm g_{lt} = -\underbrace{\frac{1}{p}\sum_{m=1}^p \bfm A_{mt}(\bfm c_{z,mt})\nabla_c\Psi_m(\bfm c_{z, mt})vec(\widehat\bfm c_{z,mt}- \bfm c_{z,mt})h_{t,ml}}_{first-step effect: (\bfm a)}+\underbrace{\bfsym \Gamma_t'\bfm P_{t,l}}_{second-step effect: (\bfm b)}+ o_P(\frac{1}{\sqrt{k_np}}) $$ where $\nabla_c\Psi_m( \bfm c_{z, mt})$ is the gradient of $\Psi_m(\bfsym \beta_{mt},\bfm c_{z, mt})$ with respect to $vec(\bfm c_{z,mt})$.
The first new phenomenon is that the first-step effect ($\bfm a$) is a cross-sectional average of the estimation errors for $\{\widehat\bfsym \beta_{mt}-\bfsym \beta_{mt}: m=1,..., p\}$, whose rate of convergence is $O_P((k_np)^{-1/2})$ and is in fact negligible when $\operatorname{Var}(\bfsym \Gamma_t|\bfm X_{lt})$ is bounded away from zero, even though the second moment condition is not Neyman orthogonal. This is essentially due to the fact that in the presence of many moment conditions, the first-step effect can be “cross-sectionally averaged out" and thus can be dominated by the second-step effects when the latter has a slower rate of convergence.
The second new phenomenon, which is also the unique feature in our model, is that whether the second-step moment condition $\bfm E(\bfsym \beta_{lt}|\bfm X_{t})= \bfm g_{lt}$ is “noise-free" is unknown, in the sense that it is possible that the magnitude of the eigenvalues of $\operatorname{Var}(\bfsym \Gamma_t|\bfm X_{lt}) $ can be arbitrary in their parameter space $[0, C]$, and may vary across time. Especially, when $\operatorname{Var}(\bfsym \Gamma_t|\bfm X_{t}) $ is near the boundary of the parameter space (zero), which occurs when $\bfm X_{t}$ has nearly full explanatory power on $\bfsym \beta_{lt}$, the first-step effect becomes the only leading term in the asymptotic expansion.
The aforementioned two unique features of our model call for a different asymptotic analysis for the two-step estimator considered here, and similar to the discussions in the linear factor model, lead to an unknown rate of convergence and a discontinuity in the limiting distribution.
To formally present our theory, we assume that the following conditions hold uniformly over a class of DPG's: $\mathbb P\in\mathcal P$.
Next, for each $m\leq p$, $\Psi_m(\bfsym \beta, \bfm c)$ is linear in $\bfsym \beta$, so $\nabla_\beta\Psi(\bfm c)$ does not depend on $\bfsym \beta$. When $\bfsym \beta_{mt}$ is evaluated at the true value, we simply write $ \Psi_m( \bfm c ):= \Psi_m(\bfsym \beta_{mt}, \bfm c )$ and $\nabla_c\Psi_m( \bfm c ):=\nabla_c\Psi_m(\bfsym \beta_{mt}, \bfm c )$ to explicitly make them as functions of $\bfm c$. Here $\nabla_\beta,\nabla_c$ respectively denote the gradient operators with respect to $\bfsym \beta$ and $\mbox{vec}(\bfm c)$.
Assumption (ref) is a high-level assumption on the effect of estimating a large number of quadratic variations. It requires two fundamental conditions. The first line requires that the moment function $\Psi_m( \bfm c)$ should be well approximated by linear functions locally; the second line additionally requires the smoothness of $ \bfm A_{mt}(\bfm c)$. A loose upper bound using $\frac{1}{p}\sum_{m=1}^p \|\widehat\bfm c_{z,mt}-\bfm c_{z,mt}\|^2=O_P(k_n^{-1})$ based on the Cauchy-Schwarz inequality can be achieved, so both conditions can be simply verified so long as $p=o(k_n)$. On the other hand, sharper upper bounds can also be achieved to allow for a much larger $p$, by noting that both conditions are regarding high-dimensional cross-sectional averages of the effects of $\widehat \bfm c_{z,mt}- \bfm c_{z,mt}$, and intuitively, these should be “averaged out" as $p\to\infty, $ and can be verified case-by-case for specific models. For instance, powerenhancement and bai2017inferences provide technical arguments to verify (ii) when $\Psi_m(\bfm c)$ is linear in $\bfm c^{-1}$. In fact, when proving Theorem (ref) in the linear factor model, we verify these conditions directly and apply our general result theorem (ref).
We now describe the asymptotic variance of $\widehat \bfm g_{lt}$. For $m=1,..., p$, let $$ \ensuremath{\boldsymbol{\xi}}_{mt,s}:=\bfm A_{mt}(\bfm c_{z,mt})\nabla_c\Psi_m( \bfm c_{z, mt})\mbox{vec}\left(\frac{1}{s}(\bfm Z_{m,t+s}-\bfm Z_{mt}) (\bfm Z_{m,t+s}-\bfm Z_{mt})'-\bfm c_{z,mt } \right). $$
The asymptotic distribution of $\widehat\bfm g_{lt}$ is jointly determined by $\ensuremath{\boldsymbol{\xi}}_{mt,s}$ and $\bfsym \gamma_{mt}$. Let
Generally, we have the following theorem.
Therefore, when the strength of $\bfm V_{\gamma, t}$ is unknown, the general two-step GMM estimation in the current context yields an unknown rate of convergence $O_P(a_{np})$, where $a_{np}$ may vary on the range $[(pk_n)^{-1/2}, p^{-1/2}]$. This leads to a discontinuity on its limiting distribution.
We extend the cross-sectional bootstrap described in Section 3 to the more general context, in order to achieve the uniform inference. Importantly, note that the main issue that causes the unknown limiting distribution is from $ {\bfsym \Gamma_t'\bfm P_{t,l}} $ in the expansion of $\widehat\bfm g_{lt}-\bfm g_{lt}$. This term comes from the second-step estimation. Hence the first-step regressions for $\bfsym \beta_{t}$, which only depends on time-series estimations, is not needed to be repeated in the bootstrap steps.
As such, we directly resample $\{ (\widehat\bfsym \beta_{mt}^*, \bfm X_{mt}^*): m=1,... , p \}$ with replacement, where
and $\widehat\bfsym \beta_{mt}$ is the estimated $\bfsym \beta_{mt}$ in step 1. Here $\{m_1,..., m_p\}$ is a simple random sample with replacement from $\{1,..., p\}$. We then let $\bfsym \Phi_t^*=(\bfsym \phi_{m_1,t},...,\bfsym \phi_{m_p,t})'$, $\bfm P_t^*=\bfsym \Phi_t^*(\bfsym \Phi_t^{*'}\bfsym \Phi_t^*)^{-1}\bfsym \Phi_t^{*'}$, and $\widehat\bfsym \beta_t^*=(\widehat\bfsym \beta_{m_1, t}^*,..., \widehat\bfsym \beta_{m_p, t}^*)'. $ When we are interested in $\bfm g_{lt}$ for given $l$, we always fix the index of the $l$-th resampled cross-sectional unit being $l$, that is, $m_l=l$. Hence $\widehat\bfsym \beta_{lt}^*=\widehat\bfsym \beta_{lt}$. Define $$ \widehat\bfm G_t^*= \bfm P_t^*\widehat\bfsym \beta_{t}^* $$ and let $ \widehat \bfm g_{lt}^*$ be the $l$-th column of $ (\widehat\bfm G_t^*)'$.
When $\bfm g_{lt}$ is multidimensional, we present the confidence interval for a linear transformation $\bfm v'\bfm g_{lt}$.
For the first-order validity of the bootstrap, we impose a high-level assumption as Assumption (ref) in the bootstrap sampling space. In this assumption, $ \{\bfm A_{mt}^*(\bfm c), \Psi_m ^*(\bfm c), \widehat \bfm c_{z,mt} ^*, \bfm c_{z,mt} ^* : m\leq p\}$ denote the bootstrap samples of $ \{\bfm A_{mt}(\bfm c), \Psi_m (\bfm c), \widehat \bfm c_{z,mt} , \bfm c_{z,mt} : m\leq p\}$; and $h_{t,ml}^*=\bfsym \phi_{lt}^{*'}(\frac{1}{p}\bfsym \Phi_t^{*'}\bfsym \Phi_t^*)^{-1}\bfsym \phi_{mt}^*$.
Next, recall that $ \ensuremath{\boldsymbol{\xi}}_{mt,s}:=\bfm A(\bfm c_{z,mt})\nabla_c\Psi_m( \bfm c_{z, mt})\mbox{vec}\left(\frac{1}{s}(\bfm Z_{m,t+s}-\bfm Z_{mt}) (\bfm Z_{m,t+s}-\bfm Z_{mt})'-\bfm c_{z,mt } \right). $
We would like to emphasize that Assumption (ref) does not assume that $\{\bfm Z_{mt}: m\leq p\}$ should be cross-sectionally uncorrelated, because in many applications it doe not hold for $\bfm Z_{mt}$, but in fact holds for $\ensuremath{\boldsymbol{\xi}}_{mt}^0$. In the linear regression model ((ref)) for instance, $\bfm Z_{mt}=(Y_{mt}, \bfm F_{mt})$, $m\leq p$, then $\{Y_{mt}: m\leq p\}$ would not be uncorrelated if $\bfm F_{mt}$ contains common regressors (e.g., factors). On the other hand, it can be directly verfied that in this case,
where $\ensuremath{\boldsymbol{\delta}}_{mt,s} $ satisfies Assumption (ref) (ii), so is negligible; $\bfm c_{FF,mt}$ denotes the quadratic variation of $\bfm F_{mt}$. In the idiosyncratic volatility model (Example (ref)), it is also straightforward to verify that $$ \ensuremath{\boldsymbol{\xi}}_{mt,s}^0=-\bfm c_{FF,t}^{-1} \left(\frac{1}{s}(U_{m,t+s}-U_{mt})^2 -c_{uu,mt}\right)h_{t,ml}. $$ So $\ensuremath{\boldsymbol{\xi}}_{mt,s}^0$ are cross-sectionally uncorrelated given that $\{U_{m,t+s}-U_{mt}: m\leq p\}$ are cross-sectionally uncorrelated.
We use the price data of stocks from the S&P 500 index constituents for the period from July 2006 through June 2013. We collect intraday transactions data of each stock from the TAQ database and construct returns every five minutes. We drop the overnight returns for excluding stock splits and dividend issuances, and abnormal prices that bounced back within a few seconds. Stocks with missing price data are also dropped. Therefore there are in total 380 stocks in our dataset. In addition, we construct the Fama-French four factors with five-minute frequency by first generating the five-minute returns of each common stocks on the NYSE, the AMEX, and the NASDAQ in the CRSP database and then following the method described in FF. These factors are: the market factor (Mkt), the small-minus-big market capitalization (SMB) factor, high-minus-low book to market ratio (HML) factor, and the profitability factor (RMW), the difference between the returns of firms with high and low operating profitability.
We also collect fundamentals of those stocks from the Compustat database over the same period to construct firm characteristics. We consider four characteristics for each stock: size, value, momentum, and volatility as in CMO. The annual size and value characteristic of each stock is the logarithm of the market value and the ratio of the market value to the book value in the previous June respectively. The monthly momentum and volatility characteristic of each stock is the cumulative returns of the last twelve months including the previous month, and the standard deviation of the last twelve months, including previous month respectively.
Our estimation is based on the 5-minute frequency, and intervals are taken as daily windows. Hence there are $k_n=78$ observations for each stock each day. We use the linear sieve basis, which are simply the standardized values of the four characteristics, and estimate the spot characteristic and idiosyncratic beta for each company on each trading day.\footnote{We also tried B-splines with degree 3 eilers1996flexible, particularly for estimating and plotting the $\bfm g_t(\cdot)$ functions. Results obtained a very similar.} We divide the assets into three categories: large, medium, and low, based on either the firms' size or the volatility characteristics. Figure (ref) plots the cross-sectional average of the characteristic betas corresponding to each of the four factors, classified by either the size or the volatility. Both size and volatility have noticeable effects on at least one of the factor betas. The cross-sectionally averaged characteristic betas for the SMB factor are noticeably different across three size groups, and in the long run, companies with larger size (market value) tend to be less sensitive to the SMB factor than companies with smaller size. As shown in the fifth panel of Figure (ref), companies with smaller volatilities tend to be less sensitive to the market factor than companies with larger volatilities. While both phenomena have been documented in the literature, the characteristic betas, however, capture long-run movements in beta driven by structural changes in the economic environment and in firm- or industry-specific conditions, so demonstrate long-run patterns in betas from these figures.
In addition, we also estimate the cross-sectional variations in the idiosyncratic betas, measured by $\frac{1}{|G|}\sum_{j\in G}\widehat \gamma_{lt,k}^2$ for $k=$ Mkt, HML, SMB and RMW factors. In the lower panel firms are grouped by volatility: $G\in\{\text{small vol}, \text{medium vol}, \text{large vol} \}$. So the computed measure shows the cross-sectional variations among firms of small, medium and large volatilities. Figure (ref) plots $\frac{1}{|G|}\sum_{j\in G}\widehat \gamma_{lt,k}^2$ over time. There are substantial differences on the cross-sectional variations among firms with difference sizes. In particular, the strength of $\bfsym \Gamma$ is the strongest and also the most volatile across time for firms of large volatilities, is the weakest but least volatile across time for firms of small volatilities. This measures different prediction power of the characteristics on betas among firms of different level of volatilities. Figure (ref) also demonstrates that the strength of $\bfsym \Gamma$ can be represented by various asymptotic sequences across time, so the uniform inference is very essential.
We construct 95% construct confidence intervals for each of the firms' characteristic betas on a daily base, and report and compare them among three groups (by either size or volatility). On each trading day we construct the confidence intervals and calculate the proportion of positive/negative significances among firms in each group. Then we average these (cross-sectional) proportions over all days within a fixed year, leading to the “averaged proportion of significance" for each group.
When the groups are formed by size, Table (ref) reports the results of 2006, and we find that results of other years (2007 through 2012) demonstrate similar patterns: (1) All stocks have significantly positive characteristic betas loading on the market factor. In fact, most of the characteristic betas for the market factor are larger than one. (2) There is a substantial difference in the characteristic betas on the SMB factor between firms of small/medium size and firms of large size. Only 4.7% of firms of large size have positive significance, but this proportion is as high as 87% for firms of small size. On the other hand, more than fifty percent of firms of large size have negative significance, but there are less than one percent of firms of small size. This shows that the in-firm conditions and characteristics produce a long-run mechanism making small firms positively exposed and large firms negatively exposed to the SMB systematic risk. It becomes more interesting when we compare the results with the proportions of $\bfsym \Gamma$ and $\bfsym \beta$. We find that for SMB, the proportion of positive $\bfsym \beta$ is 37% for large firms, and 71% for small firms, while the proportion of negative $\bfsym \beta$ is 62% for large firms, and 28% for small firms. In contrast, these proportions respectively become 51% and 48% for positive $\bfsym \Gamma$, and 52% for negative $\bfsym \Gamma$, so the difference among firms of large and small sizes in $\bfsym \Gamma$ is much less noticeable. This suggests that the characteristic beta is the main driving horse to determine the sign of $\bfsym \beta$, while the idiosyncratic beta is more related to beta's cross-sectional variations. (3) As the size becomes larger, there is also a decreasing pattern on the negative significance of the HML beta (more noticeable on the SMB betas).
When we group firms by the volatility, however, the pattern demonstrates noticeable variations over years. The results are given in Table (ref). Results of 2010 are similar to 2011, and results in 2007 are similar to 2006 so are not presented. Firms with larger volatility tend to be more positively exposed to the HML factors than firms with smaller volatility, who are more negatively exposed to HML. This pattern appears in 2006, 2007, 2010 and 2011, but is reversed during the crisis period in 2008-2009, and European debt crisis 2012.
We now focus on two individual stocks' confidence intervals. We take the two firms that have the highest frequency to be respectively classified in the “large group” and the “small group” by size, and call them “large” and “small”. Figure (ref) plots the estimated characteristic betas and the associated confidence intervals of the two firms over time. As for the beta associated with the market factor, while both are positively significant, the characteristic betas of the firm with smaller size are constantly larger than one, making it more sensitive to the changes of market risks than the firm with the larger size. In addition, the pattern shown by the characteristic beta of the SMB factor is similar to Table (ref): in the long run, the smaller firm is positively exposed and the larger firm is negatively exposed to the SMB systematic risk.
We now test the relevance of characteristics on factor betas, an important research question in asset pricing. We consider the linear specification $\bfm g_{lt}=\bfm X_{lt}'\ensuremath{\boldsymbol{\theta}}_t$ for a common coefficient matrix $\ensuremath{\boldsymbol{\theta}}_t$ and test the relevance of each of the characteristics $\bfm X_{lt}=$ (size, value, momentum and volatility) on the betas. Note that the $(i,k)$-th element of $\ensuremath{\boldsymbol{\theta}}_t$, denoted by $\theta_{t,ik}$, represents the effect of characteristic $i$ on the $k$ the factor beta. Although this specifies a linear function $\bfm g_{t}(\cdot)$, $\bfm X_{lt}$ could include nonlinear (sieve) transformations of each characteristic. Our test is uniformly valid over $\bfsym \gamma_{lt}$. To estimate $\ensuremath{\boldsymbol{\theta}}_t$, we use the linear sieve $\bfsym \Phi_t=(\bfm X_{1t},...,\bfm X_{pt})'$ and $\bfm P_t=\bfsym \Phi_t(\bfsym \Phi_t'\bfsym \Phi_t)^{-1}\bfsym \Phi_t'$. Then $$ \widehat\ensuremath{\boldsymbol{\theta}}_t=\left( \bfsym \Phi_t'\bfsym \Phi_t \right)^{-1} \bfsym \Phi_t'\widehat \bfm G_t . $$
We construct the bootstrap confidence intervals for each component of the estimated $\ensuremath{\boldsymbol{\theta}}_t$ on each trading day, and calculate the proportion of positive (and negative) significance each year. These results are reported in Table (ref). For most of the period, the volatility has a significantly positive effect on the market factor, the value characteristic has a significantly positive effect on the HML factor, and the size characteristic has a significantly negative effect on the SMB factor. These results are consistent with the fitted $\bfm g_t(\cdot)$ functions in Figures (ref)-(ref) (in the appendix). Also note that size has insignificant effects on the market beta. We explain this from two aspects: on one hand, the market beta is mostly affected by the volatility instrument, and once it is conditioned, the size is no longer significant. On the other hand, we focus on firms that constitute to the S&P 500 index, whose sizes are relatively large, and are therefore not essential in explaining the market betas.
Finally, addtional numerical results are presented in the appendix, where we plot the estimated $\bfm g_t(\cdot)$ functions fitted by B-splines.
This paper studies a conditional factor model with a large number of individuals for high-frequency data. One of the key features of our model is that we specify the factor betas as functions of time-varying observed characteristics that pick up long-run beta movements driven by structural changes in the economic environment and in firm- or industry-specific conditions, plus a remaining (idiosyncratic) component that captures high-frequency movements in beta, which picks up short-run fluctuations in beta in periods of high market volatility. The two components capture different aspects of market beta dynamics. We show that the model can be extended to a more general continuous-time many moment conditions setting, and estimated using two-step GMM setting.
The limiting distribution of the estimated characteristic effect on the betas has a discontinuity when the strength of the idiosyncratic beta is near zero. We provide a uniformly valid inference using a cross-sectional bootstrap procedure for the characteristic betas, and do not need to pretest to know whether or not the idiosyncratic beta exists, or their strengths.
The proposed framework can be extended for inference about general coefficients in other important econometric models, such as discrete time conditional factor models, panel data models with varying coefficients, and nonlinear panel models. In these models, the structural parameter may be decomposed into the sum of a characteristic driven component plus an orthogonal component. An important example would be linear panel data model with varying coefficients. In these models the asymptotic distribution of the estimated characteristic effect would also have a discontinuity, and thus standard “plug-in" inference fails to hold uniformly. The proposed bootstrap framework would then be very useful in these models. We shall leave these studies for future research.
\singlespacing