Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
194,365 characters · 34 sections · 90 citation commands
Beta-Sorted Portfolios
Keywords: Beta pricing models, portfolio sorting, nonparametric estimation, partitioning, kernel regression, smoothly-varying coefficients, Fama-MacBeth variance estimator.
\thispagestyle{empty}
\doublespacing \setcounter{page}{1} \pagestyle{plain}
\pagestyle{plain}
Deconstructing expected returns into idiosyncratic factor loadings and corresponding prices of risk for interpretable factors is an evergreen pursuit in the empirical finance literature. When factors are observable, there are two workhorse approaches that continue to enjoy widespread use. The first approach, Fama-MacBeth two-pass regressions, have been extensively studied in the financial econometrics literature.\footnote{See, for example, jagannathan1998asymptotic, chen2004finite, shanken2007estimating, kleibergen2009tests, AngLiuSchwarz2010, gospodinov2014misspecification, adrian2015regression, bai2015fama, bryzgalova2015spurious, gagliardini2016ecmta, chordia2017cross, kleibergen2019identification, RRZ2020, giglio2021asset and many others. For a recent survey, see gagliardini2020estimation.} The second approach, which we refer to as beta-sorted portfolios, has received scant attention in the econometrics literature despite its empirical popularity.\footnote{The empirical literature using beta-sorted portfolios is extensive. For a textbook treatment, see BEM2016, and for a few recent papers see, for example, boons2020time, ChenHanPan2021, Eisdorferetal2021, GoldbergNozawa2021, FanLondonoXiao2022, CDHW2023, and chen2024investor.}
The implementation of beta-sorted portfolios entails the following two-step procedure, which incorporates a beta-adaptive portfolios construction (see, e.g., BEM2016). In a first step, time-varying risk factor exposures are estimated through (backward-looking) rolling window time-series regressions of asset returns on the observed factors. The most popular implementation uses rolling window regressions, often with a choice of a five-year window. In a second step, the estimated factor exposures, based on data up to the previous period, are ordered and used to group assets into portfolios. These portfolios then represent assets with a similar degree of exposure to the risk factors, and the degree of return differential for differently exposed assets is used to assess the compensation for bearing this common risk. Most frequently this is achieved by differencing the portfolio returns from the two most extreme portfolios. Finally, an average over time of these return differentials is taken to infer whether the risk is priced unconditionally---whether the portfolio earns systematic (and significant) excess returns. Notwithstanding the simple and intuitive nature of the methodology, little is known of the formal properties of this estimator and its associated inference procedures.
We provide a comprehensive framework to study the economic and statistical properties of beta-sorted portfolios. We first translate the two-step estimation algorithm with beta-adaptive portfolio construction into a corresponding econometric model. We show that the model has key features which are important to consider for valid interpretation of the empirical results. Notably, no-arbitrage conditions are not imposed and instead imply testable hypotheses. Furthermore, our framework precisely clarifies how the dynamics of the risk factors relate to the functional form of conditional expected returns. Within this framework, we introduce general sampling assumptions allowing for smoothly-varying factor loadings, persistent (possibly nonstationary) factors, and conditional heteroskedasticity across time and assets. We then study the asymptotic properties of the beta-sorted portfolio estimator and associated test statistics in settings with large cross-sectional, $n$, and time-series, $T$, sample sizes (i.e., $n,T\to \infty$).
This paper provides a host of new methodological and theoretical results. First, we introduce conditions that ensure validity and asymptotic normality of the beta-sorted portfolios estimator. We characterize precise conditions on the size of the window, $H$, of the first-stage rolling regression estimator, and the number of portfolios, $J$, of the second-stage estimator, relative to the sample sizes $n$ and $T$. We show that the rate of convergence of the estimator depends on the value of beta. For beta values closer to zero the rate of convergence is faster and is slower otherwise; in fact, for values of beta away from zero we show that the rate of convergence of the estimator is only $\sqrt{T}$, despite an effective sample size of the order $nT$, reflecting specific properties of the setting of interest. However, we also show that certain features of expected returns such as the discrete second derivative---which represents a butterfly spread trade---can be estimated with higher precision through faster rates of convergence for all values of beta, namely, $\sqrt{nT/J}$ for a single risk factor. This result also accommodates more powerful tests of the null hypothesis of no-arbitrage.
In addition, we uncover some limitations along with layers of nuance in the current empirical practice employing the beta-sorted portfolios methodology. First, as with all nonparametric estimators, the choice of tuning parameters, $H$ and $J$, are key to successful performance and are dependent on the sample sizes $n$ and $T$. In contrast, empirical practice often chooses window length in the first step and total portfolios in the second step irrespective of the sample size at hand. Second, we show how valid inference depends critically on the object of interest. Using our framework, we distinguish between the population and sample average of conditional expected returns and argue that the latter object should be the estimand of interest. Moreover, when the focus is on the sample average of conditional expected returns, we show that the widely-used FamaMacBeth1973 variance estimator is not consistent, in general, but instead is upward biased. However, we show that the Fama-MacBeth variance estimator still leads to valid, albeit possibly conservative, inference.
To address this limitation, we propose a new variance estimator and provide an empirical implementation which also produces valid, albeit possibly conservative, inference. That said, in our empirical application we show that our new variance estimator provides much sharper inference than the Fama-MacBeth variance estimator. Finally, we also provide results on the limitations of the beta-sorted portfolio estimator for a fixed time period. We show that differential returns for each period, often used as inputs for assessing the time-series properties of conditional expected returns, are contaminated by an additional term when risk factors are serially correlated. However, we demonstrate that some features of conditional expected returns for a single time period can be illuminated by utilizing butterfly spread trades.
From a theoretical perspective, beta-sorted portfolios present a number of technical challenges originating from the two-step estimation algorithm with beta-adaptive portfolio construction, since it relies on two nested nonparametric estimation steps together with a portfolio construction based on a first-step nonparametric generated regressor. More precisely, the first-stage nonparametrically estimated factor loadings enter directly into the non-smooth partitioning scheme further complicating the analysis.\footnote{For analysis of partitioning-based nonparametric estimators see Cattaneo-Farrell-Feng_2020_AoS and references therein. Partitioning-based estimators with random basis functions have been recently studied in cattaneo2020characteristic and Cattaneo-Crump-Farrell-Feng_2024_AER, but in those papers the conditioning variables are observed, while here the conditioning variable is generated using a preliminary time-series smoothly-varying coefficients nonparametric regression, and therefore prior results are not applicable to the settings considered herein.} To our knowledge, we are the first to prove validity of such an approach.
This paper is most related to the large literature studying asset pricing models with observable factors.\footnote{See, for example, goyal2012empirical, nagel2013empirical, GospodinovRobotti2013, or gagliardini2020estimation for surveys. A related literature endeavors to jointly estimate factor loadings and latent risk factors. See, for example, ConnorLinton2007, connor2012efficient, fan2016projected, kelly2019characteristics, connor2021dynamic, FTLN2022, and BCLT2024.} Given our focus on conditional asset pricing models with large panels in both the cross-section and time-series dimension, this paper is most closely related to gagliardini2016ecmta gagliardini2020estimation. The authors introduce a general framework and econometric methodology for inference in large-dimensional conditional factors under no-arbitrage restrictions. They allow for risk exposures, which are parametric functions of observable variables and provide conditions to consistently estimate, and conduct inference on the prices of risk. Although the statistical model under study shares important similarities with the setup of gagliardini2016ecmta, there are substantial differences, and the models explored previously in the literature do not nest our setup. For example, the classical beta-sorted portfolio estimator implies a data-generating process that does not (necessarily) exclude arbitrage opportunities and supposes risk exposures which are smoothly-varying. Furthermore, we show that valid estimation and inference can be achieved without requiring an assumption of the functional form of the conditional expectation of the risk factors. This is in contrast to the existing literature (e.g., adrian2015regression, gagliardini2016ecmta) where such an assumption is utilized.
Our paper is also related to the financial econometrics literature on nonparametric estimation and inference. In particular, the two steps of the beta-sorted portfolio algorithm align individually with ang2012testing, who study kernel regression estimators of time-varying alphas and betas, and cattaneo2020characteristic who study portfolio sorting estimators given observed individual characteristic variables. However, the individual results from each of these papers cannot be applied in our multi-step setting. Furthermore, the linkage between the two steps, through the role of the generated (nonparametrically estimated) regressor in the second-stage nonparametric partitioning estimator, represents a substantial technical challenge and has not been studied before. Furthermore, we offer a host of new results characterizing the properties of beta-sorted portfolios.
We demonstrate the practical usefulness of our estimation and inference results for beta-sorted portfolios in a substantive empirical application with a novel methodological contribution. More precisely, we introduce a new risk factor---a measure of the business credit cycle---and show that it is strongly predictive of both the cross-section and time-series behavior of U.S. stock returns. We also show the effectiveness of our new variance estimator as inference is much more informative relative to inference employing the widely used Fama-MacBeth variance estimator.
In summary, this paper makes a number of contributions to understanding the underlying foundational properties and practical use of beta-sorted portfolios: we introduce an econometric framework to study the beta-sorted portfolio estimator and clarify its properties (Section (ref)); we provide asymptotic theory for the estimator accommodating the nonparametric first and second steps (Sections (ref) and (ref)); we characterize the properties of the commonly-used Fama-MacBeth variance estimator along with a new plug-in variance estimator (Section (ref)); we provide results for joint inference across multiple values of beta, including a new test of no-arbitrage restrictions in Section (ref); and we introduce a novel risk factor and demonstrate its desirable properties in an empirical application (Section (ref)). Proofs of our theoretical results are relegated to a Supplemental Appendix (hereafter, SA) to streamline the exposition, and may be of independent technical interest.
For a constant $k \in N$ and a vector $v = (v_1, \ldots, v_d)^\top \in \mathbb{R}^d,$ we denote $|v |_k = (\sum_{i=1}^d |v_i|^k)^{1/k}$, and $|v|_\infty = \max_{1\leq i \le d} |v_i|$. For a random variable $V$, let $\|V\|_q = ( \mathbb{E}[ | V |^q ] )^{1/q}$. We set $(a_n: n\ge 1)$ and $(b_n: n\ge 1)$ to be positive number sequences. We write $a_n=O(b_n)$ or $a_n\lesssim b_n$ (resp. $a_n\asymp b_n$) if there exists a positive constant $C$ such that $a_n/b_n\leq C$ (resp. $1/C \le a_n / b_n\leq C$) for all large $n$, and we denote $a_n=o(b_n)$ (resp. $a_n\sim b_n$), if $a_n/b_n\to 0$ (resp. $a_n/b_n\to C$). $\operatorname*{plim}_{n\to\infty} X_n=X$ means that $X_n \to_\mathbb{P} X$. $\to_{\mathcal{L}} $ denotes convergence in law. Define $X_n = O_\mathbb{P}(a_n): \exists N_{\varepsilon}>0, \delta_{\varepsilon}>0 \quad \text { such that } \mathbb{P}(|X_n| \geq \delta_{\varepsilon}) \leq \varepsilon \quad \forall n>N_{\varepsilon}$. Define $X_n = o_\mathbb{P}(a_n): \forall \varepsilon, \delta >0 \quad \exists N_{\varepsilon, \delta} >0 \quad \text { such that } \mathbb{P}(\left|X_n\right| \geq \delta) \leq \varepsilon \quad \forall n>N_{\varepsilon, \delta}$. We also let $X_n \lesssim_\mathbb{P} a_n$ to mean $X_n = O_\mathbb{P}(a_n)$; furthermore, if $X_n \lesssim_\mathbb{P} a_n $ and $a_n \lesssim_\mathbb{P} X_n$ then we write $X_n \asymp_{\mathbb{P}} a_n$. Limits are taken as $n,T,J,H \to\infty$ unless otherwise stated explicitly. Set $a\vee b=\max(a, b)$ and $a \wedge b=\min(a, b)$.
Following wu2005nonlinear, CSW2021, and han2023probability, we consider the following measure of time-series dependence throughout.
Let $Y^*_{\bullet,n,T}(\ell)$ be a NTSS with $Y_{t,n,T}^*(\ell)$ a random variable with $\xi_{t-\ell}$ replaced by $\xi_{t-\ell}^*$, where $\xi_{t}^*$ is an i.i.d. copy of $\xi_{t}$ for each $t\in\mathbb{Z}$. Then, the dependence adjusted norm of $Y_{\bullet,n,T}$ is
with $v\geq0$ and $q\geq1$. When the $Y_{\bullet,n,T}$ does not depend on $n$ and $T$, we simplify the notation to $\Theta_T(Y_{\bullet};q,v)$ for the possibly nonstationary NTSS $Y_{\bullet}:= (Y_{t}: t\in\mathbb{Z})$, where $Y_{t} = g_t(\cdots, \xi_{t-2}, \xi_{t-1}, \xi_t)$ with the function $g_t(\cdot)$ no longer a function of $n$ and $T$. Furthermore, if $g_t(\cdot)$ is not a function of $t$, then $Y_{\bullet}$ is a stationary NTSS and we write $\Theta(Y_{\bullet};q,v)$.
We introduce a general econometric model of asset returns, and show how it aligns with the two steps that comprise the beta-sorted portfolio algorithm. We discuss the relevant properties of the model with respect to the potential presence of arbitrage opportunities.
Let $R_{it}$ denote the return of asset $i$ at time $t$, and $f_t$ a vector of observable risk factors with $f_t \in \mathbb{R}^d$. We assume that asset returns are generated by the linear stochastic coefficient model,
where $\alpha_{it}$ and $\beta_{it}$ are smoothly-varying, random coefficients with $\alpha_{it}, \beta_{it} \in\mathcal{F}_{t-1}$, $ \varepsilon_{it}$ is an idiosyncratic error term, and $\mathcal{F}_{t}$ is the information set up to time $t$. To be more precise, we define the sigma field,
where $\aleph_{1,i}$, $\aleph_{2,t}$, and $\aleph_{3,it}$ are vectors of auxiliary random variables, which are possibly observed, and will be discussed further in later sections. Since $\alpha_{it}$ and $\beta_{it}$ are both $\mathcal{F}_{t-1}$-measurable we have that $\mathbb{E}[ \left. R_{it} \right| \mathcal{F}_{t-1},f_t] = \alpha_{it} + \beta_{it}^\top f_t$. The sigma field $\mathcal{F}_{t}$ will, in general, depend on $n$ and $T$ but we suppress this dependence for notational convenience. Finally, notice that equation (ref) accommodates an unbalanced panel. Although each $n_t$ may be different, we assume that they all grow at the same rate which ensures that each cross-section contributes to the asymptotic properties of the estimator (i.e., $n_t \asymp n $ for $t = 1,\cdots, T$). For an alternative example of a random coefficient model tailored to a financial application, see BarrasGagliardiniScaillet2022.
To obtain the structural form of our model, we assume that there exists a non-random function $\mu_0(\cdot)$ such that
Equation (ref) restricts conditional expected returns to be a function of $\beta_{it}$ only, but it allows the functional form to be random and to vary with past realizations of $i$-invariant random variables. This restriction captures the notion that nonzero expected returns must reflect the compensation investors require for exposure (as measured by $\beta_{it}$) to the risk factors, $f_t$. We will make precise assumptions on $\mu_t(\cdot)$ in later sections but, loosely, one can think of $\mu_{t}(\cdot)$ as being a sufficiently smooth random function of $\beta_{it}$.
Using equations (ref) and (ref) we obtain,
which clarifies the restriction on $\alpha_{it}$ implied by equation (ref). Equation (ref) is a nonparametric analog to a reduced-rank restriction in parametric models as it imposes cross-equation restrictions between the linear coefficients in equation (ref).
Finally, combining equations (ref) and (ref), we arrive at the structural form
We assume throughout that $R_{it}$ represents excess returns, but $\mu_t(0)$ may be interpreted as the zero-beta rate at time $t$ in the case when $R_{it}$ represents raw returns. Equation (ref) may be compared to the standard beta pricing model Cochrane2005 and generalizations thereof cochrane1996jpe,adrian2015regression,gagliardini2016ecmta. The most noteworthy difference is the presence of the (possibly) nonlinear, time-varying function $\mu _{t}\left( \beta _{it}\right)$ in equation (ref). When $R_{it}$ represents excess returns then the no-arbitrage restriction implies that $\mu _{t}\left( \beta _{it}\right) =\beta _{it}^{\top }\lambda _{t}$ for some $\lambda _{t}$ gagliardini2016ecmta. Our model nests, but does not require, the imposition of the absence of arbitrage opportunities so that
The presence of this additional term representing the deviation from no-arbitrage restrictions can be motivated by appealing to structural models which feature violations of the law of one price. Such a setup as in equation (ref) could arise, for example, in the margin-constraints model of garleanu2011margin under the assumption that the security's margin is a nonlinear function of its (past) beta.
Throughout the paper we set $d=1$ to simplify the notation and the exposition. All results can be generalized to the $d>1$ case. To see why equation (ref) rationalizes the beta-sorted portfolio algorithm, let us revisit the standard two steps:
\paragraph{Step 1: Estimation of $\alpha_{it}$ and $\beta_{it}$.} For each individual asset, we calculate the rolling window (local constant) regression estimator for $\alpha_{it}$ and $\beta_{it}$ as,
where $X_{t} = (1, f_t)^{\top}$ and $H$ is the window length. This construction purposely does not have “look-ahead bias”, as neither the estimators of $\widehat{\alpha}_{it_0}$ or $\widehat{\beta}_{it_0}$ use data from time $t_0$ (or after) in their construction (a “leave-one-out” estimator). For each cross-sectional unit, this estimation of the time-varying random coefficients can be interpreted as a leave-one-out kernel regression of equation (ref) using a uniform kernel and a choice of bandwidth $h$ which satisfies $H = \lfloor Th\rfloor$, where $\lfloor . \rfloor $ denotes the floor function. In the SA we provide all proofs for an arbitrary one-sided kernel which generalizes the results we present in the following sections. $\blacksquare$
\paragraph{Step 2: Sorting portfolios using estimated $\beta_{it}$.} To see that this step comprises a cross-sectional nonparametric estimation, observe that, for fixed $t$, equation ((ref)) is the conditional mean of interest. Now, to cement intuition, assume (temporarily) that $\beta_{it}$ is observed and takes on a finite number of values, one of which is $\bar{b}$. Then, to estimate $\mathbb{E}[ R_{it} | \beta_{it} = \bar{b}]$ we need only calculate the sample mean of returns for assets $i$ for which $\beta_{it}=\bar{b}$:
When $\beta_{it}$ can take on a continuum of values, we can estimate $\mathbb{E}[ R_{it} | \beta_{it} = \bar{b}]$ by averaging returns for those assets $i$ such that $\beta_{it}$ is “near” $\bar{b}$. Forming portfolios is one way to operationalize this local averaging approach. Then, the portfolio return representing assets with values of $\beta_{it}$ closest to $\bar{b}$ is our estimate for this conditional expectation.
We define $\mathcal{B} = [ \beta _{l},\beta _{u}] $, with $\beta _{l}$ and $\beta _{u}$ fixed constants, as the support of the possible realizations of $\beta _{it}$ across $i$ and $t$. Since, in practice, $\beta_{it}$ is unobserved, we may obtain estimates $\widehat{\beta}_{it}$ from Step 1 and, for each $t=1,\ldots ,T$, we can define the beta-adaptive partition of $\mathcal{B}$ as
where $\widehat{\beta}_{(\ell) t}$ denotes the $\ell$th order statistic of the estimated betas in the first step, across $i$ for fixed $t$, i.e., the order statistics of $\{\widehat{\beta}_{it}:i=1,\ldots ,n_{t}\}$, and where we set $\widehat{\beta}_{(0)t} = \beta_l$ and $\widehat{\beta}_{(n_t)t} = \beta_u$ for simplicity. The intervals, $\widehat{P}_{jt}$, allow us to partition assets into portfolios based on their estimated beta, $\widehat{\beta}_{it}$. The number of portfolios $J_t$, and their random structure (i.e., breakpoint positions based on estimated $\beta_{it}$), vary for each time period but we assume that $J_t \asymp J$. Finally, we can construct portfolio returns by averaging the returns of the assets in each portfolio. $\blacksquare$
Given the two-step construction outlined above, suppose we would like to estimate $\mathbb{E}[ R_{it} | \beta_{it} = \beta_l ]$. The estimator is simply the portfolio return for the first portfolio, $\widehat{P}_{1t}$, which we write as $\widehat{\mu}_t(\beta_l)$. Similarly, we can estimate $\mathbb{E}[ R_{it} | \beta_{it} = \beta_u ]$ with the portfolio return for the last portfolio, $\widehat{P}_{J_t t}$, as $\widehat{\mu}_t(\beta_u)$. We can then average the differential portfolio returns for these two extreme portfolios across all available time periods as,
This is exactly the estimator that is used in practice.
A few comments are in order. First, the above two steps are completely in line with the empirical finance literature. Importantly, at no point in the two-step algorithm is there a requirement to estimate the conditional expectation of the risk factors, $\mathbb{E}[f_{t}\vert \mathcal{F}_{t-1}]$, and so the researcher remains agnostic about the dynamics of these risk factors. We will revisit this issue in later sections. Second, the practice of using moving-window regressions to accommodate time variation in $\beta_{it}$ suggests a slowly-varying coefficient model as previously used in finance applications such as in ang2012testing and adrian2015regression. However, in contrast to these previous formulations, we do not condition on the realizations of the random processes $\alpha_{it}$ and $\beta_{it}$. Instead, we retain the randomness in these objects so that the second-stage beta-sorted portfolio estimator can have a well-defined limit as $n,T \to \infty$. Third, an alternative to the smoothly-varying coefficients approach is to specify $\beta_{it}$ as a function of individual characteristics and possibly also of economy-wide variables gagliardini2020estimation. Our approach can accommodate such settings by modifying the rolling window regressions (kernel regressions) appropriately.
Although we are motivated by empirical practice, which has been replicated exactly in equation ((ref)), we introduce and work with the more general estimation approach,
Let $\widehat{P}_{{j}^\star_t t}$ be the portfolio that contains the value $\beta$ at time $t$. Then, $\widehat{\mu}_t(\beta)$ is simply the return of $\widehat{P}_{{j}^\star_t t}$. Importantly, the value of ${j}^\star_t$ will generally change over time. For example, it may be that assets with values of $\beta$ near $1/2$ fall in the sixth portfolio at times and the fifth portfolio at other times and so on. Thus, this more general estimation approach does not constitute spurious generality. The conventional implementation of beta-sorted portfolios relies on a constant choice of $J_t=J \,\, \forall t$ and so averages $J$ portfolios across all time periods. However, if the cross-sectional distribution of the $\beta_{it}$ are changing over time then there is no guarantee that a chosen portfolio represents assets with sufficiently similar betas. Therefore, the conventional estimator will be, in general, both more biased and more variable than the estimator given in equation (ref), all else equal. This is of special importance when we are interested in expected returns for intermediate values of betas and also in situations where tests of monotonicity or shape restrictions are of interest. We discuss these issues further in later sections.
The first step in the estimation procedure involves a sequence of rolling window time-series regressions. To establish consistency (and, in the SA, asymptotic normality) of $\widehat{\beta}_{it}$ we require technical, but relatively standard, assumptions on the underlying data generating process. We first restrict the behavior of the idiosyncratic error terms, $\varepsilon_{it}$, in equation (ref).
Assumption (ref) imposes moment conditions on the idiosyncratic error term, $\varepsilon_{it}$ along with regularity conditions controlling the rate of decay of the time-series dependence. We define $\mathbb{E}[\varepsilon_{it}^2| \mathcal{F}_{t-1},f_t] = \sigma_t^2$ which Assumption (ref) ensures is bounded and bounded away from zero. It is important to emphasize that for the first step we need only impose that equation (ref) holds (without imposing equation (ref)). We now characterize the behavior of the risk factors, $f_t$.
Assumption (ref) imposes some structure on the time series properties of the factor $f_t$ but is quite general and allows for certain forms of nonstationary behavior. We could relax some of these assumptions to allow for even more complex time-series properties at the expense of more detailed notations and proofs. We next restrict the behavior of $\alpha_{it}$ and $\beta_{it}$.
Assumption (ref) ensures that the alphas and betas, although random, are sufficiently smooth over time (i.e., satisfying a Lipschitz-type condition). This is the formal assumption which justifies the common empirical approach of using rolling window regressions.
To provide intuition for our proof approach recall that $X_{t}= (1, f_{t})^{\top}$ so that for a fixed time period $t_0\in [H+1 ,T]$ we can rewrite equation (ref) as \[R_{it_0} = X_{t_0}^{\top}(\alpha_{it_0},\beta_{it_0}) + \varepsilon_{it_0}.\] Then, by Assumption (ref) and since $\alpha_{it}$ and $\beta_{it}$ are measurable with respect to $\mathcal{F}_{t-1}$, $\alpha_{it_0}$ and $\beta_{it_0}$ can be identified as \[\mathbb{E}(X_{t_0} X_{t_0}^{\top}|\mathcal{F}_{t_0-1})^{-1}\mathbb{E}(X_{t_0} R_{it_0}|\mathcal{F}_{t_0-1}).\] We can use the estimator from equation (ref) to obtain $ (\widehat{\alpha}_{it_0}, \widehat{\beta}_{it_0} )$. In order to accommodate the random coefficients we exploit the fact that $\sum^{H}_{s=1} X_{t_0-s}X_{t_0-s}^{\top}$ and $\sum^{H}_{s=1} X_{t_0-s}R_{i(t_0-s)}$ are close, in the appropriate sense, to $\sum^{H}_{s=1} \mathbb{E}[ X_{t_0-s}X_{t_0-s}^{\top}|\mathcal{F}_{t_0-1}]$ and $\sum^{H}_{s=1} \mathbb{E}[ X_{t_0-s}R_{i(t_0-s)}|\mathcal{F}_{t_0-1}]$. This follows because their difference are summands of martingale difference sequences.
We first provide a (uniform) consistency result of our estimator $\widehat{\beta}_{it_0}$ over $i$ and {$t_0$}. This result generalizes the time varying coefficient analyses in zhang2012inference. We require this result to precisely control the effect of estimating $\beta_{it}$ in the first step when entering the second-step estimator. We establish this consistency on a compact interval of a trimmed support $[H+1,T]$.
Theorem (ref) provides uniform (over $t$ and $i$) rates of convergence for the first-stage rolling-window (kernel) estimators of the betas. Naturally, these rates depend on $n$, $T$, and $H$ but are also directly dependent on $q$ which represents the number of bounded moments of the idiosyncratic error term and the observed factors. For sufficiently large $q$, consistency of the rolling regression estimator can attain the usual nonparametric (optimal) convergence rate. In the SA we also provide conditions that ensure asymptotic normality of $\widehat{\beta}_{it}$, which may be of independent interest.
The second step of the estimation procedure is to sort assets into portfolios based on their value of $\widehat{\beta}_{it}$, which is obtained from the procedure described in the previous section. Then, returns for each portfolio are constructed and used to assess the relation between exposure to the risk factors and subsequent asset returns. In this section, we formalize the properties of this estimator and provide conditions for its validity. We also discuss the importance of the cross-sectional restrictions implied by equation (ref) for our main results. Before proceeding, however, we will first introduce two different objects of interest which will be useful for clarifying our results, and discussing empirical practice more broadly.
Let us first define the sample average conditional expected returns (SACER) as:
We can compare the SACER directly to equation ((ref)) to observe that this would, at first glance, be the natural candidate for our “estimand”. However, it is important to point out that the SACER is a sample average of random quantities and, hence, is random itself. The randomness arises from the presence of $\mu_t(\cdot)$ which depends on the $i$-invariant sigma field $\mathcal{G}_{t-1}$. We can contrast SACER to its population counterpart. To do so we need the following assumption:
\setcounter{assumption}{15}
\setcounter{assumption}{3}
Assumption (ref) imposes an ergodicity condition on the sample average of conditional expected returns. When $H/T=o(1)$ (as we impose) then Assumption (ref) also implies that $\sup_{\beta\in\mathcal{B}}| \bar{\mu}_T(\beta;H) - {\mu}(\beta) | = o_\mathbb{P}(1)$ so that the PACER can be thought of as the probability limit of the SACER. In the special case where $\mu_t(\cdot)=\mu(\cdot)\,\, \forall t$, then the SACER and PACER are equal.
The SACER and the PACER represent two distinct objects of interest. To make things concrete, suppose that $\mu_t(\cdot)$ is random only through a finite collection of strictly stationary state variables for the economy, say, {$\mathsf{S}_{t-1}\in \mathcal{G}_{t-1}$}. Then, when we study the SACER we are learning about average conditional expected returns for the realizations of these state variables over a specific time period only. Equation (ref) allows for more generality; however, we form our discussion around a finite set of state variables governing the economy for simplicity of exposition and to cement intuition. Consequently, for any particular sample, it may be that the SACER and PACER are “far away” from each other.
We recommend that empirical researchers focus on the SACER rather than the PACER for a few reasons. First, even if Assumption (ref) fails to hold, the SACER is well-defined under our remaining assumptions. The PACER may not exist if, for example, some state variables driving $\mu_t$ are not stationary. Second, and more importantly, the SACER is a more intuitive and interpretable object of interest as we have a wealth of data summarizing what occurred over our sample. For example, we observe the behavior of macroeconomic aggregates, changes in the legal and regulatory landscape, and the presence of unusual events (e.g., financial crises or natural disasters). Conversely, interpretation of the PACER will generally hinge on the ergodic distribution for $\mu_t(\cdot)$ which itself depends on how the (subset of) economy-wide state variables drive conditional expected returns. Characterizing the properties of this distribution would be challenging (e.g., assessing the probability of recessions or financial crises) and the appeal of the approach herein is that we can avoid making specific assumptions about the functional form of $\mu_t(\cdot)$. Third, inference on the PACER is necessarily at least as challenging as inference on the SACER since, in practice, we only observe a finite sample. We formalize this intuition in Theorem (ref). Despite our preference for the SACER, in Section (ref) we provide a further discussion of issues related to inference on the PACER.
We impose the following assumption to justify the second step of the standard empirical approach.
Assumption (ref) formally imposes the economic restrictions on conditional expected returns discussed in Section (ref). We restrict the conditional expectation of the factors to be a function of random variables in the smaller sigma field $\mathcal{G}_t$. Taken together, Assumption (ref) provides the foundation for the validity of the beta-sorted portfolio estimator. Let us first define systematic realized returns as
where the second equality follows by Assumption (ref). Next, to provide intuition, temporarily assume that the $\beta_{it}$ are observed. Then, using equation (ref), our model can be written as
and, under Assumption (ref), we have that $\mathbb{E}({\varepsilon}_{it}| \mathcal{F}_{t-1},f_t) = 0$ (recall that $\beta_{it} \in \mathcal{F}_{t-1}$). The second equality in equation (ref) makes clear that, for a fixed time period, we can only nonparametrically estimate the unknown function $M_t(\cdot)$ rather than the direct object of interest $\mu_t(\cdot)$. However,
The second term has summands, $\beta(f_t - \mathbb{E}[f_t|\mathcal{G}_{t-1}])$, which are a martingale difference sequence with respect to $\mathcal{G}_{t}$ and so we would expect this sample average to converge to zero in probability; consequently, this would ensure that $(T-H)^{-1} \sum_{t=H+1}^T {M}_t(\beta) $ and $(T-H)^{-1} \sum_{t=H+1}^T \mu_t(\beta)$ are close in probability for large $T$. A further complication, of course, is introduced by using an estimated $\beta_{it}$ in the second-stage nonparametric regression. Nevertheless, later in this section, we will make these arguments rigorous and provide appropriate conditions for valid estimation and inference methods based on the beta-sorted portfolio estimator.
Assumption (ref) allows us to highlight another important issue. As discussed in Section (ref), the first-stage estimator requires smoothly-varying random coefficients $\alpha_{it}$ and $\beta_{it}$. However, the cross-sectional restrictions in Assumption (ref) impose additional structure on these random coefficients. The combination then implies restrictions on the functional form of conditional expected returns. For example, if $\mathbb{E}[f_t|\mathcal{G}_{t-1}]$ is constant for all $t$, then, for $\alpha_{it}$ to be smooth over time, we could have $\mu_t(\cdot)$ be constant over time or, alternatively, a smooth random function over time. However, consider a more realistic example. Let $\mathsf{S}_t \in \mathcal{G}_t$ again be a vector of strictly stationary state variables, but further assume $\mathbb{E}[f_t|\mathcal{G}_{t-1}] = \zeta + \Lambda \mathsf{S}_{t-1}$, $\Lambda\neq0$ (as in, e.g., adrian2015regression or gagliardini2016ecmta). Then, since equation (ref) holds by Assumption (ref), we must have that $\mu_t(\cdot)$ varies over time or else we violate Assumption (ref) and the first-stage estimator of $\beta_{it}$ will not be consistent in general. Moreover, when $\mu_t(\cdot)$ varies over time, the SACER and PACER will not be equal, and this directly affects the interpretation of the estimation and inference results in any particular empirical application; see Theorem (ref) and associated discussion. Results such as these underscore the importance of providing a formal framework for interpreting the beta-sorted portfolio estimator.
The remainder of this section presents our main results, culminating in the asymptotic normality of the beta-sorted portfolio estimator under appropriate conditions. We first restrict the cross-sectional behavior of the idiosyncratic error term, $\varepsilon_{it}$.
This type of sampling assumption was introduced in Andrews2005 and has been utilized in the financial econometrics literature by gagliardini2016ecmta and cattaneo2020characteristic. The assumption is quite general and allows for higher-order dependence in the distribution of $\varepsilon_{it}$; for example, it accommodates conditional heteroskedasticity in the innovations over time as a function of the factors $f_t$ and the variables in $\aleph_{2,t}$.
We next impose restrictions on the $\beta_{it}$, which serve as the covariates in the second-stage nonparametric regression.
Assumption (ref) imposes regularity conditions on the conditional distribution of the $\beta_{it}$ ensuring that it is sufficiently well behaved. Specifically, the assumption ensures that the partitioning estimator is well defined with the probability of empty portfolios vanishing asymptotically. Furthermore, if we define $\Phi_{i,j,t} = \mathds{1}(F^{-1}_{\beta,t}((j-1)/J_t) \leq \beta_{it}< F^{-1}_{\beta,t}(j/J_t))$ then $q_{jt}=\mathbb{E}[\Phi_{i,j,t}|\mathcal{G}_{t-1}]$ is of the order $J^{-1}$. The conditional i.i.d. assumption in Assumption (ref) is similar to that of Assumption (ref) and is quite general allowing, for example, a nonlinear factor structure in $\beta_{it}$. As a concrete example, let $\tilde{\aleph}_{1,i}$ and $\tilde{\aleph}_{3,i(t-1)}$ be selected i.i.d. variables from $\aleph_{1,i}$ and $\aleph_{3,i(t-1)}$, respectively. Then, $\beta_{it} = g_{\beta,0}(\frac{t}{T}; g_{\beta,1}(\tilde{\aleph}_{1,i},\tilde{\aleph}_{3,i(t-1)})^\top g_{\beta,2}(\aleph_{2,t-1},\ldots,\aleph_{2,t-k_\aleph}, f_{t-2}, \ldots, f_{t-k_f}))$ for sufficiently smooth $g_{\beta,0}(\cdot)$, measurable (non-constant) $g_{\beta,1}(\cdot)$ and $g_{\beta,2}(\cdot)$, and fixed lag lengths $k_\aleph$ and $k_f$, satisfies Assumption (ref).
Finally, the last assumption we require is that our object of interest, $\mu_t(\beta)$, is sufficiently smooth in $\beta$ to accommodate nonparametric estimation.
This assumption is standard in the nonparametric literature and rules out discontinuities and other pathologies that would invalidate standard nonparametric estimation approaches. As discussed in Section (ref), portfolio sorting can be interpreted as a nonparametric estimator of a conditional mean. In order to operationalize the estimator let us define $F_{\widehat{\beta},n,t}(\cdot) =\frac{1}{n_t}\sum_{i=1}^{n_t}\mathds{1}(\widehat{\beta}_{it} \leq \cdot)$ and $F^{-1}_{\widehat{\beta},n,t}$ as the empirical CDF and empirical quantile function for the estimates of the $\beta_{it}$ (recall that $\widehat{\beta}_{(s)t} = F_{\widehat{\beta},n, t}^{-1}(s /n_t)$ for $s=1,\ldots,n_t$), respectively. Then we can define
In words, $\widehat{\Phi}_{i,j,t}$ is a binary variable which takes on the value of one when asset $i$ is in portfolio $j$ at time $t$, and zero otherwise. We can stack these binary variables into the $J_t \times n_t$ matrix $\widehat{\Phi}_t$ and obtain
where $\widehat{p}_{t}(\beta) = (\widehat{p}_{1t}(\beta), \ldots, \widehat{p}_{J_t t}(\beta))^{\top}$ with $\widehat{p}_{jt}(\beta) = \mathds{1}(\beta \in \widehat{P}_{jt})$, $\widehat{a}_t = (\widehat{\Phi}_t \widehat{\Phi}_t^{\top} )^{-1} \widehat{\Phi}_t R_t = (\widehat{a}_{1t},\cdots,\widehat{a}_{J_t t})^\top$, and $R_t = (R_{1t},\ldots, R_{n_t t})^{\top}$. Equation ((ref)) shows that $\widehat{\mu}_t(\beta)$ requires two inputs. The first input is which (unique) portfolio the evaluation point, $\beta$, resides in. In equation ((ref)) this translates to the choice of $j^\star$ which gives $\widehat{p}_{j^\star t}(\beta)=1$. The second input is the return on the $j^\star$th portfolio. This is simply given by the $j^\star$th element of $\widehat{a}_t$.
In small samples there is always the possibility that some portfolios are empty so that the inverse of $\big(\widehat{\Phi}_t \widehat{\Phi}_t^{\top} \big/ n_t \big)$ may not exist. However, under our assumptions, the results in the SA show that $\big(\widehat{\Phi}_t \widehat{\Phi}_t^{\top} \big/ n_t \big)$ exists and is finite with probability approaching one. Throughout, we assume we are on the event that this inverse exists but we suppress this from the notation and main results for simplicity of exposition. The following theorem characterizes the leading terms of the beta-sorted estimator.
Lemma (ref) introduces some key properties of the grand mean estimator, $\widehat{\mu}(\beta)$. First, under our assumptions, we may ignore the generated errors of the first-stage estimation of the $\beta_{it}$ when analyzing the second stage portfolio sorting estimator. Second, the theorem shows that when estimating the SACER the leading term is comprised of two elements: the first term, {equation (ref)}, would appear in any generic nonparametric problem whereas the second term, equation (ref), is specific to the asset pricing setup. Importantly, the second term is of the order $O_\mathbb{P}(T^{-1/2})$ representing the summation of the product of the conditional beta and the deviation of the factor from its conditional mean, $(f_t - \mathbb{E}[f_t|\mathcal{G}_{t-1}] )$. Thus, despite an approximate sample size of $nT$ in concert with a nonparametric procedure with tuning parameter $J$, the grand mean estimator, for all values of $\beta$ except zero, achieves only a $\sqrt{T}$ rate of convergence. That said, for $\beta$ evaluated at zero, there is a discontinuity and the second term becomes degenerate and the first term dominates leading to a faster rate of convergence, namely, $O_\mathbb{P}\big(\sqrt{\frac{J}{nT} + \frac{1}{TJ^2}} \big)$.\footnote{To see that this rate of convergence is faster, note that $J/n \to 0 $ is required to ensure that the probability of an empty portfolio is vanishing asymptotically.} This is reflected in the remainder term, $\mathscr{R}(\beta)$, in Lemma (ref). Finally, we see that the bias of the estimator at time $t$, $\mathscr{B}_t(\beta)$, and also for the grand mean, $\mathscr{B}(\beta)$, are of the order $O_\mathbb{P}(J^{-1})$. In words, the sample average across time does not alter the order of the bias since $\widehat{\mu}(\beta)$ is a sample average of nonparametric estimates taken one cross-section at a time.
Next we provide a central limit theorem for $\widehat{\mu}(\beta)$ which allows us to conduct inference on the estimator of the SACER. Our result employs a martingale central limit theorem hall2014martingale combined with the dependence structure introduced earlier.
Theorem (ref) characterizes the limiting distribution of the beta-sorted portfolio estimator when centered at the SACER (plus the smoothing bias, $\mathscr{B}(\beta)$). A few remarks are in order. First, despite the differential rates of convergence in the numerator depending on the value of $\beta$ shown in Lemma (ref), once properly scaled, the limiting distribution is unaffected. This comes about because the denominator adapts to the appropriate rate of convergence. Second, Theorem (ref) provides the basis of a feasible inference procedure based on a new plug-in variance estimator, discussed in Section (ref). Third, we require restrictions on the dependence properties of the data, using our concept of time-series dependence introduced in Section (ref), to ensure the convergence in distribution holds.
Finally, to provide some intuition for the rate conditions given in Theorem (ref) we give the following example for $\beta \neq 0$. Let $n = T^{\gamma_1}$, $h =T^{-\gamma_2}$ and $J= T^{\gamma_3}$, for $\gamma_{1}, \gamma_2,\gamma_3> 0$, $\gamma_2<1$ and $\gamma_3<\gamma_1$. If we ignore log terms, the rate conditions in the above theorem require that $2\gamma_1 + q\gamma_2 < q-2$ and $1 < 2\gamma_3 < \gamma_1$. Thus, even for large $\gamma_1$ (e.g., $\gamma_1=2$), our rate requirements conform with a choice of bandwidth that produces meaningful undersmoothing in the first step relative to, for example, the MSE optimal rate of $T^{-1/3}$. Intuitively, this comes about because variation in the first-stage estimates is beneficial for estimating $\mu_t(\cdot)$ since $\beta_{it}$ are the regressors in the second-stage nonparametric estimation problem. As a consequence the trade-off between bias and variance that would arise if the first stage was considered in isolation is no longer the case in the combined procedure: the cost of bias relative to variance rises.
We have established conditions which ensure asymptotic normality of the beta-sorted portfolio estimator when centered at the SACER. In this section, we investigate the properties of two feasible inference procedures. The first is the so-called Fama-MacBeth (FM) variance estimator which is ubiquitous in existing applications of beta-sorted portfolios. The FM variance estimator can be motivated by the classical sample variance estimator, and is constructed as
Thus, $\widehat{\mu}_t(\beta)$ for $t = H+1, \ldots, T$ serves as the sample “observations” and $\widehat{\mu}(\beta)$ serves as the sample “mean.”
The second feasible inference procedure is based on a new plug-in (PI) variance estimator which can be constructed using the results in Lemma (ref) and Theorem (ref). To see the logic of our approach, first consider the (unrealistic) case where we observe $\varepsilon_{it}$ and $\mathbb{E}[f_t|\mathcal{G}_{t-1}]$. We can exploit the fact that the two leading terms shown in Theorem (ref) are uncorrelated and so the natural plug-in variance estimator is
and $\widehat{q}_{jt} = n_t^{-1} \sum_{i=1}^{n_t} \widehat{\Phi}_{i,j,t} $. Of course, $\tilde{\sigma}^2_{\mathtt{PI}}(\beta)$ is infeasible. As a feasible alternative, consider
Here $\widehat{\varepsilon}_{it} = R_{it} - \widehat{\mu}_t(\beta_{it}) $ and $\widehat{\mathbb{E}[f_t|\mathcal{G}_{t-1}]}$ is a feasible estimator of $\mathbb{E}[f_t|\mathcal{G}_{t-1}]$ using only information contained in $\mathcal{G}_{t-1}$. For example a parametric approach would lead to $\widehat{\mathbb{E}[f_t|\mathcal{G}_{t-1}]} = h_{t-1}( {\widehat{\vartheta}})$ for some $h_{t-1}(\vartheta)$ with $h_{t-1}( \widehat{\vartheta})$ a corresponding feasible estimate. As a simple example, in the case of an AR(1) we would obtain $h_{t-1}( {\vartheta}) = \vartheta_0 + \vartheta_1 f_{t-1}$, and we would estimate the parameters, $(\vartheta_0,\vartheta_1)$, recursively. The estimator, $\widehat{\mathbb{E}[f_t|\mathcal{G}_{t-1}]}$, may not be correctly specified but we will demonstrate in Theorem (ref) below that we can still conduct valid, albeit possibly conservative, inference.
Finally, it is important to emphasize that, just as for the beta-sorted portfolio estimator itself (see Section (ref)), both the FM and PI variance estimators require an accounting, at each time $t$, of which portfolio $\beta$ resides in.
Before stating the pertinent properties of the two variance estimators, it will be useful to define the following object
which is the sample variance of the SACER. In addition, define the following (infeasible) version of the FM variance estimator where the individual elements have had their smoothing bias removed,
We then have the following result.
Theorem (ref) characterizes the (differential) asymptotic properties of the PI and the FM variance estimators. The intuition behind these results is specific to each of the variance estimators. For the PI variance estimator, consistency can be achieved when the specification of $\widehat{\mathbb{E}[f_t|\mathcal{G}_{t-1}]}$ is correct. However, even if this estimator is misspecified $\widehat{\sigma}_{\mathtt{PI}}(\beta)$ will be biased upward only. This conservativeness arises by the standard minimum mean-square error property of the conditional expectation with $s_{NT}$ representing the contribution arising from the misspecification. For the FM variance estimator, we can understand the result from the following decomposition of the summands of the estimator:
It is the first term that plays the role of capturing the contribution of the variance from $\sigma^2_f(\cdot)+\sigma^2_\varepsilon(\cdot)$, while the second term captures the contribution of $\sigma^2_\mu(\cdot)$. The third term, instead, is asymptotically negligible. When $\beta = 0$, in order to use the FM variance estimator, the bias from each $\widehat{\mu}_t(\beta)$ must be removed. This is why we replace $\widehat{\sigma}^2_\mathtt{FM}(\beta)$ by $\widetilde{\sigma}^2_\mathtt{FM}(\beta)$. As discussed in Remark (ref), in practice for $\widehat{\sigma}^2_\mathtt{FM}(\beta)$ we de-bias the portfolio which contains $\beta=0$, so the associated FM variance estimator will inherit this adjustment.
We can use the results of Theorem (ref) as the basis for our feasible inference procedures.
Since the SACER is a random object, we can use the results in Corollary (ref) to form prediction intervals for $\bar{\mu}_T(\beta; H)$. With either choice of variance estimator, inference is asymptotically valid but may be conservative. When using the PI variance estimator we obtain asymptotic coverage of exactly $1-2\alpha$ when the conditional expectation is specified correctly, whereas for the FM variance estimator this occurs when $\mu_t(\beta)$ does not vary over time.
Although we have argued that the SACER should be the preferred estimand for the beta-sorted portfolio estimator, it is illuminating to compare the conditions required to obtain results for this alternative estimand. First note that the beta-sorted portfolio estimator is consistent for the PACER so long as Assumption (ref) holds (i.e., the PACER exists):
The first term is $o_\mathbb{P}(1)$ under the assumptions given in Section (ref) and the second term is $o_\mathbb{P}(1)$ if Assumption (ref) also holds.
Although consistency is ensured with only a mild strengthening of the assumptions necessary for our results for the SACER, asymptotic normality requires substantially more structure on the properties of the data. This is because, for the PACER, the martingale difference property no long holds and so our previous approach to obtaining asymptotic results cannot be used. By equation (ref) it follows that the variance when estimating the SACER is smaller (by the magnitude $\sigma^2_\mu(\beta)$; see Theorem (ref)) than when estimating the PACER with equality only when $\mu_t(\beta)$ is constant over time. Furthermore, by Theorem (ref) the FM variance estimator will consistently estimate this larger variance under appropriate regularity conditions.
Establishing asymptotic theory and feasible inference procedures for the PACER is challenging because of the highly nonlinear time series dependence and complicated nonparametric structure of the beta-sorted portfolio estimator, and therefore we leave this task for future work.
In our main results we have laid a foundational framework and associated assumptions which ensure consistency, asymptotic normality, and feasible inference for the beta-sorted portfolios estimator of the SACER. These results immediately imply valid procedures for joint inference over finitely many values of $\beta$. In this section we discuss how to conduct joint inference for the SACER over multiple beta values along with how to formally test for the presence of profitable high-low and “butterfly” trading strategies.
In order to construct joint intervals for multiple values of $\beta$ we require estimators of the covariance of $\widehat{\mu}(\beta_1)$ and $\widehat{\mu}(\beta_2)$ for $\beta_1,\beta_2 \in \mathcal{B}$. For the FM variance estimator we have,
whereas for the PI variance estimator we can use,
The results we present are valid for either the FM or PI approaches and so for simplicity we present the results with the generic notation $\widehat{\sigma}(\beta_1,\beta_2)$ as a stand-in for either estimator with $\widehat{\sigma}(\beta,\beta) = \widehat{\sigma}^2(\beta)$.
Finally, the following notation will be used throughout this section. Let $\mathcal{B}_G = \{b_1, b_2, \ldots, b_G \} \subset \mathcal{B}$ be a finite collection of $G$ evaluation points of $\beta$. Our theoretical results imply the existence of a $G$-variate Gaussian limiting distribution with the covariance matrix, $\Sigma_G$, following from Theorem (ref). More precisely, let $\widehat{\mu}_G = (\widehat{\mu}(b_1), \ldots, \widehat{\mu}(b_G))^\top$ and $\mu_G = (\mu(b_1), \ldots, \mu(b_G))^{\top}$, so that $(\widehat{\mu}_G - \mu_G) \sim_a \Sigma_G^{1/2} Z$ where $Z \sim \mathcal{N}(0,I_G)$.
A valid joint prediction interval for SACER can be constructed via $[\widehat{L}_T(\beta ), \widehat{U}_T(\beta )]$ where $\widehat{L}_T(\beta ) = \widehat{\mu}(\beta) -\widehat{\sigma}(\beta) \mathsf{q}_{1-\alpha}/{\sqrt{T}}$, and $\widehat{U}_T(\beta ) = \widehat{\mu}(\beta) + \widehat{\sigma}(\beta) \mathsf{q}_{1-\alpha}/{\sqrt{T}}$, and $\mathsf{q}_{1-\alpha}$ is obtained as the $1-\alpha$ quantile of the distribution of $\big| \Sigma_{G, \mathrm{corr}}^{1/2} Z \big| _\infty$ where $\Sigma_{G,\mathrm{corr}} = \mathrm{diag}(\Sigma_G)^{-1/2} \Sigma_G \mathrm{diag}(\Sigma_G)^{-1/2}$ and $\mathrm{diag}(A)$ denotes a diagonal matrix with diagonal entries the same as the diagonal entries of $A$. In practice, we replace $\Sigma_G$ with $\widehat{\Sigma}_G$ using either the FM or PI covariance estimator introduced above. Then, with a pre-specified coverage level $1>2\alpha>0$, \[\liminf_{N,T\to\infty} \mathbb{P}\Big(\bar{\mu}_T(\beta; H) \in [\widehat{L}_T(\beta ), \widehat{U}_T(\beta )], \text{ for all } \beta\in\mathcal{B}_G\Big) \geq 1- 2\alpha.\] Intuitively, if $0$ is not contained in this joint prediction interval then we can reject the null that there exists values of $\beta \in \mathcal{G}$ which do not earn (expected) returns for the exposure to the factor. We provide an algorithm to implement the joint inference procedure below.
The most popular object of interest in the empirical finance literature is to compare the time-average of returns from the two extreme portfolios (i.e., the portfolios which encompass the evaluation points $\beta_l$ and $\beta_u$) as discussed in Section (ref). The goal is to assess whether a long-short portfolio trading strategy earns statistically significant returns, i.e., has a nonzero unconditional risk premium. However, we can use our general framework and new theoretical results to formulate a more effective inference procedure to assess the properties of expected returns. Rather than focus on $|\mu(\beta_u) - \mu(\beta_l)|$ we instead consider $\max_{\beta\in\mathcal{B}_G}\mu(\beta) - \min_{\beta\in\mathcal{B}_G}\mu( \beta)$. In words, we study the most profitable long-short strategy available. We refer to this as the generalized high-minus low strategy. In the special case when $\mu(\beta)$ is monotonic (and the grid $\mathcal{B}_G$ includes $\beta_l$ and $\beta_u$), then the two expressions are equivalent. Thus, we nest the popular high minus low portfolio inference approach as we test for the presence of any profitable long-short strategy in $\mathcal{B}_G$.
This class of high-minus-low statistics can be re-expressed in the following form,
where we denote $\beta_{\max}$ as the point attaining $\max_{\beta\in\mathcal{B}_G} (\frac{1}{T}\sum_{t=1}^T\sum^{J_t}_{j=1}\widehat{p}_{j,t}(\beta)\widehat{a}_{jt})$ and $\beta_{\min}$ as the point attaining $\max_{\beta\in\mathcal{B}_G} (-\frac{1}{T}\sum_{t=1}^T\sum_{j=1}^{J_t}\widehat{p}_{j,t}(\beta)\widehat{a}_{jt}))$. Algorithm (ref) below shows how to construct a joint prediction interval $[\widehat{L}_T, \widehat{U}_T] $. Intuitively, if $0$ is outside of this interval then there exists at least one profitable long-short strategy across a pair of values in $\mathcal{B}_G$.
We introduce one final joint inference procedure which is motivated by a “butterfly” trading strategy corresponding to a discrete second derivative (see Remark (ref); the long-short trading strategy of the previous subsection can be thought of as a discrete first derivative). Furthermore, under the assumption of the absence of arbitrage opportunities, then the “butterfly” trading strategy we formulate should have zero (conditional) expected returns. Thus, we can also directly conduct inference on whether the data are consistent with a no-arbitrage assumption. In words, under no-arbitrage, there does not exist an ex-ante profitable trade where one goes long one unit of each of two assets (one with $\beta _{1}$ and one with $\beta _{3}$) and short two units of an asset (with $\beta _{2}$). Naturally, we will focus on the estimator: \[ \max_{\beta_1+\beta_3=2\beta_2} \textnormal{\textsc{Bfly}}(\beta_1,\beta_2,\beta_3; \widehat{\mu}) = \max_{\beta_1+\beta_3=2\beta_2} \frac{1}{T} \sum_{t=1}^T \textnormal{\textsc{Bfly}}_t(\beta_1,\beta_2,\beta_3; \widehat{\mu}_t).\]
To construct asymptotically valid prediction intervals we can utilize the results of Lemma (ref) which yields the following leading term expansion,
As we discussed in Remark (ref), the “butterfly” estimator enjoys a better rate of convergence for all values of $\beta$, namely, $\sqrt{nT}/\sqrt{J}$ and so we require an alternative variance estimator to form prediction intervals. The asymptotic variance of $T^{-1}\sum_{t=1}^T\sqrt{Tn_{t}/J_{t}}\cdot\textnormal{\textsc{Bfly}}_t(\beta_1,\beta_2,\beta_3; \widehat{\mu})$ may be estimated by
Let $\beta_{1,2,3}$ be an abbreviation for $\beta_1, \beta_2,\beta_3$ and $\beta_{1,2,3}'(\neq \beta_{1,2,3}) $ as an abbreviation for $\beta_1', \beta_2',\beta_3'$. We define the corresponding covariance estimator as
For an equal-spaced choice of gridpoints, $\mathcal{B}_G$, we can then collect all ordered triplets $(b_1,b_2,b_3)$ which satisfy $2b_2 = b_1 + b_3$ $(b_1\neq b_3)$. Then, we can construct the $G^\star \times G^\star$ estimated variance-covariance matrix $\widehat{\Sigma}_{\mathrm{DiD}}$ by using $\widehat{\sigma}_{\mathrm{DiD}}^2(b_1,b_2,b_3)$ for the diagonal elements and $\widehat{\sigma}_{\mathrm{DiD}}(b_{1,2,3},b_{1,2,3}')$ for the off-diagonal elements.\footnote{When $G$ is even, then $G^\star = G(G-2)\big/4$ and when $G$ is odd, then $G^\star = (G-1)^2\big/4$.}
Using Algorithm (ref) below we can construct the appropriate prediction interval. Intuitively, if zero is outside the prediction interval, then there exists at least one “butterfly” strategy based on $\mathcal{B}_G$ which earns nonzero expected returns. Furthermore, if zero is outside of the interval then we cannot reject the hypothesis that the no-arbitrage condition holds in the data.
In this section, we introduce a novel risk factor and show that it is strongly predictive of both the cross-section and time-series behavior of U.S. stock returns. We also utilize this application to illustrate the practical advantages of the novel theoretical results presented earlier in the paper.
Our risk factor is a new measure of the business credit cycle. The business credit cycle is commonly evaluated by means of ratios of credit aggregates to measures of output. Although theoretically appealing, a drawback to these approaches is that it is difficult to parse out movements in credit ratios that are arising from composition changes in the aggregates as compared to all other movements. Here we take a different approach. We rely on the Federal Reserve's Senior Loan Officer Opinion Survey\footnote{The properties of the SLOOS were first studied in SchreftOwens1991, LMR2000, and LM2002, LM2006. See also CrumpLuck2023.} (SLOOS) as our proxy for the “credit” portion of the ratio and the ISM Manufacturing Index as our measure of the “output” portion. This has three distinct advantages. First, as the SLOOS and ISM are both diffusion indices, they have uniform behavior across their history even in the face of changes in the structure of the economy. Second, they are much more timely than credit aggregates and national accounts data which tend to be released with a substantial lag. Third, they are not subject to revision. Thus we have a timely factor which we can evaluate in real time with no look-ahead bias.
Our factor is simply constructed as
where $\mathrm{SLOOS}_t$ is the net percentage of large domestic banks tightening standards for commercial and industrial loans to all firms and $\mathrm{ISM}_t$ is the ISM index.\footnote{The Senior Loan Officer Opinion Survey is currently conducted on a quarterly basis. To construct a monthly series we keep the SLOOS value constant until a new value is available. For the period from 1984m1 through 1990m1, the credit standards question was not included in the SLOOS. For this period we use as a replacement the net willingness to make consumer installment loans by large domestic banks.} Although both the SLOOS and the ISM are diffusion indices, they are scaled differently and so the affine transformation of the SLOOS is implemented so that they are both on the same scale (between 0 and 100). To understand why this is (the inverse of) a credit-to-output type measure note that a fall in the SLOOS corresponds to easier lending standards (higher credit growth) and a fall in the ISM to less output. Thus, when the CCW variable is low, credit-to-output is high. A similar logic applies for when the CCW variable is high. Our factor is available starting in January 1965 when the first SLOOS was implemented.
As a preliminary check for the validity of our factor we assess its ability to predict future market returns. Specifically, an implication of our setup (see equation ((ref))) is that if the factor is serially correlated then lagged values should be predictive of future equity returns. To show this, we consider the standard predictive regression setup and run predictive regressions of the form,
We utilize the standard predictors obtained from GoyalWelch2008 as a benchmark comparison along with our risk factor. In Table (ref) we present in-sample $R^2$ from predictive regressions for forecast horizons of 1, 3, 6, and 12 months ahead. The first fourteen rows present the results for the benchmark predictors investigated in GoyalWelch2008. The next row, labeled “CGP” reports results using only the SLOOS portion of our risk factor as in CGP2015. Finally, The last row, labeled “CCW” provides the results for our new risk factor. The results are stark. The in-sample $R^2$ from our new risk factor far outstrips that of the other predictors considered. To ensure our results are not a consequence of overfitting, in Table (ref) we present out-of-sample $R^2$ results using a training sample up to the end of 1989. Again, the results are stark with our risk factor outperforming each of the other predictors by a wide margin.
We can now investigate how our risk factor performs in explaining the cross-section of equity returns. We implement our estimators as described in Sections (ref) and (ref). We use monthly data from the Center for Research in Security Prices (CRSP) over the sample period January 1926 to December 2019. We restrict these data to those firms listed on the New York Stock Exchange (NYSE), American Stock Exchange (AMEX), or Nasdaq and use only returns on common shares (i.e., CRSP share code 10 or 11). To deal with delisting returns we follow the procedure described in BEM2016. When forming market equity we use quotes when closing prices are not available and set to missing all observations with $0$ shares outstanding. For our risk factor we use a measure of the business credit cycle described in equation ((ref)). We utilize five-year rolling regressions to estimate betas and we choose the number of portfolios as $J_t = J_1 \cdot (\frac{n_t}{\max_{1\leq t\leq T} n_t})^{\frac{1}{2}}$ where $J_1 = 10$. The latter choice can be motivated by appealing to cattaneo2020characteristic as the optimal choice of portfolios under the simplifying assumption that all $\beta_{it}$ were known.
Figure (ref) presents our estimate of the grand mean, $\widehat{\mu}(\beta)$ in the black line. There is a clear downward slope in the relationship between $\beta$ and expected returns -- although it does not appear to be linear. To address the differential behavior of the beta-sorted portfolio estimator for the portfolio containing $\beta=0$, we fit a linear regression in this portfolio only and use constant fits for all other portfolios. The grey vertical lines in Figure (ref) depict confidence intervals at each selected point in the support of $\beta$. The top chart in Figure (ref) uses the PI variance estimator we introduced in equation ((ref)) whereas the bottom chart uses the FM variance estimator. To implement our plug-in variance estimator we use an (expanding window) AR(1) specification in our risk factor. We can clearly see the difference in the precision for drawing inferences from the data. The confidence intervals based on our new PI variance estimator are substantially shorter than those of the FM variance estimator. Although we showed in Section (ref) that both of these estimators are conservative, in general, for inference on the SACER, the FM is much larger in this application. By Remark (ref) we can use the difference between the FM variance estimate and the PI variance estimate as an estimate of the lower bound on $\sigma^2_\mu(\beta)$. Inspecting Figure (ref) we see that this lower bound can be informative as the length of the 95% confidence interval based on the FM variance estimator are at least three times as long as that based on the PI variance estimator.
We can see the difference between the two variance estimators even more clearly in Table (ref) where we present the point estimate for selected values of $\beta$ along with lower and upper bounds for confidence intervals constructed with the two different variance estimators. The results are striking. Across all values of $\beta$ and for both nominal coverage rates, the confidence intervals formed using our plug-in variance estimator are approximately 30% of the length of those using the FM variance estimator. We can also illustrate the difference by considering confidence intervals on the high-minus-low estimator, $\widehat{\mu}(\beta_u) - \widehat{\mu}(\beta_l)$, which has a point estimate of $0.5$. When using the PI variance estimator the 95% confidence interval is $[0.14,0.86]$ whereas for the FM variance estimator it lengthens to $[-0.69, 1.69]$.
Beta-sorted portfolios are a commonly used empirical tool in asset pricing. In a first step, time-varying factor exposures are estimated by weighted regressions of asset returns on an observable risk factor to ascertain how returns co-move with the variable of interest. In a second step, individual assets are grouped into portfolios by similar factor exposures and differential returns are assessed as a function of differential exposures. Yet the simple and intuitively appealing algorithm belies a more complicated statistical setting involving a two-step estimation procedure where each stage involves nonparametric estimation.
We provide a comprehensive statistical framework which rationalizes this commonly-used estimator. Armed with this foundation we study the theoretical properties of beta-sorted portfolios linking directly to the choice of estimation window in the first step and the number of portfolios in the second step which serves as the tuning parameters for each nonparametric estimator. We introduce conditions that ensure consistency and asymptotic normality for a single cross-section and for the grand mean estimator. We also introduce a new variance estimator and characterize the properties of the ubiquitous FM variance estimator. We provide joint inference procedures for different hypotheses of interest in financial applications including tests of the no-arbitrage assumption. Finally, we also discover some limitations of current practices and provide new guidance on appropriate implementation and interpretation of empirical results.
\setcounter{equation}{0}
Define $b_{it_0} = (\alpha_{it_0},\beta_{it_0})^{\top} = \mathbb{E}(X_{t_0} X_{t_0}^{\top}|\mathcal{F}_{t_0-1})^{-1}\mathbb{E}(X_{t_0} R_{it_0}|\mathcal{F}_{t_0-1})$. We consider a slightly generalized first step estimator that allows for (one-sided) kernel weighting:
where $H = \lfloor Th\rfloor$, $X_t=(1,f_t^\top)^\top$, and $K(.)$ is a kernel function satisfying the following assumption.
For a $(m\times n)$-dimensional matrix $A = (a_{i j})_{1\le i\le m, 1\le j \le n}$, we define the induced matrix norms $|A|_1 = \max_{1\le j \le n} |\sum_{i=1}^m a_{i,j}|$, $|A|_2 = \max_{|v| = 1} |A v|_2$, and $|A|_{\infty} = \max_{1\le i\le m}|\sum_{j=1}^n a_{i,j}|$.
To save notation, we define the following quantities:
and
The following theorem is a generalization of Theorem (ref) in the paper.
The rate condition $T^{2/q-1} n^{2/q} / h \to 0$ implies that $\frac{T^{1/q}}{Th} \lesssim \sqrt{\frac{\log T}{Th}}$, and it favors higher moments. For example, it implies $\frac{n^{1/4}}{T^{3/4}h} \to 0$ when $q=8$.
We provide an asymptotic normality result for $\widehat{b}_{it_0}$, assuming conditional homoskedasticity (but possibly time-varying) for simplicity. This result is not used in the main paper, but is reported here for completeness.
A plug-in consistent estimator of the asymptotic variance can be constructed using estimated residuals:
where $\widehat{u}_{it}$ denote the estimated residuals.
{We assume that $d=1$ without loss of generality.} For the first result, we have
where
In addition, Assumption (ref) implies that (i) $\{z_{t-s}^2-\mathbb{E}[z_{t-s}^2]\}_{s}$, $\{z_{t-s}\tau\big(\frac{t - s}{T}\big)\}_{s}$, and $z_t$ are mean zero and square-integrable, and (ii) $\Theta(z^2_{\bullet};q,v) \lesssim 1$ with $v > 1/2-2/q$. Therefore, Lemma A.3 in zhang2012inference implies that $\mathfrak{R}_1 + \mathfrak{R}_2 + \mathfrak{R}_3 \lesssim_{\mathbb{P}} \frac{T^{1/q}}{Th} + \sqrt{\frac{\log T}{Th}}$.
For the second result, we have
where
These terms form a martingale difference sequence. Thus, we let $u_s$ be equal to either $z_{t_0-s}^2-\mathbb{E}[z_{t_0-s}^2]$, $(z_{t_0-s}-\mathbb{E}[z_{t_0-s}|\mathcal{F}_{t_0-s-1}])\tau\big(\frac{t_0 - s}{T}\big)$, or $z_{t_0-s}-\mathbb{E}[z_{t_0-s}|\mathcal{F}_{t_0-s-1}]$, to save notation. Then, using summation by part
because
by the Lipschitz continuity of the kernel function. Let $\lambda$ be a sufficient large positive constant. Then, hall2014martingale implies that
Thus, it suffices to look at $H $ blocks of observations. We have
and applying Freedman's inequality freedman1975tail and the union bound, we can verify that $\mathfrak{R}_4 + \mathfrak{R}_5 + \mathfrak{R}_6 \lesssim_{\mathbb{P}} \frac{T^{1/q}}{Th} + \sqrt{\frac{\log T}{Th}}$.
The last two conclusions follow analogously, using martingale methods and a union bound over $i$. For example, consider the fourth conclusion. We have
The term $\sum_{s=1}^H K\big(\frac{s}{Th}\big) X_{t_0-s} \varepsilon_{it_0-s}$ a martingale difference sequence, and therefore proceeding as above we verify the desired result. The second term of the upper bound in the preceding display is bounded similarly. \qed
We have $\widehat{b}_{it_0}-b_{it_0} = A(t_0)^{-1}B_i(t_0)-\tilde{A}(t_0)^{-1}B_i(t_0)+\tilde{A}(t_0)^{-1}B_i(t_0)-\tilde{A}(t_0)^{-1}\tilde{A}(t_0)b_{it_0}$, and therefore
where
because $\max_{\lfloor Th \rfloor+ 1\leq t_0\leq T} \big| \tilde{A}(t_0)^{-1}\big|_{\infty} \lesssim_\mathbb{P} 1$.
For the first term, we have
For the second term, we have
This completes the proof.\qed
From previous results, the bias is
and therefore we have the decomposition:
with
Similar to the proof of Theorem (ref), we have
under the rate conditions imposed.
Next, observe that
because
Therefore, putting the results above together and proceeding as before, we have
Furthermore, using prior results, it follows that
uniformly over $t_0$ and $i$. The stochastic linear approximation on the right-hand-side of the equal sign is a martingale difference sequence, and thus the proof can now be easily completed by applying Corollary 3.1 in hall2014martingale. \qed
This section presents the proofs of the main results reported in the paper. It also provides additional results that are either discussed heuristically in the paper or are not given there to streamline the presentation.
Let $k_{jt} = \lfloor n_t j/J_t \rfloor$ and $\kappa_{j,t} =j/J_t$. Recall
and
The following lemma presents some basic properties of the estimated quantiles and the resulting partitioning scheme.
The next lemma controls the convergence of the Gram matrix, score vector, and other related quantities underlying our estimator. Recall that $Q_t$ is a diagonal matrix with elements $\{q_{jt}: j=1,\ldots,J_t\}$.
Finally, the following lemma gives the cross-sectional convergence rate of $\widehat{\mu}_t(\beta)$ to $M_t(\beta)$.
We begin with the elementary decomposition
where
and
with
where $a_{t}^\circ = (\mathbb{E}[ \Phi_{i,t} \Phi_{i,t}^{\top} | \mathcal{G}_{t-1} ])^{-1} \mathbb{E}[ \Phi_{i,t} R_{it} | \mathcal{G}_{t-1} ]$.
By previous results, the term $(\widehat{\Phi}_t \widehat{\Phi}_t^{\top}/n_t )^{-1}$ exists and is finite with probability approaching one, and on that event, $\mathbb{E}[\mathfrak{R}_1(\beta)]=0$. Thus, on that event, using the martingale structure,
Proceeding analogously, for the second term, we verify
because $J\mathsf{L}_{nT}\to 0$.
For the third term, first consider the case when $\beta\neq0$. Proceeding as above, we have
When $\beta=0$, we obtain the faster upper bound
because $J\mathsf{L}_{nT} \to 0$.
For the fourth term, first consider the case when $\beta\neq0$. Using the same logic as before,
since $J\mathsf{L}_{nT}\to0$. Likewise, when $\beta=0$,
because $J\mathsf{L}_{nT} \to 0$.
Finally, for the bias term, we verify that $\mathscr{B}(\beta) \lesssim_\mathbb{P} J^{-1}$. Define
where $\mu_{t}(\beta_{t}) = (\mu_{t}(\beta_{1t}), \ldots, \mu_{t}(\beta_{n_tt}))'$. Then,
where the last line follows by Lemma (ref), Bernstein's inequality and since $\big| \tilde{a}_{t}^{\circ} \big|_\infty$ is bounded on the event that $({\Phi}_t {\Phi}_t^{ \top}/n_t )^{-1}$ exists and is finite which occurs with probability approaching one. Therefore, $\mathscr{B}(\beta) \lesssim \mathscr{B}_1(\beta) + \mathscr{B}_2(\beta)$ where
For the first term $\mathscr{B}_1(\beta)$ we have
The third term is $O_\mathbb{P}(J^{-1})$ by above calculations and because $|\widehat{p}_{t}(\beta)|_\infty=1$. Next note that the first term is an average of $\mu_t(\beta_{it})$ in each $\widehat{P}_{jt}$ which contains $\beta$ whereas the second term is an average of $\mu_t(\beta_{it})$ in each $P_{jt}$ which contains $\beta$ where $P_{jt}$ are the portfolios constructed using the sample quantiles of $\beta_{it}$. Thus, the first two terms are bounded by
by our smoothness assumptions on $\mu_t(\cdot)$ and by Lemma (ref). The second term $\mathscr{B}_2(\beta)$ follows by similar steps.\qed
By Lemma (ref), we have
where
forms a martingale difference sequence adapted to the filtration $\mathcal{F}_{t-1}$, and because
and
Thus, the proof is completed by employing the martingale central limit theorem of hall2014martingale, whose conditions are implied by the following two conditions:
and
For the first condition (ref), we have
where
For the first term we have
and, similarly, $\mathfrak{R}_2 \lesssim_\mathbb{P} \frac{1}{T}$ for both cases of interest ($\beta\neq0$ and $\beta=0$).
For the second condition (ref), first define
and because $\mathsf{Z}_{t}(\beta)$ is mean-zero and adapted to the filtration $\mathcal{F}_t$ and therefore
Then, we have
because, for $k\geq0$, we have
and therefore
where
and
To complete the proof, observe that under our imposed assumptions,
and
Then the conclusion follows. \qed
It will be convenient to define the rate $r_{n,T}(\beta) = \frac{J}{Tn}+\frac{1}{TJ^2}$ for $\beta = 0$ and $r_{n,T}(\beta) = \frac{1}{T}$ for $\beta\neq 0.$ We start with the result when $\beta \neq 0$,
where
For $\mathfrak{R}_1$, we can follow similar steps as for the proof of Lemma (ref) to obtain that
Next, note that $\mathfrak{R}_2=o_\mathbb{P}(\frac{1}{TJ} + \frac{1}{T^2})$ by Lemma (ref) and Theorem (ref) so that $\mathfrak{R}_2=o_\mathbb{P}\Big(\frac{1}{T}\Big)$. Finally, for $\mathfrak{R}_3$ we have that
By similar steps as in the proof of Lemma (ref), and using the fact that $| \mu_t(\beta) |$ is bounded, the first term is $O_\mathbb{P}(\frac{1}{TJ} + \frac{\sqrt{r_{n,T}(\beta)}}{T})$ and by Lemma (ref) and Theorem (ref) the second term is $O_\mathbb{P}(\frac{\sqrt{r_{n,T}(\beta)}}{T})$. Thus, $\mathfrak{R}_3 = o_\mathbb{P}\Big(\frac{1}{T}\Big) $, and the result for $\beta \neq 0$ follows.
Now consider the $\beta=0$ case. We have that,
where
For $\mathfrak{R}_1$, we can follow similar steps as for the proof of Lemma (ref) to obtain that $\mathfrak{R}_1 = o_\mathbb{P}(r_{n,T}(\beta))$. Next, note that $\mathfrak{R}_2=o_\mathbb{P}(r_{n,T}(\beta))$ directly by Lemma (ref) and Theorem (ref). Finally, for $\mathfrak{R}_3$, using the Cauchy–Schwarz inequality, we have
By the same steps as for $\mathfrak{R}_1$ we have that the first factor is $O_\mathbb{P}(r_{n,T}(\beta))$. By assumption, we have that ${\sigma}^2_\mu(\beta) = o_\mathbb{P}(r_{n,T}(\beta))$ and so $\mathfrak{R}_3 = o_\mathbb{P}(r_{n,T}(\beta))$. \qed
Recall that we define the rate $r_{n,T}(\beta) = \frac{J}{Tn}+\frac{1}{TJ^2}$ for $\beta = 0$ and $r_{n,T}(\beta) = \frac{1}{T}$ for $\beta\neq 0$ from the Proof of Theorem (ref)$(i)$. We first decompose $\widehat{\sigma}^2_{f,\mathtt{PI}}(\beta)$ as
where
and
Clearly, $s_{nT}(\beta) \geq 0$ for all $\beta \in \mathcal{B}$. For $\mathfrak{R}(\beta)$, note that the summands form a martingale difference sequence with respect to $\mathcal{F}_{t-1}$ so that, when $\beta\neq 0$,
When $\beta = 0$, we can follow similar steps and obtain that $\mathbb{V}\mathrm{ar}[\mathfrak{R}(\beta)]=o({r_{n,T}(\beta)^2})$ using Lemma (ref). Finally, we need only show that $\big| \widetilde{\sigma}^2_{f,\mathtt{PI}}(\beta) - \sigma^2_f(\beta) \big| = o_\mathbb{P}({r_{n,T}(\beta)})$ which holds under the conditions of Theorem (ref) and given the results in Lemma (ref) to Lemma (ref).
We next prove that $|\widehat{\sigma}^2_{\varepsilon,\mathtt{PI}}(\beta) - \widetilde{\sigma}^2_{\varepsilon,\mathtt{PI}}(\beta)|= o_\mathbb{P}({r_{n,T}(\beta)})$. We have that,
where
For the first term,
For the second term,
For the last two terms we have
Finally, we need only show that $\big| \widetilde{\sigma}^2_{\varepsilon,\mathtt{PI}}(\beta) - \sigma^2_\varepsilon(\beta) \big| = o_\mathbb{P}({r_{n,T}(\beta)})$. The above statement holds under the conditions of Theorem (ref) and given the results in Lemma (ref) to Lemma (ref). This completes the proof. \qed
For the first result, note that by Theorem (ref) and the assumptions imposed,
where
Note that $\mathfrak{R}_{1,t}\lesssim \mathfrak{R}_{11,t} + \mathfrak{R}_{12,t} $ with
and it follows by standard arguments that
Next, we have
Finally, for $\mathfrak{R}_{3,t}$, we proceed as for $\mathfrak{R}_{1,t}$ but taking into account the first-step estimation. Thus, $\mathfrak{R}_{3,t} \lesssim \mathfrak{R}_{31,t} + \mathfrak{R}_{32,t} + \mathfrak{R}_{33,t}$ with
with arbitrary high probability for $n$ and $T$ large enough, uniformly in $t$, and where $\mathcal{B}=[\beta_l,\beta_u]$. We set $f_{\beta,t}(F_{\widehat{\beta},t}^{-1}(\kappa_{j,t}))=f_{\beta,t}(\beta_l)$ if $F_{\widehat{\beta},t}^{-1}(\kappa_{j,t})<\beta_l$, and $f_{\beta,t}(F_{\widehat{\beta},t}^{-1}(\kappa_{j,t}))=f_{\beta,t}(\beta_u)$ if $F_{\widehat{\beta},t}^{-1}(\kappa_{j,t})>\beta_u$. Recall also that we set $\widehat{\beta}_{(0)t} = \beta_l$ and $\widehat{\beta}_{(n_t)t} = \beta_u$ for simplicity.
For $\mathfrak{R}_{31,t}$, using standard results, for any $c\in\mathbb{R}$, we have
where $Z_{n,\beta,t}(y) = \sqrt{n_{t}}(F_{\beta,n,t}(y)-F_{\beta,t}(y))$. The first term in the upper bound vanishes due to the modulus of continuity of the empirical distribution function stute1982oscillation, the uniform inequality in massart1990tight with the union bound, and our imposed rate conditions. For the second term in the upper bound, first note that
where $\tilde{c}$ is some point between $\beta_{it} - \widehat{\beta}_{it} + F_{\widehat{\beta},t}^{-1}(\kappa_{j,t}) + c/{\sqrt{n_t}}$ and $\beta_{it} - \widehat{\beta}_{it} + F_{\widehat{\beta},t}^{-1}(\kappa_{j,t})$. Therefore,
Therefore, we have
Similarly, for $\mathfrak{R}_{2,t}$ and $\mathfrak{R}_{3,t}$, by the modulus of continuity of the empirical distribution function stute1982oscillation, Assumption (ref), and the uniform inequality in massart1990tight with the union bound, we obtain
and
For the upper bound of the second result, define $U_{(k_{(j-1)t}),t} = F_{\beta,t}(\beta_{(k_{(j-1)t}),t})$, and note that $U_{(k_{jt}),t} - U_{(k_{j-1 t}),t}$ follows a Beta distribution with parameter $(k_{jt}- k_{j-1 t}, n_t+1 - (k_{jt}- k_{j-1 t}))$ conditional on $\mathcal{G}_{t-1}$, and thus $\mathbb{E}[U_{(k_{jt}),t} - U_{(k_{j-1 t}),t}|\mathcal{G}_{t-1}]= (k_{jt} -k_{j-1 t})/(n_t+1)$. Employing a Taylor series expansion of $F^{-1}_{\beta,t}(b)$ and following B.10 in bobkov2001hypercontractivity, we verify
for a positive constant $C$. It follows that the upper bound in the last display goes to $0$ under the rate conditions imposed. The lower bound in the second result of the lemma is proven similarly using the results in skorski2023bernstein.
For the third results, it follows the same steps as for $\max_{1\leq j\leq J_t-1}\big|\widehat{\beta}_{(k_{jt}), t} - {\beta}_{(k_{jt}), t} \big|$ except for a additional differencing step:
We shall break to the following term,
Similar to the previous step, we have $ \max_{\lfloor Th \rfloor + 1 \leq t\leq T} \mathfrak{R}_{1,t} \lesssim_p \sqrt{\frac{\log(nT)}{nJ}}$.
Next, $$ \max_{\lfloor Th \rfloor + 1 \leq t\leq T} \mathfrak{R}_{2,t} \lesssim \mathsf{R}_{nT}/J.$$
Then,
Thus the conclusion holds.
For the next result,
where
For the term $\mathfrak{R}_{4,t}$,
where
and
To see the above result, notice that
and employing standard empirical process theory we also verify that
Finally, for the term $\mathfrak{R}_{5,t}$, we have
where
and
and the proof is completed using the same logic as before.
Finally, the last result follows by Bernstein's inequality and standard calculations. \qed
For the first result, we have
For the second result, using the union bound, Markov's inequality, the conditional on $\mathcal{F}_{t-1}, f_t$ i.i.d. property of $\varepsilon_{it}$, and Bernstein's inequality,
where $\vartheta_{j,t} =\max_{1\leq i\leq n_t} \mathbb{E}[\Phi_{i,j,t}^2 \varepsilon_{it}^2|\mathcal{F}_{t-1},f_t] \lesssim_\mathbb{P} J^{-1}$. Thus, setting $M=\sqrt{nJ^{-1}/\log(nT)}$ and $u=C\sqrt{\frac{\log(nT)}{nJ}}$ for $C$ large enough, the result follows.
The third result is proven similarly using $ \max_{1\leq i\leq n_t} \mathbb{E}[(\widehat{\Phi}_{i,j,t} -\Phi_{i,j,t})^2 \varepsilon_{it}^2|\mathcal{F}_{t-1},f_t] \lesssim_\mathbb{P} \mathsf{L}_{nT}$.
For the fourth result,
using Bernstein's inequality conditional on $\mathcal{G}_{t-1}$.
The fifth result follows similarly to the third result.
For the sixth result,
The last result follows similarly to the proof of the previous terms. \qed
We have
where
with $a_{t}^\circ = (\mathbb{E}[ \Phi_{i,t} \Phi_{i,t}^{\top} | \mathcal{G}_{t-1} ])^{-1} \mathbb{E}[ \Phi_{i,t} R_{it} | \mathcal{G}_{t-1} ]$, and $\mathscr{R}_t(\beta) = \mathfrak{R}_{1t}(\beta) + \mathfrak{R}_{2t}(\beta) + \mathfrak{R}_{3t}(\beta) + \mathfrak{R}_{4t}(\beta)$ with
Proceeding as in the proof of Lemma (ref),
and
This completes the proof.\qed
\singlespacing
\begingroup \endgroup