Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
111,170 characters · 21 sections · 84 citation commands
Stochastic Discount Factors with Cross-Asset Spillovers
\thispagestyle{empty}
\setcounter{page}{1}
The central objective of empirical asset pricing is to identify firm-level signals that explain the cross-section of expected stock returns—whether through exposure to risk factors or persistent mispricing. The dominant paradigm, grounded in the assumption of self-predictability, asserts that a firm’s own characteristics forecast its own returns (see, e.g., cochrane2011presidential, harvey2016and). Complementing this view is a growing literature on cross-predictability—the idea that the characteristics or returns of one asset can help forecast the returns of others (see, e.g., lo1990contrarian, hou2007industry, cohen2008economic, cohen2012complicated, huang2021psychological, huang2022frog). A key mechanism underpinning this phenomenon is the presence of lead–lag effects, whereby price movements or information from one firm precede and predict those of related firms. Such effects can stem from staggered information diffusion, peer influence within industries, supply chain linkages, or correlated trading by institutional investors that induces price pressure across related assets.
Despite recent methodological advances in modeling cross-stock predictability, several foundational questions remain unresolved. Chief among them is how a mean–variance investor can analytically integrate multiple predictive signals when returns are interconnected across assets. Equally crucial is developing a framework that jointly captures both the relevance of individual signals and the structure of return spillovers—enhancing portfolio performance while preserving interpretability.
This paper addresses these questions by proposing a unified and systematic framework for constructing maximum–Sharpe ratio strategies. We combine firm-level signals through a flexible weighting vector (the signal-aggregation vector $\Lambda$) and model cross-asset spillovers using a structured connection matrix (the spillover matrix $\Psi$). The resulting optimal strategy admits a transparent analytical characterization. This formulation naturally connects to the stochastic discount factor (SDF; see hansen1991implications, cochrane2009asset, back2017asset), which, in this context, takes the form of a single factor that prices the cross-section of returns.
An important distinction in the asset pricing literature lies between conditional and unconditional Sharpe ratio optimization. As emphasized by hansen1987role, conditional optimization targets the best return–risk trade-off at each point in time using the information then available, whereas unconditional optimization maximizes this trade-off in expectation using long-run moments.\footnote{See also lewellen2006, who emphasize the distinction between conditional and unconditional beta pricing.} Our framework follows the latter approach: while it incorporates time-varying signals—such as firm characteristics and cross-asset linkages—the stochastic discount factor is optimized to perform well on average over time. This orientation prioritizes long-horizon performance over period-by-period efficiency, yielding strategies that are transparent, robust, and empirically grounded.
While the analytical formulation provides a population-level characterization of the Sharpe-optimal SDF, our empirical implementation uses a regression-based procedure tailored for high-dimensional applications. We build on the approach of britten1999sampling and employ ridge-type regularization—with a single tuning parameter $\lambda$ chosen by five-fold cross-validation—to estimate both the signal weights and the connection matrix. This method converges to the theoretical solution in large samples while enhancing numerical stability and interpretability. Unlike expected return-maximization—which, under certain specifications, can lead to extreme concentration in a single predictor—Sharpe ratio-maximization encourages diversification across signals, thereby enhancing robustness and practical relevance.
To build intuition, we start with a low-dimensional toy example using five well-known firm characteristics and nine portfolios sorted by size and book-to-market. This simplified setting enables us to illustrate the estimated signal weights, cross-asset linkages, and resulting trading strategy in full detail. We evaluate performance with a rolling out-of-sample procedure, re-estimating the strategy each month using the prior 10 years of data. Even in this controlled environment, the maximum–Sharpe ratio strategy based on cross-stock predictability attains an annualized Sharpe ratio of 1.22, compared with 0.60 for the self-predictive benchmark—an improvement driven simultaneously by cross-asset spillovers, shifts in signal relevance, and their interaction.
We then scale the framework to a comprehensive empirical setting using 138 firm-level signals from the jensen2023there dataset. Our primary investment universe consists of 138 univariate spread portfolios spanning 1963–2023. We also consider a broader set of 544 bivariate portfolios sorted by firm size and a secondary characteristic. Applying the same rolling 10-year estimation scheme, the maximum–Sharpe ratio (MS) strategy attains annualized Sharpe ratios of 2.21 and 3.32 on the spread and bi-sort portfolios, respectively—consistently outperforming both self-predictive benchmarks and maximum-expected return (MR) strategies. Specifically, the Sharpe ratio of our cross-predictive SDF strategy exceeds that of a self-predictive Sharpe ratio–maximizing benchmark by 0.79 on spread portfolios and more than 1.26 on bi-sorted portfolios—translating into economically meaningful gains in certainty-equivalent returns. Moreover, compared to expected return–maximizing strategies, our Sharpe ratio–maximizing SDF improves risk-adjusted performance by factors of 4–10, depending on the investment universe and market regime.
To assess robustness, we evaluate performance across different market environments. We split the test sample by investor sentiment and by volatility regimes based on the VIX index. The Sharpe ratio–maximizing strategy maintains strong performance across all subsamples. For example, in high-sentiment periods, the strategy delivers a Sharpe ratio of 2.19 on spread portfolios and 3.58 on bi-sort portfolios. Even in low-sentiment or high-volatility regimes—conditions that typically challenge individual anomaly-based strategies—the strategy sustains Sharpe ratios above 2. These results contrast with the more state-dependent performance of expected return–maximizing portfolios.
The SDF defines a single factor that, ex ante, prices the cross-sectional variation in expected returns of the test assets. We evaluate whether this factor's payoffs are priced by leading asset pricing models and find sizable, statistically significant alphas relative to a broad set of benchmarks. These include the liquidity factor pastor2003, the Fama–French five-factor model fama2015five, the q-factors hou2015, the mispricing factors stambaugh2017, the behavioral factors daniel2020short, and a comprehensive fourteen-factor model. Across all specifications, the strategy delivers alphas of about 0.25% per month with $t$-statistics above 11, indicating that the return variation embedded in cross-asset spillovers is not captured by existing models.
Upon optimizing the Sharpe ratio, we uncover the underlying economic drivers of return predictability. By examining the estimated weights assigned to firm-level characteristics, we find that the most influential predictors cluster in the categories of investment, value, and profitability, with signals such as liquidity of book assets, dividend yield, and return on equity consistently receiving the highest weights. In contrast, return-based signals—including momentum, short-term reversal, and seasonality—exhibit persistently low weights. This pattern suggests that the cross-predictive SDF is primarily anchored in stable firm fundamentals rather than transitory market signals.
In optimizing the Sharpe ratio, we also obtain a connection matrix, denoted by $\Psi$, that encodes the predictive relationships across stocks. Each entry $\Psi_{i,j}$ reflects the extent to which signals from asset $i$ forecast the returns of asset $j$, while diagonal elements represent self-predictive strength. Empirically, the average off-diagonal entry is substantial—often exceeding the average diagonal—indicating that cross-asset predictive linkages carry more information than self-predictive signals alone. Aggregating rows and columns of the matrix following diebold2014network, we uncover a directional structure: certain stocks consistently act as net transmitters of predictive signals, while others serve primarily as net receivers. Transmitters are typically large and low-turnover, whereas receivers tend to be smaller, high-turnover stocks with characteristics such as value orientation, high profitability, low investment activity, and strong past returns.
It is worth noting that the Sharpe ratio of the cross-predictive strategy is time-varying and declines notably after 2000. In the 1990s, the strategy delivers exceptional performance, with Sharpe ratios exceeding 2 on spread portfolios and above 4 on bi-sort portfolios. However, performance attenuates in the post-2000 period, mirroring the broader decline in self-predictability. For instance, green2017characteristics document that many anomaly portfolios become less profitable after 2003, attributing the decline to the widespread adoption of anomaly-based strategies, improved market liquidity, and the growth of passive ETF investing.
Despite this attenuation, the proposed strategy maintains strong performance from 2000 to 2023, achieving Sharpe ratios of 1.58 (spread portfolios) and 2.21 (bi-sort portfolios)—substantially higher than those of standard benchmark factors: 0.41 (market), 0.27 (size), 0.20 (value), 0.54 (profitability), 0.43 (investment), and 0.09 (momentum). By the end of 2023, five-year trailing Sharpe ratios decline to approximately 1.2 for the spread and bi-sort strategies, yet both remain consistently superior to traditional benchmarks even in recent years.
The paper proceeds as follows. Section (ref) presents the econometric framework. Section (ref) outlines the estimation methodology. Section (ref) describes the data. Section (ref) reports the empirical findings. Section (ref) concludes.
We consider an investment universe consisting of \( N \) risky assets. At each time \( t \), the investor observes a signal matrix \( S_t \in \mathbb{R}^{N \times M} \), where each row corresponds to one asset and contains \( M \) predictive characteristics (e.g., size, valuation, profitability, investment, past returns). Each column of \( S_t \) is cross-sectionally standardized to have zero mean and unit variance. Although our framework allows for a time-varying number of assets, the empirical analysis focuses on a fixed cross-section of sorted portfolios. We define \( t = 1 \) as the first period in which signals are observed, and \( t = T + 1 \) as the final period in which asset returns are realized.
A linear strategy that incorporates multiple signals and cross-predictability is specified as
where \( \omega_t \in \mathbb{R}^N \) denotes the portfolio weights, \( \Lambda \in \mathbb{R}^M \) assigns loadings to each signal, and \( \Psi \in \mathbb{R}^{N \times N} \) encodes how signals from one asset influence positions across all assets. Specifically, the weight on asset \( i \) is determined by multiplying \( \Lambda' \), \( S_t' \), and the \( i \)th column of \( \Psi \), allowing all signals in \( S_t \) to contribute to each asset’s position. The element \( \Psi_{i,j} \) quantifies the predictive impact of asset \( i \)’s signals on asset \( j \).
Relative to brandt2009parametric, who model portfolio weights as a function of firm-specific attributes, our framework generalizes the approach by allowing economically meaningful cross-asset spillovers to shape portfolio allocations. Moreover, although we focus on linear strategies, the framework readily accommodates nonlinear extensions by enriching the signal matrix with polynomial or Fourier-based transformations. For instance, one can construct an expanded signal matrix of dimension \( N \times MP \), where the first \( N \times M \) block corresponds to the original \( S_t \), the second to its elementwise square, and subsequent blocks to higher-order transformations up to the \( P \)th power. Importantly, such extensions preserve the dimension of the $\Psi$ matrix, while the $\Lambda$ vector expands accordingly to accommodate the enlarged set of predictors—including higher-order powers of the original signals. We leave the formal development and empirical implementation of such nonlinear extensions to future research.
We construct managed-portfolio returns in excess of the risk-free rate by interacting future returns with the current values of predictive signals:
where $\Pi_s$ is an $N^2 \times M$ matrix of managed-portfolio returns, $I_N$ is the $N \times N$ identity matrix, $r_s$ is a vector of $N$ excess returns realized at time $s > t$, and $\otimes$ denotes the Kronecker product.
The expected returns on these managed-portfolios are then defined as $ \Pi = E\bigl[\Pi_s\bigr]$. Additionally, define
so that \(\Phi\in\mathbb{R}^{N^2}\). The vectorized \(\Phi\) and the matrix $\Pi$ streamline later expressions for portfolio outcomes.
To aid interpretation, limit extreme equity positions, and stabilize estimation, we impose Euclidean norm constraints on key parameters. Specifically, we set
where the Euclidean norm constraint on the vector $\Phi$ is equivalent to a Frobenius norm constraint on the matrix $\Psi$. From a Bayesian perspective, these constraints correspond to zero-mean Gaussian priors on $\Lambda$ and $\Phi$, inducing ridge-type regularization that penalizes large parameter values.
Proposition (ref) formulates the realized return of the strategy in a convenient form, along with the expected return and Sharpe ratio. Appendix (ref) provides the proof.
We offer several remarks regarding Proposition (ref).
First, our empirical analysis primarily focuses on maximizing the Sharpe ratio, using expected return-maximization as a benchmark for comparison. While both objectives rely on the same expressions for expected returns, they lead to different optimal estimates for the signal weight vector $\Lambda$ and the vectorized connection matrix $\Psi$. In particular, expected return-maximization reduces to a bilinear optimization problem with closed-form solutions, whereas Sharpe ratio-maximization entails solving a generalized eigenvalue problem via an iterative procedure.
Importantly, maximizing the squared Sharpe ratio necessitates the use of both representations of the Sharpe ratio provided in Proposition (ref) when estimating the optimal values of $\Lambda$ and $\Phi$. Explicit solutions for both the expected return and Sharpe ratio-maximization problems are presented later in the paper.
Second, Proposition (ref) makes extensive use of the vectorized form of $\Psi$, which fully retains the cross-predictive structure embedded in $\Psi$. As a result, the information content relevant for cross-predictability is entirely preserved in $\Phi$, ensuring that the resulting strategy remains grounded in the same underlying predictive relationships.
Third, the expression for investment return offers an intuitive economic interpretation of our trading strategy. Recall that $\Pi$ denotes the matrix of managed-portfolio expected returns, with each of its $N^2$ rows representing the expected value of one asset’s return multiplied by one of the $M$ signals across the $N$ assets. Under the normalization $\mathbb{E}[S_t]=0$, $\Pi$ simplifies to the covariance matrix between future asset returns and contemporaneous signal values. If characteristic $m$ of stock $j$ helps predict the future return of stock $i$, the corresponding element of $\Pi$ will be nonzero, reflecting this predictability.
Thus, in this framework, $\Lambda$ assigns relative weights to signals, $\Phi$ encodes cross-asset interactions, and together they operate on the matrix $\Pi$ to optimize investment metrics.
Fourth, the expected return of the trading strategy can alternatively be expressed as
where \(\mu_m = \sum_{p=1}^{N^2} \Pi_{pm} \Phi_p\) represents a weighted combination of portfolio expected returns, with \(\Pi_{pm}\) denoting the expected return of the corresponding managed-portfolio and \(\Phi_p\) capturing the strength of the \(p\)-th relationship within the strategy.
This expected-return expression is informative because it demonstrates that, whether subject to an $L_1$ constraint or left unconstrained, the optimal solution is a corner solution: the trading strategy is entirely driven by the predictor with the largest absolute value of $\mu_m$, denoted predictor $j$, with $|\Lambda_j| = 1$ and all other elements of $\Lambda$ equal to zero. In contrast, under an $L_2$ constraint, the optimal $\Lambda$ (given $\Phi$) is proportional to the $M$-vector that collects the $\mu_m$ values. By comparison, Sharpe ratio-maximization effectively harnesses the benefits of diversification across predictors, assigning meaningful weight to multiple signals.
In the context of expected return-maximization, he2024PPMulti extend the principal portfolios framework of kelly2023principal from a single-signal to a multi-signal setting by introducing a three-dimensional prediction tensor. Our study should not be viewed as a multi-predictor extension of principal portfolios. Rather, we propose a framework that differs in both econometric structure and economic objective. From a modeling standpoint, we focus on a two-dimensional matrix $\Pi$, where one dimension captures multiple signals and the other encodes cross-predictive relationships across assets. From an economic perspective, the proposed methodology is explicitly designed to flexibly optimize the Sharpe ratio.
Fifth, the realized return $\pi_s$ of the maximum-Sharpe ratio portfolio is proportional to the stochastic discount factor (SDF), as implied by the fundamental asset pricing identity hansen1991implications, cochrane2009asset, back2017asset:
where $M_s$ denotes the pricing kernel and $\omega$ is the vector of slope coefficients. Identifying the true $\omega$ is challenging in finite samples due to the “limits to learning” highlighted by didisheim2024apt. While the literature has proposed various estimators of the SDF, our approach introduces a novel proxy that explicitly captures cross-asset spillovers, distinguishing it from prior work.
Sixth, kelly2023principal focus on expected return-maximization and propose an alpha-beta decomposition: the antisymmetric and symmetric components of the prediction matrix yield the principal alpha and principal exposure portfolios, respectively. Although our expected return-maximizing strategy can be cast within this framework, our Sharpe ratio-maximizing strategy—by construction—excludes alpha, consistent with the SDF interpretation in Equation (ref).
Empirically, we demonstrate that expected return-maximizing and Sharpe ratio-maximizing strategies—both accounting for cross-asset spillovers—lead to substantially different outcomes. The Sharpe ratio-maximizing strategy consistently delivers significantly higher Sharpe ratios across the full sample, as well as during both expansion and contraction periods.
Next, sorting assets by the estimated weights surfaces the portfolio’s informational backbone: it ranks assets by how much they raise the strategy’s risk-adjusted payoff. High-ranked (large-weight) assets are those that sharpen the payoff of the pricing kernel in three complementary ways: they carry economically meaningful fundamentals; they occupy advantageous positions in the web of cross-asset co-movements that let the portfolio harness spillovers; and they help balance the residual risks created elsewhere in the strategy. An asset can rank highly even if its own return is not strongly predictable—when it acts as a conduit that improves how the portfolio captures cross-asset structure or when it completes the diversification needed to express valuable payoff directions more cleanly. Lower-ranked assets contribute less to efficiency either because their information is largely redundant or because they add volatility without commensurate benefit.
Finally, as shown in Appendix (ref), the connection matrix $\Psi$ closely aligns with the projection of stock returns onto the distinct elements of the signals.
Up to this point, we have only imposed norm constraints on the strategy's positions. However, empirical asset pricing typically requires a trading strategy, factor, or anomaly to take the form of a long-short portfolio—that is, to be zero-cost with total leverage equal to two.
The following proposition imposes this zero-cost constraint on the strategy.
Notice that \( \omega_t' \iota_N = 0 \), where \( \iota_N \) is an \(N\)-vector of ones. Fortunately, all previous derivations remain valid under the zero-cost constraint.
The necessary modifications are as follows. Define \( \Pi_{si} = \Theta (r_{s} S_{it}') \) for each \( i = 1, 2, \ldots, N \), and construct \( \Pi_s \) by vertically stacking \( \Pi_{si} \). All investment metrics in Proposition (ref) can then be re-derived under the zero-cost constraint.
In Appendix (ref), we demonstrate that the zero-cost constraint reduces the expected profitability of the trading strategy. However, this constraint is essential for ensuring comparability across strategies.
In our empirical analyses, we primarily focus on zero-cost strategies, where the long and short positions are of equal magnitude by construction. To further ensure comparability, we rescale these positions so that total portfolio leverage equals two. This adjustment aligns our strategies with standard practice in the literature fama1993common.
We provide methods for estimating the unknown parameters underlying the trading strategy.
The following proposition presents the solution for the strategy that maximizes expected return.
These estimates correspond to the first singular vectors from the matrices \(V\) and \(U\), respectively. This choice ensures that the optimal trading strategy leverages the directions that capture the greatest variance in the prediction matrix \(\Pi\), thereby extracting the most informative signal structure. Importantly, \(\hat{\Lambda}\) and \(\hat{\Phi}\) are obtained from the singular value decomposition of the sample-based matrix \(\Pi\), and should therefore be interpreted as empirical estimators rather than population parameters.
The following propositions formulate the estimates that maximize the squared Sharpe ratio. Appendix (ref) provides the proof and detailed derivations.
In this way, we utilize both alternative expressions for the Sharpe ratio in Proposition (ref) to iteratively estimate the optimal parameters $\Lambda$ and $\Phi$. However, the eigenvalue problems in (ref) and (ref) require computing the inverse of large matrices, which is challenging in high-dimensional settings. To address this, we propose the following proposition to iteratively estimate $\Lambda$ and $\Psi$.
We highlight several key aspects of Sharpe ratio-maximization.
First, the preceding propositions recast the problem as a managed-portfolio optimization, yielding the optimal weights for the tangency portfolio—or equivalently, for the stochastic discount factor (SDF).
Second, we impose a common ridge penalty $\lambda$ when estimating both $\Lambda$ and $\Phi$, thereby enforcing uniform shrinkage across all components. This shared regularization parameter simplifies exposition, enhances replicability, and mitigates overfitting in finite samples. We implement the five-fold cross-validation scheme to select $\lambda$ dynamically.
Third, although the generalized eigenvalue solution provides a population-level characterization of the Sharpe ratio-maximizer, in practice we replace the unknown moment matrices with their sample analogues and apply the same ridge penalty. Rather than solving a generalized eigenvalue problem directly, we cast the estimation as a single ridge-penalized regression. This approach recovers the optimal SDF direction in finite samples, improves numerical stability by shrinking weights on weak or collinear signals, and avoids the computational burden of eigendecomposition. The resulting weight vector exactly coincides with the theoretical maximizer under the ridge-regularized sample formulation.
Fourth, the solution to the Sharpe ratio-maximizing strategy can be interpreted as a regularized linear combination of the principal components (PCs) of the matrix \( \Pi \), with both \( \Lambda \) and \( \Phi \) estimated via ridge regressions on projected versions of \( \Pi \). Unlike expected return-maximizing approaches that primarily load on the leading PCs, this strategy optimizes portfolio weights across the full spectrum of PCs. As a result, it captures predictive signals even in low-variance directions---consistent with the findings of kelly2025universal---and achieves superior risk-adjusted performance.
Fifth, our methodology for estimating the stochastic discount factor offers a distinct contribution to recent advances that emphasize firm characteristics (e.g., kelly2019characteristics, lettau2020factors, chen2024deep, feng2024deep, didisheim2024apt, cong2025growing, liu2025genetic). Unlike these approaches that treat assets independently, we incorporate structured cross-asset dependencies. This not only enhances empirical performance in out-of-sample tests but also yields a more interpretable economic narrative for how information propagates across securities.
Cross-asset dependencies are also central to recent transformer-based approaches in asset pricing, which leverage multi-headed attention mechanisms to extract and aggregate predictive signals across assets. For instance, cong2022alphaportfolio introduce AlphaPortfolio, a deep reinforcement learning framework with cross-asset attention networks (CAAN) that model interdependencies among securities. Similarly, the AIPM framework of Kelly2024aipm embeds transformer architectures within the SDF, showing that nonlinear information sharing across assets can significantly improve empirical performance.
While these transformer models offer substantial modeling flexibility, our framework provides a complementary linear alternative that emphasizes transparency and interpretability. We capture cross-asset spillovers through a connection matrix $\Psi$, where each element $\Psi_{i,j}$ quantifies the predictive influence of asset $i$’s signals on asset $j$’s returns. Although related to the linear surrogate of the transformer, our approach differs in a key respect. The linear transformer models the attention matrix as a function of asset-level signals, requiring estimation of $O(M^3)$ parameters, where $M$ denotes the number of characteristics. In such setups, signal relevance and cross-asset dependencies are entangled within the signal space.
By contrast, we disentangle these two components: signal relevance is captured by a vector $\Lambda$, while cross-asset connections are modeled separately via $\Psi$. This separation reduces parameter complexity to $O(M + N^2)$, promotes computational efficiency, enables straightforward replication, and delivers an economically interpretable decomposition of predictive strength and cross-asset signal spillovers.
Our dataset combines monthly stock returns from the Center for Research in Security Prices (CRSP), accounting variables from Compustat, and analyst coverage and earnings forecasts from the Institutional Brokers' Estimate System (IBES). We assume that quarterly and annual financial statements from Compustat become publicly available four months after the end of the corresponding fiscal quarter. The full sample spans January 1963 to December 2023. Out-of-sample evaluation begins in February 1973, with estimation windows based on rolling samples of the most recent 120 monthly observations.
We employ 138 firm-level signals across 13 characteristic themes: Accruals, Debt Issuance, Investment, Leverage, Low Risk, Momentum, Profit Growth, Profitability, Quality, Seasonality, Size, Short-Term Reversal, and Value. These signals originate from jensen2023there.\footnote{We use the “Global Stock Returns and Characteristics” dataset under “Contributed Data Forms” on WRDS: \url{https://wrds-www.wharton.upenn.edu/pages/get-data/contributed-data-forms/global-factor-data/}. Table IA.II of jensen2023there details the signal definitions and references. Of the original 153 signals, we exclude 15 that begin after 1963 to satisfy the sample-coverage requirements of kelly2023principal. We apply standard filters to retain only observations with: (i) excntry = “USA”, (ii) CRSP shrcd $\in\{10,11\}$, (iii) CRSP exchcd $\in\{1,2,3\}$, and (iv) non-missing monthly excess return (ret_exc) and next-month excess return (ret_exc_lead1m). Each characteristic is standardized to have a mean of zero and a standard deviation of one.}
For each of the 138 signals, we sort stocks into terciles each month and compute high-minus-low factor returns. To form factor-level signals, we aggregate stock-level signals into corresponding factor portfolios. Returns and signals are value-weighted by market equity, with individual market-equity weights winsorized at the 80th percentile of NYSE capitalization, following the data providers' recommendations.
We also construct bivariate sorted portfolios to serve as alternative investment universe. First, stocks are sorted into two size groups (big vs.\ small) based on market equity. Independently, each signal sorts stocks into three groups (high, medium, low). Cross-classifying these sorts produces six portfolios; we retain only the high and low portfolios for each size group, resulting in four portfolios per signal. As with the spread portfolios, returns and signals are capped-value-weighted by winsorized market equity. We omit the bivariate portfolios for the characteristic ami_126d due to missing returns in 2023. Moreover, since size already plays a role in the sorting procedure, we consider a total of $136 \times 4 = 544$ portfolios.
Thus, we consider two investment universes: one constructed from univariate sorts comprising 138 spread portfolios, and the other from bivariate sorts comprising 544 portfolios. Each portfolio is associated with a time series of returns and 138 signal observations.
To build intuition for the proposed framework, we construct a low-dimensional toy dataset comprising five firm characteristics—market equity (ME), book-to-market ratio (BM), operating profits to lagged book equity (OP), asset growth (INV), and 12-month momentum (MOM)—and nine portfolios formed by a $3 \times 3$ sort on ME and BM (ranging from ME1×BM1 to ME3×BM3).
This simplified setup allows us to explicitly report the estimated low-dimensional parameters $\Lambda$ and $\Psi$, as well as the weight vector $\omega$. It also enables a comparison of key performance metrics for: (i) strategies subject to unit-norm constraints without an explicit zero-cost requirement; and (ii) zero-cost strategies with total leverage constrained to two.
We implement expected return and Sharpe ratio-maximizing strategies, as formulated in sections (ref) and (ref). These strategies, which target different objectives, yield notable differences in parameter estimates and performance outcomes. Table (ref) summarize the monthly average returns, monthly standard deviations, and annualized Sharpe ratios for each strategy over the out‐of‐sample period from February 1973 to December 2023.
The first two rows of the table consider the case in which the zero-cost assumption is not imposed. The results show that the strategy maximizing expected return (MR\,Cross) delivers a high average monthly return of 5.56%, but with substantial volatility (standard deviation of 61.9%), yielding a Sharpe ratio of just 0.31.
The Sharpe ratio-maximizing strategy (MS\,Cross) attains a mean return of 2.36% and a much lower volatility (9.78 %), yielding a Sharpe ratio of 0.84. Consequently, a mean--variance investor would find the Sharpe ratio‐maximizing strategy considerably more attractive, whereas an investor solely targeting expected returns would prefer the expected return‐maximizing strategy. Thus far, the out‐of‐sample performance aligns closely with the ex ante investment objectives.
Next, we consider a strategy that maximizes the Sharpe ratio using self‐prediction to isolate the incremental contribution of cross‐predictive relations relative to self‐predictive relations. The key distinction between these two strategies lies in the structure of the connection matrix $\Psi$. Under cross‐prediction, $\Psi$ is a full $9\times9$ matrix, capturing all pairwise interactions among the characteristics and returns of assets. In contrast, under self‐prediction, $\Psi$ is restricted to its diagonal terms.
The second and third rows of Table (ref) report the performance of the Sharpe ratio–maximizing strategies under cross‐prediction and self‐prediction, respectively. The cross‐prediction strategy (MS\,Cross) delivers a Sharpe ratio of 0.84 with a mean return of 2.36%, whereas the self‐prediction variant (MS\,Self) achieves a lower Sharpe ratio of 0.60 and the lowest mean return of 1.31% . This gap in both risk‐adjusted and absolute returns illustrates the incremental benefit of incorporating cross‐predictive relationships beyond self‐prediction alone, underscoring the pivotal role of cross‐predictive dynamics in enhancing portfolio performance.
To provide further economic perspective on the value of accounting for cross-stock predictability, we compute the certainty equivalent return of the investment strategies. The certainty equivalent is defined as $CE = \mu - \frac{\gamma}{2} \sigma^2$, where $\mu$ and $\sigma$ are the expected return and volatility of the strategy, respectively, and the risk aversion parameter $\gamma$ is set to 2. Accounting for cross-predictability, the certainty equivalent rate of return is approximately 16.84% per year—8.00% higher than that of self-predictability—indicating economically significant gains.
We next maximize expected return and Sharpe ratio under the zero‐cost and leverage‐two constraints. The fourth and fifth rows of Table (ref) report these constrained strategies, confirming that imposing the zero‐cost restriction reduces expected returns for both objectives. Nevertheless, even with zero cost and fixed leverage, the Sharpe ratio–maximizing strategy outperforms the expected‐return–maximizing strategy, delivering a higher mean return (0.50% vs.\ 0.49%) and a substantially higher Sharpe ratio (1.22 vs.\ 0.53).
To provide additional insight into cross-prediction and self-prediction strategies, Table (ref) reports the estimated values of $\Lambda$, $\Psi$, and $\omega$ for each approach without imposing the zero-cost constraint. The estimation window spans 120 months, from December 2003 to November 2023, covering our last out-of-sample period. Panel A presents the Sharpe ratio-maximizing cross-prediction strategy; Panel B presents the Sharpe ratio-maximizing self-prediction strategy; and Panel C reports the differences in the portfolio weights $\omega$ between the two.
In Table (ref), Panel A shows that the estimated $\Lambda$ coefficient for book‐to‐market equity (BM) is $-0.34$, whereas the coefficients for the other four characteristics are all positive, with the smallest value at $0.21$. This suggests that the Sharpe ratio–maximizing strategy with cross‑prediction is well balanced across the five characteristics. The full $9\times9$ matrix $\Psi$ exhibits substantial values both on and off the diagonal: the average absolute value of its diagonal entries is $0.0068$, compared to an average absolute off‑diagonal entry of $0.0805$, indicating that cross‑predictive relationships play an even more substantial role in defining the trading strategy.
Panel B of Table (ref) shows that under self‐prediction the estimated $\Lambda$ coefficients exhibit greater dispersion---asset growth (INV) even turns negative---while $\Psi$ is constrained to its diagonal (average absolute value of $0.0288$, all off‐diagonals zero). This contrast highlights the structural effect of cross-predictability: including cross-predictive terms not only yields nonzero off-diagonal elements of $\Psi$ but also shifts the estimated $\Lambda$ coefficients, altering the relative importance of characteristics.
Panel C reports how the optimal weights $\omega$ shift between cross‐ and self‐prediction: under cross‐prediction, long exposures to ME3\,BM1 decrease, and shorts in ME1\,BM2 deepen. For example, the ME1\,BM2 position is $-0.27$ under cross‐prediction---driven by off‐diagonal $\Psi$ entries of $0.24$ and $0.25$---whereas it is substantially smaller under self‐prediction.
As noted earlier, the optimal trading strategy that accounts for cross-predictability delivers a 8.00% higher certainty equivalent return, suggesting that the estimated $\Lambda$ and $\Psi$, which determine the portfolio weights $\omega$, differ to an economically significant degree when cross-predictability is incorporated, relative to the benchmark case of self-predictability.
In summary, the results in Tables (ref) and (ref) confirm that incorporating cross‑predictive relationships is valuable for constructing robust investment strategies, even in a low‑dimensional illustrative setting.
Table (ref) reports the performance of linear cross‑predictive strategies implemented as zero‑cost, leverage‑two portfolios, comparable to common factor and anomaly implementations. MR and MS denote the strategies that maximize expected return and the Sharpe ratio, respectively. In Panel A, we consider an investment universe with 138 spread portfolios detailed in the data section. Over the full sample period, MR achieves a monthly average return of 0.42% with an annualized Sharpe ratio of 0.45, whereas MS records a lower monthly average return of 0.29% but a substantially higher annualized Sharpe ratio of 2.21.
We further analyze performance during evolving market states by splitting the out‑of‑sample period at the median of the investor sentiment index baker2006investor.\footnote{The sentiment data spans July 1965 to December 2023 and is obtained from the variable `SENT” on Jeffrey Wurgler's website: \url{https://pages.stern.nyu.edu/ jwurgler/data/SENTIMENT.xlsx}.} During high‑sentiment regimes, MR delivers an average monthly return of 0.73%, while in low‑sentiment regimes its return falls to 0.11%. The MS strategy exhibits robust Sharpe ratios across both regimes: 2.19 in high‑sentiment periods and 2.22 in low‑sentiment periods.
In Panel B, we evaluate investments in 544 bi‑variate sorted portfolios as detailed in the data section. Over the full out‑of‑sample period (January 1973–December 2023), MR delivers a monthly average return of 0.45% and an annualized Sharpe ratio of 0.52, whereas MS achieves an exceptionally high annualized Sharpe ratio of 3.32. In sub‑period analyses, MR’s average return increases during high‑sentiment regimes, while MS maintains Sharpe ratios above 3 in both high‑ and low‑sentiment periods.
In Panels C and D, we split the period January 1990–December 2023 at the median of the VIX index.\footnote{The VIX data spans 1990 to 2023 and is obtained from the CBOE: \url{http://www.cboe.com/products/vix-index-volatility/vix-options-and-futures/vix-index/vix-historical-data/}.} In Panel C (spread portfolios), MR's average return is 0.59% during high‑VIX regimes and 0.07% during low‑VIX regimes (0.33% full sample), while MS records Sharpe ratios of 2.02 and 1.98 in high‑ and low‑VIX regimes (1.92 full sample).
In Panel D (bi‑variate sorted portfolios), MR attains a monthly average return of 0.39% and an annualized Sharpe ratio of 0.42, while MS achieves a Sharpe ratio of 2.90. MR’s return remains higher in high‑VIX regimes, and MS sustains Sharpe ratios around 3 in both high‑ and low‑VIX regimes.
In summary, MR strategies deliver high expected returns during high‑sentiment and high‑VIX regimes, but considerably lower expected returns otherwise. In contrast, MS strategies consistently achieve superior Sharpe ratios across all market states.
We compare the principal portfolio-based trading strategies of kelly2023principal with our own over the out-of-sample period from 1973 to 2019, as in their study. The results are reported in Table (ref).
Panel A, row 1 (PP--ME), reports the performance of the first principal portfolio on the market-equity signal: a 3.27% monthly expected return, a 0.51 annualized Sharpe ratio, and a sum of absolute equity positions equal to 23.22. Rows 2 and 3 present the first principal portfolios for the book-to-market and momentum signals, which achieve Sharpe ratios of 0.60 and 0.48, respectively, with similarly high leverage. The principal portfolio can be applied to only one signal at a time. We also consider to take the 1/N equal-weighted strategy of the first principal portfolios across all 138 signals, namely the PPEW strategy, which delivers a 2.83% monthly expected return and a 0.56 annualized Sharpe ratio. Notably, the leverage of PPEW is only 1.35, reflecting substantial diversification benefits by equal weighted average across predictors.
Our maximum-expected return strategy achieves an 135.14% monthly expected return and a 0.52 annualized Sharpe ratio, with leverage of 537.70. Overall, the maximum-expected return strategy slightly underperforms the principal portfolios in Sharpe ratio, albeit remains reasonably close to them.
By contrast, the MS strategy harnesses multiple predictors to diversify exposures and optimize risk-adjusted returns, achieving an annualized Sharpe ratio of 2.22 with a leverage factor of 438.01. While the principal portfolio approach targets expected return subject to a volatility constraint, our strategy is derived directly from Sharpe ratio-maximization. As a result, it places greater emphasis on balancing return and risk, leading to improved performance on risk-adjusted metrics in our empirical setting.
To enhance implementability, we impose zero‐cost and leverage‐two constraints on both strategies. Panel B of Table (ref) reports the resulting performance. Under these constraints, the maximum‐expected‐return strategy (Row 1) achieves a 0.46% monthly expected return and an annualized Sharpe ratio of 0.51, while the maximum‐Sharpe ratio strategy (Row 2) attains a 0.30% monthly expected return and an annualized Sharpe ratio of 2.33. In both cases, the portfolios maintain zero net cost and a constant leverage of two in every period.
Overall, the maximum‑Sharpe ratio strategy remains highly competitive—delivering superior risk‑adjusted performance relative to a range of recent approaches, including principal portfolios. Accordingly, we focus our subsequent analyses to the constrained max‑SR strategy.
The existing literature on SDF estimation has predominantly focused on self-predictive frameworks, where each asset’s signals are used solely to forecast its own returns. kelly2019characteristics propose Instrumented PCA with the belief that the factor loadings on SDF are determined by assets characteristics, overcoming the limitations of static loading in PCA. lettau2020factors find that the SDF estimated on Risk-Premium PCA is more highly correlated with the true SDF than those estimated on PCA. luo2025sdf estimate the SDF with observable characteristics-based factors with L1-penalized SDF regression; whereas, didisheim2024apt apply the L2-penalized SDF regression on observable and Random-Fourier-Feature generated factors. All of these papers have been working on high-dimensional characteristics-based portfolios to estimate the SDF, where the belief of self-prediction are embedded the portfolios.
By contrast, our framework utilizes managed-portfolios inherently reflecting the belief of cross-prediction: $\pi_s$ (ref), $\chi_\Phi$ (ref), and $\chi_\Lambda$ (ref). Whether cross-predictive strategies—where signals from one asset help predict the returns of others—can systematically outperform self-predictive ones in high-dimensional settings remains an open question. To investigate this, we construct the self-predictive strategy by restricting the matrix $\Psi$ to its diagonal, thereby eliminating all cross-asset interactions.
Panel A of Table (ref) reports the empirical performance of the Sharpe ratio–maximizing strategies on the 138 spread portfolios. The self-predictive strategy achieves a Sharpe ratio of 1.42, while the cross-predictive counterparts attain Sharpe ratios of 2.08 without zero-cost requirement and 2.21 with zero-cost and leverage-two constraints. This more than 0.60 difference in Sharpe ratio underscores the incremental value of incorporating cross-asset predictive signals.
Panel B reports results for the 544 bivariate sorted portfolios. The self-predictive strategy achieves a Sharpe ratio of 2.06, while the cross-predictive variants reach 3.32 and 3.00 under constrained and unconstrained implementations, respectively. This gap of more than 1.00 in Sharpe ratio highlights the significant contribution of off-diagonal elements in $\Psi$ to improved portfolio performance.
Overall, the evidence confirms that cross-predictive strategies materially enhance the estimation and performance of stochastic discount factors—particularly in richer portfolio universes and longer samples.
We conduct a series of factor-spanning tests to assess whether the established asset pricing factors fully explain the expected returns of the Sharpe ratio-maximizing strategies. Table (ref) reports monthly alphas (%), factor loadings, and associated $t$-statistics. Panel A presents the Sharpe ratio-maximizing strategy on the spread portfolios, while Panel B reports for the bivariate sorted portfolios.
We first evaluate the fama2015five five-factor model (FF5). The strategy on spread portfolio exhibits modest loadings on Market ($\beta = -0.01$, $t = -2.56$) and SMB ($\beta = 0.01$, $t = 1.78$) but insignificant exposures to the other four factors, while delivering a highly significant monthly alpha of 0.29% ($t = 13.29$). This suggests that the strategy's returns are largely orthogonal to the FF5 factors. Also, we augment the FF5 model with momentum (UMD), short-term reversal (REV), and liquidity (LIQ) factors pastor2003. In this expanded specification, the strategy shows significant loadings on UMD and REV but not on LIQ, while its alpha remains economically and statistically significant at 0.26% ($t = 11.54$). These findings indicate that momentum and reversal effects partially explain the strategy's performance, with little role for liquidity risk.
Next, we then examine the hou2015 four-factor model, which incorporates investment (R_IA) and profitability (R_ROE) factors alongside market and size factors. The strategy displays negligible loadings on R_IA and R_ROE, while maintaining a highly significant alpha. The stambaugh2017 mispricing factors---MGMT and PERF---also fail to subsume the strategy's returns: The strategy shows minimal exposures to both factors, with an alpha of 0.28% ($t = 10.20$). Then, we assess the daniel2020short model, which includes the market factor and two behavioral factors, PEAD and FIN. While the strategy loads significantly on PEAD, its alpha remains robust at 0.29% ($t = 11.09$), and it shows no meaningful exposure to FIN. Finally, in a comprehensive regression incorporating all fourteen factors, The strategy maintains an alpha of 0.26% ($t = 8.04$), with statistically significant but economically small loadings on SMB, UMD, REV, LIQ, FIN, and R_IA. These results collectively demonstrate that the strategy's expected returns cannot be fully explained by existing factor models.
Panel B corroborates these findings. The strategy on bivariate sorted portfolio displays statistically significant but economically modest loadings on RMW, CMA, REV, PERF, R_IA, MGMT, and PERF. Notably, it maintains a monthly alpha of 0.25% ($t = 11.36$) even after controlling for all fourteen factors, further supporting the strategy's robustness to established factor models.
Across all specifications—including the Fama–French five-factor model with UMD, REV, and LIQ augmentations, Hou–Xue–Zhang, Stambaugh–Yuan, and Daniel–Hirshleifer–Sun frameworks, and even the comprehensive fourteen-factor regression—the Sharpe ratio-maximizing strategies on the spread and bivariate sorted portfolios exhibit persistently large and highly significant alphas with only moderate loadings on existing factors. This suggests that conventional models may miss the cross-asset return predictability captured by our strategy. Below, we further analyze the pricing content of the Sharpe ratio-maximizing strategies.
To assess the persistence and evolution of risk‐adjusted returns over time, Figure (ref) Panel A shows the ten‐year trailing Sharpe ratios of our three maximum-Sharpe ratio strategies—MS Spread, MS BiSort and MS BiSort fixed—alongside those of the market and momentum factors for comparison. \footnote{The shrinkage parameter $\lambda$ for MS Spread and MS BiSort are selected via cross-validation. Appendix (ref) reports the selected parameter values of time. MS BiSort fixed uses a fixed $\lambda=1$, which is the most frequently selected value. } By smoothing over a decade window, we can observe how the trading strategies respond to changing market conditions.
These strategies deliver eye‐catching Sharpe ratios in the 1990s—MS BiSort climbs as high as 4–7 before 2000, and MS Spread approaches 4—reflecting their ability to capture persistent value-enhancing opportunities. After 2000, however, it is natural to see some attenuation: wider adoption of anomaly tradings, increased market liquidity, and a lower‐volatility regime tend to compress excess returns over time. Accordingly, by the end of 2023, the trailing Sharpe of MS Spread and MS BiSort has moderated to about 1.2. By contrast, the MS BiSort fixed seems to deliver even higher Sharpe ratios than MS BiSort, suggesting that our cross-validation scheme is conservative and provides a low bound for the out-of-sample performance.
To make more clear comparison, Figure (ref) Panel B shows the Sharpe ratio of each strategy relative to that of MS BiSort fixed. In early sample before 2000, the Sharpe ratio of MS Spread (MS BiSort) is approximately 60% (90%) that of MS BiSort fixed, and market and momentum factors have below 20% Shape ratio relative to MS BiSort. In the most recent sample, the Sharpe ratios of MS Spread, MS BiSort, MKT-RF, and UMD are 70%, 75%, 43%, and 4% that of MS BiSort fixed.
For context, both the market factor’s rolling Sharpe ratio and that of the momentum factor remain well below our strategies over the entire forty‐year span. Although the performance gap narrows in the post‐2000 era, both maximum-Sharpe ratio strategies continue to deliver robust risk‐adjusted returns relative to these benchmarks.
Table (ref) reports the (annualized) Sharpe ratios of the cross-predictive maximum-Sharpe ratio strategies, Fama-French five factors, and momentum factor for three sample periods: the whole OOS period from 1973:02 to 2023:12, before 2000:01, and after 2000:01. Although, the Sharpe ratios of our strategies attenuate after 2000, they remain competitive compared to the benchmark factors in three sample periods.
To understand the economic underpinnings of our Sharpe ratio-maximizing strategies or SDF, we examine the estimated values of $\Lambda$, which assign weights to firm-level predictive signals. These weights reflect the relative contribution of each signal to the SDF. We focus on the absolute value of these weights averaged over time to assess long-term signal importance. Table (ref) presents the ten most influential signals, ranked by their time-series average of absolute $|\Lambda|$ values, where Panel A is for spread portfolios and Panel B is for bivariate sorted portfolios.
Panel A, investing in spread portfolios, indicates that the most important signals are concentrated in the investment and value categories. For instance, the top signal---liquidity of book assets ortiz2014real---receives an average importance of 0.139, while dividend yield litzenberger1979effect, the leading signal in the value theme, ranks seventh overall with an importance of 0.126. These findings suggest that the strategy places greater emphasis on firm fundamentals linked to capital structure, financing constraints, and valuation, rather than technical or return-based indicators.
As for Panel B, profitability dominates the top ten signals, followed by size, low leverage, and low risk themes. For instance, return on equity haugen1996commonality and operating profitability-to-lagged book equity fama2015five are top signals, all belonging to \emph{profitability}. Besides, \emph{price per share} miller1982dividends emerges from the \emph{size} theme, recalling stronger size effects in the test assets sorted on size and other signals.
Figure (ref) presents the importance measures for all 138 signals, organized into 13 thematic categories (as defined in the Data section). Sub-figures (a) and (b) display theme-level importance for spread portfolios and bivariate-sorted portfolios, respectively.\footnote{We provide the time-varying signal-level importance measures in Figure (ref) of the Appendix (ref).} The heatmap visualization employs color intensity to indicate importance levels---with red (blue) representing high (low) importance--allowing clear identification of which signals consistently influence portfolio construction.
In sub-figure (a) for spread portfolios, investment- and value-related signals dominate the red spectrum, reinforcing the role of tangible firm fundamentals. In contrast, momentum, profit growth, debt issuance, seasonality, and \emph{short-term reversal} appear consistently in the blue range, indicating minimal weight in the Sharpe ratio-maximizing SDF.
Turning to sub-figure (b) for bivariate-sorted portfolios, the profitability theme dominates the heatmap, particularly following a pronounced regime shift in the late 1980s. The size theme exhibits persistent importance throughout the sample period, reflecting the strong cross-sectional dispersion in firm size within our test assets. In contrast, accruals, profit growth, seasonality, and short-term reversal show consistently low importance over the entire sample.
Our analysis reveals that while the dominant predictive role of investment, value, profitability, and size themes remains stable over time, certain signals---particularly accruals and quality---exhibit heightened importance during high-volatility or low-sentiment periods. This time variation suggests dynamic shifts in return predictability patterns, which our framework successfully captures through its adaptive structure.
In summary, our signal importance analysis demonstrates that the cross-predictive SDF is primarily driven by stable, economically grounded predictors, with negligible dependence on transient or noisy effects. These findings not only underscore the robustness and economic interpretability of our framework but also open new avenues for investigating the fundamental drivers of cross-sectional return predictability.
To uncover the economic structure embedded in the cross-predictive matrix $\Psi$, we interpret $\Psi$ as the adjacency matrix of a directed network across $N$ assets. This representation enables us to move beyond portfolio-level effects and examine how predictive information flows through the cross-section. That is we identify assets that function as net transmitters or receivers of signals and assessing the alignment of these linkages with economic groupings such as firm size.
Following the connectedness methodology of diebold2014network, we compute three metrics for each asset $i$---outgoing connectedness ($\mathrm{FROM}$), incoming connectedness ($\mathrm{TO}$) and net connectedness ($\mathrm{NET}$)---along with a market-level overall network intensity ($\mathrm{TOTAL}$). Let $\Psi_{i,j}$ denote the predictive influence of asset $i$ on asset $j$. We define the network metrics as follows:
Here, $\mathrm{FROM}_i$ measures the total strength of predictive signals sent from asset $i$ to others, capturing how much $i$ contributes to forecasting the returns of other assets. $\mathrm{TO}_j$ measures the total strength of predictive signals received by asset $j$ from all other assets, reflecting how much $j$ is influenced by the rest of the network. $\mathrm{NET}_k$ is the difference between incoming and outgoing connectedness, indicating whether asset $k$ is a net transmitter ($< 0$) or net receiver ($> 0$) of predictive information. $\mathrm{TOTAL}$ aggregates the overall off-diagonal magnitude of $\Psi$ across all asset pairs, summarizing the average intensity of cross-asset predictive linkages in the network. The use of absolute values follows diebold2014network and ensures all measures are non-negative, thereby capturing signal strength regardless of sign.
We compute these metrics monthly for two asset universes—138 spread portfolios and 544 bivariate sorted portfolios—over $T = 611$ months. To investigate the firm-level characteristics driving variation in connectedness, we estimate monthly cross-sectional regressions:
where $\mathrm{Connectedness}_{i,t}$ is one of $\mathrm{FROM}_i$, $\mathrm{TO}_i$, or $\mathrm{NET}_i$, and $\mathrm{Char}_{i,t}$ is a vector of observable characteristics. We report time-series averages of the estimated coefficients along with Newey--West neweywest1987 $t$-statistics using a Bartlett kernel and lag length $L = 4(T/100)^{2/9} \approx 5$.
Table (ref) reports the results of monthly cross-sectional regressions of three network connectedness measures—$\mathrm{FROM}$, $\mathrm{TO}$, and $\mathrm{NET}$—on firm characteristics for two groups of test assets: spread portfolios (Panel A) and bivariate sorted portfolios (Panel B). The results reveal economically intuitive patterns linking a stock’s network role to size, valuation, profitability, investment, momentum, and several trading frictions.
In Panel A for spread portfolios, the $\mathrm{FROM}$ regressions, measuring how much a stock helps predict others, we observe that smaller stocks (low ME), high book-to-market (BM), high profitability (OP), and high momentum (MOM) stocks tend to transmit stronger signals to others. These firms—small, value, profitable, and past winners—have greater forecasting influence, possibly because they aggregate market-wide information or drive co-movements. Additionally, stocks with low illiquidity (ILL) and low turnover (TRN) exhibit higher FROM, suggesting that liquidity increase a stock’s impact to the network. Volatility (VLT), by contrast, enters positively, implying that more volatile stocks spill predictive attention. Notably, the coefficient on size (ME) becomes insignificant, once controlling five trading frictions, which means that the size effect on $\mathrm{FROM}$ is a manifestation of trading frictions but not size itself.
The $\mathrm{TO}$ regressions, which capture how strongly a stock is predicted by others, show the opposite patterns on many characteristics. Stocks with high ME, high BM, low OP, low INV, high MOM, high VLT, and high BETA receive more predictive inputs from others. This suggests that firms that are large, volatile, illiquid, and priced as value stocks appear more susceptible to being forecasted using cross-asset information. Interestingly, high-MOM stocks both receive and transmit signals, indicating they may act as informational amplifiers within the network.
The $\mathrm{NET}$ regressions, defined as $\mathrm{TO} - \mathrm{FROM}$, consolidate these effects to identify whether a stock is a net receiver or transmitter of predictive information. Stocks that are large (high-ME), low-BM, low-OP, low-INV, and low-MOM tend to be net receivers, while small, value, profitable, non-investing, and momentum-driven stocks are net transmitters. These directional patterns highlight a persistent asymmetry: small, value, strong profitability, and conservative investment firms disseminate predictive signals, while larger and illiquid firms absorb them.
In Panel B for bivariate sorted portfolio, these patterns still exist. For ease of interpretation, we focus on the $\mathrm{NET}$ regressions. We find that small, low-BM, high-OP, low-INV, and low-MOM firms are net receivers in the network, while big, value, weak-profitable, conservative-investing, and high-momentum stocks are net transmitters. After controlling five trading frictions in the regressions, the coefficient on size become significantly positive, while other four coefficients are unchanged. As for trading frictions, stocks with low volume, low volatility, high turnover, and low market-beta tend to receive spillovers from others than transmitting signals to others.
Together, two sets of test assets demonstrate significant correlations between network connectedness and asset characteristics, shedding light on that the determinants of cross-asset spillover effects. The estimated $\Psi$ matrix embeds an economically interpretable hierarchy of signal flows, shaped by firm fundamentals and market frictions. This structure supports imposing sparsity or blockwise restrictions to enhance interpretability and control overfitting—especially by limiting signal flows that contradict observed economic asymmetries. Nevertheless, the correlation between connectedness and asset characteristics depends on the choice of test assets. That is, different test assets reflect different patterns in asset pricing, see feng2020taming, avramov2025sparse. In this specific exercise, we confirm the prominent status of size as an asset characteristic in building sorted portfolios as test assets fama1993common.
Table (ref) connects to several literature. For bivariate portfolios (Panel B), we initially corroborate lo1990contrarian, finding big stocks lead small stocks (coefficient -0.15, row 1 on $\mathrm{NET}$)—a result robust to controlling for BM, OP, INV, and MOM (coefficient -0.19, row 2). However, controlling for trading frictions reverses the size coefficient, suggesting big stocks become net receivers, warranting further investigation of size's role in lead-lag effects.\footnote{For comparability, we replicate results for 1973-1987 (matching lo1990contrarian's sample end) and find consistent size coefficient signs.} Contrary to chordia2000trading, we find low-turnover stocks transmit signal to high-turnover stocks after controlling for size.\footnote{While chordia2000trading uses "Trading Volume" in their title, they actually employ daily turnover as their volume proxy.} It holds for both spread and bivariate sorts. The divergence from prior papers reflects discretion in test assets and sample periods. Moreover, the two papers focus exclusively on weekly return spillovers, whereas we incorporate multiple firm-level monthly signals, including past returns. Collectively, we demonstrate that cross-asset spillovers are fundamentally linked to asset characteristics.
Figure (ref) depicts the $\mathrm{TOTAL}$ connectedness index---the average intensity of the off-diagonal elements in $\Psi$---for both the 138 spread portfolios (dashed line) and the 544 bi-sort portfolios (solid line) over the 1973--2023 period. Four key findings emerge. First, the time-series of $\mathrm{TOTAL}$ connectedness on the spread portfolios varies markedly through time: it troughs in the mid-1980s and again after 2020, but peaks around the early 1990s and during the post financial crises, 2010s. Second, the indices for bivariate sorted portfolios share the trough in mid-1980s and peak in early 1990s, however, slight fluctuations after 2000. Overall, the average level of $\mathrm{TOTAL}$ of spread portfolios is almost equal to that of bivariate sorted portfolios before 2000, but become higher after 2000. Third, despite these episodic surges, the time series reverts to a long-run mean near 0.72, suggesting a stable baseline level of cross-asset information transmission.
Taken together, Figure (ref) demonstrates that cross-asset spillover effects intensify during turbulent periods but persist as a pervasive market feature. These findings underscore the importance of modeling the full $\Psi$ matrix---rather than restricting attention to its diagonal elements---for constructing Sharpe ratio-maximizing portfolios.
To analyze directional spillover effects in the bivariate-sorted portfolios more precisely, we decompose the $\Psi$ matrix into four blocks (A, B, C, and D) according to firm size. Figure (ref) presents the resulting predictive information flows across these partitions.
Figure (ref) presents the time series of absolute average values for each of the four blocks in $\Psi$. The results reveal consistently stronger predictive relations in Block A (Small $\rightarrow$ Small) and Block C (Big $\rightarrow$ Small) compared to Block B (Small $\rightarrow$ Big) and Block D (Big $\rightarrow$ Big), particularly during the last two decades. The time-series averages are 1.50 and 1.53 for Blocks A and C, respectively, versus 1.09 and 1.11 for Blocks B and D. Notably, the divergence between the A/C and B/D blocks has increased substantially in recent years.
These findings confirm an asymmetric predictive structure, which aligns with the $\mathrm{NET}$ regression coefficient of $-0.15$ reported in Panel B of Table (ref). This result is consistent with the evidence in lo1990contrarian, showing that large stocks tend to lead small stocks, but not vice versa. The persistent and stable nature of these patterns over time supports the economic rationale for imposing restrictions on $\Psi$, particularly by excluding small-to-large predictive links. Furthermore, the long-run regularity of these asymmetries suggests that dynamic sparsity structures---which adapt to time-varying network block strengths while maintaining economically motivated constraints---could offer significant modeling value.
In summary, the connectedness analysis reveals that the connection matrix $\Psi$ encodes economically meaningful structure. For bivariate sorted portfolios on size and other signals, big stocks act as net transmitters of predictive signals; controlling more signals, we find that low trading volume, high turnover ratio, and low-beta stocks are net transmitters. Meanwhile, value, profitable, non-investing, and high-momentum assets are more likely to be net receivers. The strength of cross-predictive relations is comparable to that of self-predictive effects. The overall network intensity fluctuates over time, but remains around a stable level. Decomposing $\Psi$ by firm size shows that predictive flows from large to small firms dominate those in the reverse direction.
This paper develops a structured framework for constructing Sharpe ratio–maximizing investment strategies using multiple firm-level signals and accounting for informational linkages across assets. By jointly estimating signal relevance and a matrix capturing cross-asset predictive relationships, our approach yields closed-form portfolio weights derived from a generalized eigenvalue decomposition. In high-dimensional settings, estimation is implemented through Ridge-SDF regressions, which offer a stable and interpretable managed-portfolio representation of the decision variables. The resulting stochastic discount factor consistently delivers high out-of-sample Sharpe ratios across a range of asset universes and market conditions, outperforming both self-predictive models and expected return-maximization. Economically, the strategy is primarily driven by fundamental characteristics related to investment, valuation, and profitability. In addition, the estimated connection matrix reveals that large and low-turnover stocks tend to act as net transmitters of predictive signals, while the overall strength of cross-asset linkages remains persistently high over time.
The paper opens several promising avenues for future research. First, the framework could be extended to other asset classes where cross-asset interdependencies are economically meaningful, such as corporate bonds, currencies, sovereign credit, or international equities. For instance, in the corporate bond market, issuer fundamentals or equity-side information may predict bond returns through industry linkages, shared ownership networks, or common analyst coverage. Similarly, in currency markets, major reserve currencies may act as informational hubs whose movements help forecast subsequent shifts in peripheral currencies. Second, incorporating economic structure into the modeling of cross-asset relationships could enhance both interpretability and predictive performance. As the number of assets and signals expands, estimating all possible interactions becomes increasingly challenging. Imposing economically motivated constraints—such as directional spillovers based on firm size or sectoral hierarchies—could provide a more structured and scalable approach.
\setstretch{1.6}