EconBase
← Back to paper

Inferential Theory for Pricing Errors with Latent Factors and Firm Characteristics

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

77,607 characters · 27 sections · 43 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inferential Theory for Pricing Errors with Latent Factors and Firm Characteristics

\def\spacingset#1{ {#1}} \spacingset{1}

\if11 \fi

\if01 \fi

abstractWe study factor models that combine latent factors with firm characteristics and propose a new framework for modeling, estimating, and inferring pricing errors. Following zhang2024testing, our approach decomposes mispricing into two distinct components: inside alpha, explained by firm characteristics but orthogonal to factor exposures, and outside alpha, orthogonal to both factors and characteristics. Our model generalizes those developed recently such as kelly2019characteristics and zhang2024testing, resolving issues of orthogonality, basis dependence, and unit sensitivity. Methodologically, we develop estimators grounded in low-rank methods with explicit debiasing, providing closed-form solutions and a rigorous inferential theory that accommodates a growing number of characteristics and relaxes standard assumptions on sample dimensions. Empirically, using U.S. stock returns from 2000–2019, we document strong evidence of both inside and outside alphas, with the former showing industry-level co-movements and the latter reflecting idiosyncratic shocks beyond firm fundamentals. Our framework thus unifies statistical and characteristic-based approaches to factor modeling, offering both theoretical advances and new insights into the structure of pricing errors.

\spacingset{1.4}

{8pt} {8pt} \setlength\intextsep{8pt} {4pt}

Introduction

The search for a parsimonious yet interpretable representation of asset returns lies at the heart of modern asset pricing. Since the seminal works of sharpe1964capital,ross1976arbitrage,fama1973risk, researchers have studied linear factor models where excess returns are driven by a small number of systematic risk factors. A dominant empirical approach to uncover these factors has been statistical, relying on principal component analysis (PCA) to extract latent sources of common variation chamberlain1982arbitrage,connor1986performance, connor1988risk. While such latent-factor models effectively capture the covariance structure of returns, they often lack clear economic interpretation and are static in nature, making them ill-suited for conditional or time-varying risk exposures.

In parallel, a large literature in empirical finance has emphasized firm characteristics as the basis for factor construction, most prominently through the portfolio-sorting tradition that culminated in the Fama--French family of factor models fama1993common. By anchoring factors in observable firm fundamentals, these models yield interpretable risk premia and direct economic meaning. However, ad hoc portfolio sorts can sacrifice statistical efficiency, discarding variation that is captured by latent statistical factors. Consequently, two lines of research, statistical factor extraction and characteristic-based portfolio construction, have developed largely in parallel, each offering distinct advantages but limited integration.

Recent advances in conditional and high-dimensional asset pricing have sought to bridge these approaches by allowing latent factor structures to depend explicitly on firm characteristics. fan2016projected introduced projected PCA; kelly2019characteristics proposed Instrumented PCA (IPCA), in which factor loadings and pricing errors are modeled as functions of firm characteristics; and kim2021arbitrage and zhang2024testing further refined this framework by relaxing identification restrictions and improving estimation. A complementary literature has incorporated nonlinear and machine-learning-based representations of characteristics, including deep factor and autoencoder models bryzgalova2019forest, gu2021autoencoder, feng2024deep, which demonstrate that firm fundamentals can efficiently span the space of risk exposures. At the same time, econometric work on high-dimensional factor models has developed a rigorous asymptotic theory for latent-factor estimation and inference bai2003inferential, fan2016projected, chernozhukov2023inference, chen2023semiparametric. Yet despite this progress, a unified framework that combines the interpretability of characteristic-based models with the inferential rigor of modern econometrics remains elusive.

Two methodological gaps are particularly salient. First, the IPCA model of kelly2019characteristics assumes that pricing errors (alphas) are fully explained by characteristics, violating the orthogonality condition between alphas and factor loadings required by the Arbitrage Pricing Theory (APT). This undermines the economic interpretation of estimated “pricing errors”, as they may inadvertently load on systematic factors. zhang2024testing highlighted this issue and proposed a decomposition of alphas into components inside and outside the span of characteristics. However, Zhang’s formulation depends on arbitrary choices of orthonormal bases and is not invariant to the rescaling of characteristics, raising concerns about robustness and interpretability. Moreover, the approach remains algorithmic: estimation relies on iterative numerical procedures with bootstrap-based inference but without accompanying asymptotic theory, leaving the econometric underpinnings incomplete.

This paper develops a general econometric framework that addresses these limitations and formally unifies latent-factor and characteristic-based approaches. Building on advances in low-rank and debiased estimation, we propose a model that decomposes pricing errors into two orthogonal components: inside alpha, the portion of mispricing attributable to firm characteristics but orthogonal to factor exposures; and outside alpha, the residual component orthogonal to both factors and characteristics. This decomposition restores theoretical consistency with APT while allowing a richer economic interpretation of both components. By deriving closed-form estimators and explicit bias corrections, we obtain tractable estimators that admit Gaussian inference even as the number of characteristics grows with the sample size. Specifically, our contributions are fourfold:

\paragraph{Modeling.} We provide a new decomposition of pricing errors that is basis-free, unit-invariant, and consistent with the orthogonality implied by APT. The decomposition generalizes zhang2024testing and extends the IPCA framework of kelly2019characteristics to accommodate both characteristic-driven and residual mispricing components, allowing for richer dynamics and greater interpretability of both components.

\paragraph{Methodology.} Using recent developments in low-rank and debiased estimation fan2022structural, chernozhukov2023inference, we derive closed-form estimators that are computationally efficient and theoretically grounded, and well suited for high-dimensional panels. Unlike previous iterative procedures, our estimators ensure valid orthogonality between pricing errors and factor betas and incorporate debiasing steps that are essential for inference.

\paragraph{Theoretical Contributions.} We establish a full inferential theory for characteristic loadings, inside alphas, and outside alphas. We relax the conventional assumption on the relative size of the cross-sectional dimension ($N$) and time-series length ($T$), and introduce bias-correction techniques that allow inference without requiring the restrictive assumption that $T/N\to \infty$ and the number of characteristics is finite, extending the asymptotic theory of high-dimensional factor models bai2003inferential, fan2016projected, chen2023semiparametric. These results place our framework on a firmer statistical footing than previous approaches and make it applicable to a wide range of empirical settings.

\paragraph{Empirical Findings.} Applying our methodology to U.S. stock returns and the same 36 firm characteristics considered by kelly2019characteristics and zhang2024testing from 2000 to 2019, we uncover new insights into the structure of pricing errors. We find strong evidence of both inside and outside alphas. Inside alphas exhibit persistent industry-level co-movements associated with fundamental drivers such as technology or finance sector shocks, while outside alphas capture transitory, firm-specific deviations consistent with behavioral or liquidity-based anomalies.

In summary, our framework unifies statistical and characteristic-based approaches, yielding both methodological innovations and substantive insights into the nature of pricing errors. It connects recent econometric innovations in high-dimensional inference with ongoing efforts in finance to rationalize the vast number of empirical return predictors harvey2016and, hou2020replicating, offering a richer and more interpretable decomposition of pricing errors, grounding estimation in modern econometric methods with rigorous inferential guarantees, and providing new empirical evidence on the structure of mispricings in equity markets.

The remainder of this paper is organized as follows. Section (ref) introduces the model of our paper and Section (ref) discusses the estimation and debiasing procedure. Section (ref) provides the inferential theory of our estimators. Section (ref) shows how our inferential theory can be applied to infer the US stock market and presents the empirical findings of our analysis. Finally, we conclude with a few remarks in Section (ref). All proofs and simulation studies are relegated to the supplement due to the space limit.

In what follows, we use $\|\cdot\|_{\rm F}$ and $\|\cdot\|$ to denote the matrix Frobenius norm and the spectral norm, respectively. For any vector $a$, $\|a\|$ denotes its $\ell_2$ norm. For any set $\mathcal{A}$, $|\mathcal{A}|$ is the number of elements in $\mathcal{A}$. We use $\otimes$ to denote the Kronecker product. $a \lesssim b$ means $|a|/|b| \leq C_1$ for some constant $C_1>0$ and $a \gtrsim b$ means $|a|/|b| \geq C_2$ for some constant $C_2>0$. $c \asymp d$ means that both $c/d $ and $d/c$ are bounded. $a \ll b$ indicates $|a| / |b| \rightarrow 0$ and $a \gg b$ indicates $|b| / |a| \rightarrow 0$. In addition, $I_n$ denotes the $n \times n$ identity matrix, $\textbf{1}_n$ denotes the $n \times 1$ vector of $1$, and $\boldsymbol{0}_{n\times m}$ denotes the $n \times m$ matrix consisting of zeros. In addition, $e_l$ is the $l$-th column of the identity matrix.

Modeling Two Types of Mispricing

Let $R_{t+1}$ the vector of excess returns on $N$ assets from period $t$ to $t+1$. A general factor pricing model posits that $$ R_{t+1} = \alpha_t + B_t f_{t+1} + E_{t+1}, $$ where $f_{t}$ is a $K \times 1$ vector of $K$ systematic factors, $B_t$ is the $N\times K$ matrix of factor loadings, and $E_{t+1}$ is an idiosyncratic noise vector. The vector $\alpha_t$ captures pricing errors (or “alphas”) and plays a critical role: under the Arbitrage Pricing Theory (APT), alphas should be orthogonal to factor exposures, i.e., $\alpha_t^\top B_t=0$. Otherwise, what appears as mispricing could simply reflect unmodeled factor risk.

The KPS Model and Its Limitations

kelly2019characteristics, henceforth KPS, proposed an influential specification in which both factor loadings and pricing errors are modeled as linear functions of firm characteristics. Specifically, let $X_t$ denote the $N \times L$ matrix of firm characteristics observed at time $t$. The KPS model imposes: $$ \alpha_t = X_{t} \eta, \qquad {\rm and}\qquad B_t = X_{t} \Gamma, $$ for parameter matrix $\Gamma\in {\mathbb R}^{L\times K}$ and $\eta\in {\mathbb R}^L$. This setup blends the strengths of statistical factor analysis with characteristic-based portfolio construction, allowing latent factors to be systematically linked to observable firm-level information.

While elegant, as pointed out in zhang2024testing, the KPS specification suffers from two major drawbacks. First, it does not enforce the orthogonality condition $\alpha_t^\top B_t = 0$. As a result, the so-called “pricing error” may in fact load on systematic factors, undermining its interpretation as pure mispricing. Second, by constraining $\alpha_t$ to lie in the span of $X_t$, the model rules out the possibility that some pricing errors are unrelated to the chosen set of characteristics. This restriction may omit economically meaningful forms of mispricing.

A Decomposition into Inside and Outside Alphas

To address these shortcomings, we propose decomposing the pricing error into two orthogonal components: $$\alpha_t=\alpha_{I,t}+\alpha_{O,t},$$ where \paragraph{Inside Alpha ($\alpha_{I,t}$):} the component of mispricing that is both orthogonal to the factor loadings and spanned by firm characteristics. This represents pricing errors that can be systematically related to observable fundamentals. Formally, $$\alpha_{I,t}=(I_N - P_{B_t}) X_{t} \eta,$$ where $P_{B_t} = B_t \left( B_t^\top B_t \right)^{-1} B_t^\top$ is the projection matrix onto the linear space spanned by $B_t$. It is clear that for any $\eta\in {\mathbb R}^L$, there exists $\eta_\perp\in {\mathbb R}^L$ such that $\eta_\perp^\top \Gamma=0$ and $$ (I_N-P_{B_t})X_t\eta=(I_N-P_{B_t})X_t\eta_\perp. $$ Thus, without loss of generality, we shall assume in what follows that $$\alpha_{I,t} = (I_N-P_{B_t})X_t\eta, \qquad {\rm and}\qquad \eta^\top \Gamma=0.$$

\paragraph{Outside Alpha ($\alpha_{O,t}$):} the residual mispricing component orthogonal to both $B_t$ and the span of $X_t$. This captures idiosyncratic pricing errors not explained by firm characteristics. We represent it as $$\alpha_{O,t}=B_t^o\delta_{o,t},$$ where $B_t^o$ is a basis for the subspace orthogonal to $X_t$, defined by

equation[equation omitted — 171 chars of source]

where $P_{X_t} = X_{t} (X_{t}^\top X_{t})^{-1} X_{t}^\top$ and $\Omega_{N \times (N-L)}$ is some full column rank matrix like $\left[ I_{N-L} \ \ \boldsymbol{0}_{(N-L) \times L} \right]^{\top}$.

This decomposition preserves the crucial orthogonality $\alpha_{I,t}^\top B_t=\alpha_{O,t}^\top B_t=0$ for both types of alphas by construction. Economically, it disentangles mispricing attributable to observable fundamentals (inside alpha) from residual, potentially behavioral or market-friction-driven anomalies (outside alpha).

The decomposition into inside and outside alphas has important economic implications. Inside alphas capture systematic mispricing tied to firm characteristics, which may reflect persistent risk premia omitted from standard factor models or inefficiencies linked to observable fundamentals. Outside alphas, in contrast, capture residual idiosyncratic deviations that cannot be traced back to known characteristics, and may be driven by liquidity frictions, behavioral biases, or institutional trading pressures. By separating the two, our framework provides both a sharper theoretical alignment with APT and a more flexible empirical tool for studying the sources of mispricing.

Comparison with zhang2024testing

Our decomposition is inspired by the approach of zhang2024testing, who also distinguishes between pricing errors within and outside the span of firm characteristics. However, there are important differences: \paragraph{Unit Invariance.} Zhang’s model can be sensitive to the scaling of firm characteristics, meaning that changing measurement units (e.g., dollars vs. millions) can alter the representation of alphas. Our formulation is invariant to such rescaling, making it more robust for empirical implementation as noted in Appendix (ref). \paragraph{Basis Dependence.} Zhang defines inside alpha as $\alpha_{I,t} = B_t^{I} \delta_{I}$ where $B_t^{I}$ is an orthonormal basis for the subspace orthogonal to $B_t$ but within the span of $X_t$, and $\delta_I$ is time-invariant. This construction depends critically on the choice of basis, which can change over time and affect the stability of estimation. In contrast, our specification $(I_N - P_{B_t}) X_{t} \eta$ avoids this indeterminacy and ensures that inside alphas are basis-free. \paragraph{Outside Alpha Dynamics.} Zhang assumes the outside pricing error $\alpha_{O,t}=B_t^{o}\delta_{o}$ for a time-invariant $\delta_{o}$, which is restrictive and may bias inference. We allow for more flexible dynamics by modeling $$ \delta_{o,t}=\zeta + \xi_{t} $$ where $\zeta$ captures a persistent component and $\xi_t$ is a sparse, time-varying shock. This assumption balances flexibility with tractability and reflects the plausible view that idiosyncratic mispricings may occasionally shift due to market conditions or firm-specific events.

Estimation and Debiasing

In this section, we describe how to estimate the parameters of the model introduced above -- namely, the characteristic-loading matrix $\Gamma$, the latent factors $f_t$, and the pricing error components $\alpha_{I,t}$ and $\alpha_{O,t}$. Our procedure builds on low-rank estimation methods but is carefully modified to ensure identification, orthogonality, and valid inference even when the number of characteristics $L$ is large relative to the number of assets $N$.

Estimation of $\Gamma$ and Latent Factors

Model Transformation and Motivation

Starting from our model

equation*[equation* omitted — 80 chars of source]

and substituting $\alpha_{I,t} = (I_N - P_{B_t}) X_t \eta$, $\alpha_{O,t} = B^o_t \delta_{o,t}$, and $B_t = X_t \Gamma$, we obtain

equation[equation omitted — 112 chars of source]

where

equation*[equation* omitted — 165 chars of source]

Equation (ref) shows that once we account for the part of the pricing error captured by firm characteristics, the transformed return dynamics are effectively governed by a low-rank structure: $R_{t+1}$ depends linearly on $X_t \Gamma$ through a small number of latent factors $\breve f_{t+1}$.

To exploit this structure, we pre-multiply both sides of (ref) by $(X_t^\top X_t)^{-1} X_t^\top$. This step removes the cross-sectional dependence induced by $X_t$ and yields

equation*[equation* omitted — 79 chars of source]

where $\ddot R_{t+1} = (X_t^\top X_t)^{-1} X_t^\top R_{t+1}$ and $\ddot E_{t+1} = (X_t^\top X_t)^{-1} X_t^\top E_{t+1}$. Averaging over time and centering give

equation[equation omitted — 94 chars of source]

where $f^{d}_{t+1} = \breve f_{t+1} - T^{-1}\sum_t \breve f_{t+1}$, and the superscript $d$ denotes de-meaned quantities. Equation (ref) reveals that $\ddot R^d = [\ddot R_2^d, \ldots, \ddot R_{T+1}^d]$ admits a low-rank factor structure, $\ddot R^d = \Gamma F^d + \ddot E^d$, with $\mathrm{rank}(\Gamma F^d) = K$. Here $ F^d = [f^d_{2},\cdots, f^d_{T+1}] $ and $\ddot{E}^d = [ \ddot{E}^d_{2}, \cdots , \ddot{E}^d_{T + 1}]$.

Initial Estimator via Low-Rank Approximation

We obtain an initial estimator $\tilde \Gamma$ as the top $K$ left singular vectors of $\ddot R^d$. This spectral estimator parallels the principal components estimator in classical factor analysis but operates in the transformed “characteristics space,” ensuring that the estimated factors are conditionally orthogonal given $X_t$.

This estimator is $\sqrt{NT}$-unbaised when $T \ll N$, but as $T$ grows relative to $N$, it can suffer from bias due to the finite-sample correlation between estimated factors and residuals. We next correct this bias using a debiasing step grounded in recent developments in low-rank inference.

Bias and Debiasing of $\Gamma$

Given $\tilde \Gamma$, we estimate the de-meaned factor matrix as $$ \tilde{F}^d =\operatorname*{arg\,min}_A\|\ddot{R}^d-\tilde{\Gamma}A\|_{\rm F}^2= \left(\tilde{\Gamma}^\top \tilde{\Gamma}\right)^{-1} \tilde{\Gamma}^\top \ddot{R}^d=H_F F^d+\left(\tilde{\Gamma}^\top \tilde{\Gamma}\right)^{-1} \tilde{\Gamma}^\top\ddot{E}^d, $$ where $$ H_F=\left(\tilde{\Gamma}^\top \tilde{\Gamma}\right)^{-1} \tilde{\Gamma}^\top\Gamma. $$ Similarly, $$ \tilde{\Gamma}=\operatorname*{arg\,min}_A\|\ddot{R}^d-A\tilde{F}^d\|_{\rm F}^2=\ddot{R}^d\tilde{F}^{d \top} (\tilde{F}^d\tilde{F}^{d \top})^{-1}=\Gamma H_\Gamma+\ddot{E}^d\tilde{F}^{d \top} (\tilde{F}^d\tilde{F}^{d \top})^{-1}, $$ where $$ H_\Gamma=F^d\tilde{F}^{d \top} (\tilde{F}^d\tilde{F}^{d \top})^{-1}. $$ The estimation error $\tilde{\Gamma}-\Gamma H_\Gamma$ can then be expressed as

gather[gather omitted — 417 chars of source]

The sample covariance of residuals, \[ \ddot E^d \ddot E^{d\top} = \sum_{t=1}^T (X_t^\top X_t)^{-1} X_t^\top E^d_{t+1} E_{t+1}^{d\top} X_t (X_t^\top X_t)^{-1}, \] has nonzero expectation and when $T/N$ does not vanish, the second term on the right hand side introduces non-negligible bias.

To correct for this, we approximate the expectation of the noise covariance by \[ \sum_{t=1}^T \hat \sigma_{t+1}^2 (X_t^\top X_t)^{-1}, \quad \text{where } \hat \sigma_{t+1}^2 = \frac{1}{N} \sum_{i=1}^N \hat \varepsilon_{i,t+1}^2, \] and $\hat \varepsilon_{i,t+1}$ are residuals from the current fit: \[ \hat \varepsilon_{i,t+1} = r_{i,t+1} - (\hat \alpha_{O,it} + \tilde{\alpha}_{I,it} + x_{it}^\top \tilde \Gamma \tilde f_{t+1}). \] Subtracting this estimated bias yields the debiased estimator:

equation*[equation* omitted — 207 chars of source]

The corresponding debiased estimate of the latent factors is

equation*[equation* omitted — 92 chars of source]

This procedure removes the leading-order bias term of $\tilde \Gamma$ that arises when $T/N$ is not small. In Section (ref), we show that the resulting estimator admits a valid asymptotic normal distribution under mild regularity conditions, allowing for inference on both $\Gamma$ and the characteristic loadings even when the number of characteristics $L$ grows with $N$.

Estimation of Pricing Errors

Having estimated $\hat \Gamma$ and $\hat F^d$, we next turn to the estimation of inside and outside alphas.

Inside Alpha ($\alpha_{I,t}$)

By definition, \[ \alpha_{I,t} = (I_N - P_{B_t}) X_t \eta = (P_{X_t} - P_{B_t})X_t \eta, \quad (P_{X_t} - P_{B_t})(\alpha_{O,t} + B_t f_{t+1})=0. \] A direct estimator of this quantity is \[ \hat \alpha_{I,t} = (P_{X_t} - P_{X_t \hat \Gamma}) R_{t+1}. \] However, the convergence rate of this estimator is $\sqrt{L}/ \sqrt{N}$, which can be slow when $L$ is large. To obtain a more efficient estimator, we exploit the transformed model \[ \ddot R_{t+1} = \eta + \Gamma \breve f_{t+1} + \ddot E_{t+1}, \] which implies \[ (I_L - P_\Gamma) \ddot R_{t+1} = \eta + (I_L - P_\Gamma) \ddot E_{t+1}. \] Hence, we can estimate $\eta$ by \[ \hat \eta = (I_L - P_{\hat \Gamma}) \bar{\ddot{R}} \quad \text{where } \bar{\ddot{R}} = \frac{1}{T} \sum_{t=1}^T \ddot R_{t+1}. \] Finally, substituting back yields a compact expression for inside alpha: \[ \hat \alpha_{I,t} = (I_N - P_{X_t \hat \Gamma}) X_t \hat \eta = (I_N - P_{X_t \hat \Gamma}) X_t \bar{\ddot{R}}. \] This estimator enforces the orthogonality between $\alpha_{I,t}$ and factor loadings by construction and is computationally straightforward, requiring only matrix multiplications.

Outside Alpha ($\alpha_{O,t}$)

For the outside alpha, the estimation procedure consists of two steps. Note that $X_t^\top B^o_t = 0$, so projecting $R_{t+1}$ onto the orthogonal basis yields: \[ (B_t^{o\top} B_t^o)^{-1} B_t^{o\top} R_{t+1} = \delta_{o,t} + (B_t^{o\top} B_t^o)^{-1} B_t^{o\top} E_{t+1}. \] Thus, an initial estimator of $\delta_{o,t}$ is \[ \tilde \delta_{o,t} = (B_t^{o\top} B_t^o)^{-1} B_t^{o\top} R_{t+1}. \] Because we allow for a time-varying but sparse component $\xi_t$ such that $\delta_{o,t} = \zeta + \xi_t$, we estimate the persistent part $\zeta$ by time averaging: \[ \tilde \zeta = \frac{1}{T} \sum_{t=1}^T \tilde \delta_{o,t}, \] and then obtain a sparsity-regularized estimate of the transitory part via hard thresholding: \[ \tilde \xi_{t,q} =

cases\tilde \delta_{o,t,q} - \tilde \zeta_q, & if |\tilde \delta_{o,t,q} - \tilde \zeta_q| \geq \rho_t,\\ 0, & otherwise,

\] where the threshold $\rho_t$ is chosen proportional to $\sqrt{(\log NT)/N}$ according to the analysis from Section (ref). Additionally, since $\tilde \zeta$ has a bias term $\bar{\xi}= \frac{1}{T} \sum_{t=1}^T \xi_t$ in it, we further refine the estimator using $\tilde \xi_{t}$: $$ \hat{\zeta} = \tilde \zeta - \frac{1}{T} \sum_{t=1}^T \tilde \xi_{t} . $$ Similarly, we refine the estimator $\tilde \xi_{t,q}$ when $\tilde \xi_{t,q} \neq 0$: \[ \hat \xi_{t,q} =

cases\tilde \delta_{o,t,q} - \hat \zeta_q, & if \tilde \xi_{t,q} \neq 0,\\ 0, & if \tilde \xi_{t,q} = 0.

\] The final estimator of outside alpha is then \[ \hat \alpha_{O,t} = B^o_t (\hat \zeta + \hat \xi_t) . \]

Estimation Procedure

We summarize the complete estimation procedure for $\Gamma$, the latent factors, and the two pricing error components below. The procedure relies only on standard linear algebra operations (matrix multiplications, singular value decomposition, and projection), and scales well for large panels.

breakablealgorithm{\fname@algorithm \thealgorithm #Estimation and Debiasing of Conditional Factor Model} \ifx\relax#\relax\relax \addcontentsline{loa}{algorithm}{\numberline{\thealgorithm}#Estimation and Debiasing of Conditional Factor Model} \else \addcontentsline{loa}{algorithm}{\numberline{\thealgorithm}#\relax} \fi \kern2pt\hrule\kern2pt \begin{algorithmic}[1] \Require Excess returns $\{R_{t+1}\}_{t=1}^T$, firm characteristics $\{X_t\}_{t=1}^T$, number of factors $K$, threshold $\rho_t$. \State Step 1: Transformation and Initial Estimation of $\Gamma$ \State Compute $\ddot R_{t+1} = (X_t^\top X_t)^{-1} X_t^\top R_{t+1}$ and demean across $t$ to form $\ddot R^d$. \State Obtain top $K$ left singular vectors of $\ddot R^d$: $\tilde \Gamma \leftarrow \text{SVD}(\ddot R^d)$. \State Compute $\tilde F^d = (\tilde \Gamma^\top \tilde \Gamma)^{-1} \tilde \Gamma^\top \ddot R^d$. \State Step 2: Initial Estimation of $\alpha_{I,t}$ and $f_{t+1}$ \State Estimate $\tilde \eta = (I_L - P_{\tilde \Gamma}) \bar{\ddot{R}}$, $\bar{\ddot{R}} = T^{-1}\sum_t \ddot R_{t+1}$. \State Compute $\tilde \alpha_{I,t} = (I_N - P_{X_t \tilde \Gamma}) X_t \tilde \eta$. \State Compute $\tilde{f}_{t+1} = (\tilde \Gamma^\top \tilde \Gamma)^{-1} \tilde \Gamma^\top \ddot R_{t+1} + (\tilde \Gamma^\top X_t^\top X_t \tilde \Gamma)^{-1} \tilde \Gamma^\top X_t^\top X_t \tilde \eta$. \State Step 3: Debiasing of $\Gamma$ \State Compute residuals $\hat \varepsilon_{i,t+1} = r_{i,t+1} - (\hat{\alpha}_{O,it} + \tilde{\alpha}_{I,it} + x_{it}^\top \tilde \Gamma \tilde f_{t+1} )$. \State Estimate $\hat \sigma_{t+1}^2 = N^{-1}\sum_i \hat \varepsilon_{i,t+1}^2$. \State Apply bias correction: \[ \hat \Gamma = \tilde \Gamma - \Big(\sum_t \hat \sigma_{t+1}^2 (X_t^\top X_t)^{-1}\Big) \tilde \Gamma (\tilde \Gamma^\top \tilde \Gamma)^{-1}(\tilde F^d \tilde F^{d\top})^{-1}. \] \State Compute $\hat F^d = (\hat \Gamma^\top \hat \Gamma)^{-1} \hat \Gamma^\top \ddot R^d$. \State Step 4: Inside Alpha \State Repeat Step 2 with $\hat{\Gamma}$ to derive $\hat{\alpha}_{I,t}$ and $\hat{f}_{t+1}$. \State Step 5: Outside Alpha \State Construct $B^o_t = X^o_t[(X^{o\top}_t X^o_t)/N]^{-1/2}$, $X^o_t = (I_N - P_{X_t})\Omega_{N \times (N-L)}$. \State Compute $\tilde \delta_{o,t} = (B_t^{o\top}B_t^o)^{-1}B_t^{o\top}R_{t+1}$. \State Estimate $\tilde \zeta = T^{-1}\sum_t \tilde \delta_{o,t}$. \State Apply hard thresholding: \[ \tilde \xi_{t,i} = \begin{cases} \tilde \delta_{o,t,q} - \tilde \zeta_i, & |\tilde \delta_{o,t,q} - \tilde \zeta_q| \ge \rho_t,\\ 0, & \text{otherwise.} \end{cases} \] \State Refinement: estimate $\hat{\zeta} = \tilde \zeta - \frac{1}{T} \sum_{t=1}^T \tilde \xi_{t}$ and \[ \hat \xi_{t,q} = \begin{cases} \tilde \delta_{o,t,q} - \hat \zeta_q, & \text{if } \tilde \xi_{t,q} \neq 0,\\ 0, & \text{if } \tilde \xi_{t,q} = 0. \end{cases} \] \State Compute $\hat \alpha_{O,t} = B^o_t(\hat \zeta + \hat \xi_t)$. \Ensure Outputs: Debiased $\hat \Gamma$, latent factors $\hat f_{t+1}$, inside alpha $\hat \alpha_{I,t}$, outside alpha $\hat \alpha_{O,t}$. \end{algorithmic}

By transforming returns into characteristic space and exploiting low-rank structure, we obtain closed-form estimators for both $\Gamma$ and the pricing errors. The bias-correction step ensures valid inference even when $T$ is not small relative to $N$. Conceptually, our approach differs from the algorithmic methods in zhang2024testing, which iteratively solve first-order conditions without theoretical guarantees. Instead, our estimators admit clear analytical forms, are grounded in the recent theory of debiased low-rank estimation, and directly link to the inferential results in Section (ref).

Inference and Asymptotic Theory

This section develops the inferential theory for our estimators of characteristic loadings, factors, and pricing errors. While the estimation procedure in Section (ref) yields closed-form solutions, valid inference requires understanding their asymptotic behavior as both the cross-sectional and time-series dimensions grow. We show that the estimators admit standard Gaussian limits under mild regularity conditions, allowing conventional hypothesis testing even when the number of firm characteristics increases with the sample size.

Setup and Regularity Conditions

We first present a sequence of assumptions that ensure well-behaved moments, identification, and dependence properties of the data-generating process. For clarity, we group these conditions by theme.

assumption[Characteristics and Identification] Each firm $i$ at time $t$ is associated with an $L$-dimensional vector of characteristics $x_{it}$. \begin{itemize} • The second moments are uniformly bounded: $E[x_{it,l}^2] \le C$ for some constant $C>0$. • The cross-sectional covariance matrix $Q_t = N^{-1}\sum_{i=1}^N x_{it}x_{it}^\top$ has eigenvalues bounded away from zero and infinity: \[ c_1 < \psi_{\min}(Q_t) \le \psi_{\max}(Q_t) < c_2, \] for some positive constants $c_1$ and $c_2$, with probability approaching one. Here $\psi_{\min}(\cdot)$ and $\psi_{\max}(\cdot)$ are the smallest and largest nonzero eigenvalues, respectively. \end{itemize}

Assumption (ref) ensures that characteristics are sufficiently informative and non-collinear. It parallels the “pervasive” condition in classical factor models fan2016projected, chen2023semiparametric and is relatively mild since $L \ll N$ in most applications.

assumption[Factors and Loadings] Let $\Gamma$ denote the $L\times K$ matrix of characteristic loadings and $f_t$ the $K$-dimensional latent factor. \begin{itemize} • $\Gamma^\top \Gamma$ is well-conditioned: $c_1 < \psi_{\min}(\Gamma^\top\Gamma) \le \psi_{\max}(\Gamma^\top\Gamma) < c_2$ for some positive constants $c_1$ and $c_2$. • $E[\|f_t\|^4] < C_1$ for some positive constant $C_1$. • The de-meaned factor covariance satisfies $T^{-1}F^d (F^d)^\top \xrightarrow{p} \Sigma_f$, where $\Sigma_f$ is positive definite. • The eigenvalues of $(\Gamma^\top\Gamma)\Sigma_f$ are distinct. • There exists a constant $C_2>0$ such that $E[\|B_{it}\|^2] \le C_2$ for all $i,t$. • Identification: $\eta^\top\Gamma = 0$ and $\|\eta\|\le C_3$ for some constant $C_3>0$. \end{itemize}

These conditions guarantee identification of the factors and their characteristic-based loadings. Condition (i) is similar to the “pervasive” condition on factor loadings and common in the factor model literature. See, e.g., chen2023semiparametric. Conditions (ii) - (iv) ensure factor uniqueness up to rotation and are also typical in the factor model literature. See, e.g., bai2003inferential,fan2016projected,chen2023semiparametric. Condition (vi) enforces the orthogonality of inside alphas to factor loadings, which is essential for identifying pricing errors. See, also, kelly2019characteristics,kim2021arbitrage,chen2023semiparametric.

assumption[Idiosyncratic Noise] Conditional on $(x_{it},f_{t+1})$, the idiosyncratic component $\epsilon_{it+1}$ satisfies: \begin{itemize} • $E[\epsilon_{it+1}] = 0$ and $E[\epsilon_{it+1}^2] = \sigma_{t+1}^2$; • Sub-Gaussianity: $E[\exp(s\epsilon_{it+1})] \le \exp(C_1 s^2\sigma_{t+1}^2)$ for all $s\in\mathbb{R}$; • Independence across $i$ and weak dependence across $t$: $\max_{i,t}\sum_{s}|\mathrm{Cov}(\epsilon_{it},\epsilon_{is})| \le C_2.$ \end{itemize}

Assumption (ref) allows for heteroskedasticity and mild serial dependence, both prevalent in asset-return data. Sub-Gaussianity simplifies the derivations without excluding heavy-tailed behavior under weak dependence.

assumption[Sparsity of Outside Alphas] Let $\delta_{o,t} = \zeta + \xi_t$ denote the outside-alpha component. Then, for each coordinate $q$, \[ \frac{1}{T}\sum_{s=1}^T |\xi_{s,q}| \ll \sigma_{t+1} \frac{\sqrt{\log(NT)}}{\sqrt{N}}, \quad \frac{\sigma_{t+1} \sqrt{\log(NT)}}{|\xi_{t,q}|\sqrt{N}} \to 0 \text{ for } q \in D_t, \] and $(\log N / N)|D_t|\to 0$ where $D_t = \{ 1 \leq q \leq N-L : \xi_{t,q} \neq 0 \}$.

This assumption imposes sparsity on transitory mispricing shocks, consistent with the view that only a small subset of firms experience idiosyncratic pricing deviations at any given time.

assumption[Central Limit Conditions] Define \[ Q_f = T^{-1}\sum_t f_t^df_t^{d\top},\quad Q_t^B = N^{-1}\sum_i B_{it}B_{it}^\top,\quad Q_t^{a,B} = N^{-1}\sum_i a_{it}B_{it}^\top, \] where $a_{it} = \eta^\top x_{it}$. Conditioning on $(x_{it},f_{t+1})_{i\leq N, t \leq T}$, \begin{align*} & (i) \ \ \frac{1}{\sqrt{NT}} \sum_{i=1}^N \sum_{t=1}^T \left(e_l^\top Q_t^{-1}x_{it}\right) f_{t+1}^d \epsilon_{i,t+1} \to_d \mathcal{N}\left( 0, \Sigma_{xf,l} \right), \\ &(ii) \ \ \frac{1}{\sqrt{NTL}} \sum_{j=1}^N \sum_{s=1}^T g_{it,js} \epsilon_{j,s+1} \to_d \mathcal{N}\left( 0, \sigma_{I,it}^{2} \right), \\ &(iii) \ \ \frac{1}{\sqrt{N}}\sum_{j=1}^N B^o_{t,jq} \epsilon_{j,t+1} \to_d \mathcal{N}\left( 0, \sigma_{\delta,qt}^{2} \right),\\ &(iv) \ \ \sigma_{o,it}^{-1} \left( \frac{1}{NT} \sum_{j=1}^N \sum_{s=1}^T B_{t,i}^{o \top} B_{s,j}^{o} \epsilon_{j,s+1} + \frac{1}{N} \sum_{j=1}^N \left( \sum_{q \in D_t} B_{t,iq}^o B_{t,jq}^o \right) \epsilon_{j,t+1} \right) \to_d \mathcal{N}\left( 0, 1 \right),\\ &where \\ & g_{it,js} = \left[1 - \left( Q^{a,B}_t (Q_t^B)^{-1} + \bar{\breve{f}}^\top \right) \left( Q^{f}\right)^{-1} f_{s+1}^d \right] \left(x_{it}^\top Q_t^{-1} x_{js} - B_{it}^\top (Q_t^{B})^{-1} B_{js} \right) \\ & \qquad \ \ - \left( B_{it}^\top (Q^{B}_t)^{-1} (Q^{f})^{-1} f_{s+1}^d \right) \left( a_{js} - Q^{a,B}_t (Q^{B}_t)^{-1} B_{js} \right) , \end{align*} for some positive values $\sigma_{I,it}$, $\sigma_{\delta,qt}$, $\sigma_{o,it}$, and a positive definite matrix $\Sigma_{xf,l}$.

Assumption (ref) provides a high-dimensional Lindeberg-type CLT that accommodates growing $L$ and heteroskedastic, weakly dependent errors, forming the statistical backbone of our inference. Because $(g_{it,js},B^{o}_{js})_{j\leq N, s \leq T}$ are functions of $(x_{js},f_{s+1})_{j \leq N, s \leq T}$, this assumption requires a weak dependence in the noise term, $(\epsilon_{js})_{j\leq N, s \leq T}$. For example, if $\epsilon_{it}$ are independent across $i$ and $t$ with $\mathbb{E}[\epsilon_{it}^2] = \sigma_{t}^2$, the condition will be satisfied by the Lindeberg theorem with the variances:

align[align omitted — 567 chars of source]

where $\bar{\sigma}^2 = \frac{1}{T} \sum_{s=1}^T \sigma_{s+1}^2$. Because the size of $\left\VertB_{t,i}^o\right\Vert^2$ is close to $N-L$ and $B_{t,iq}^{o}$ is generally bounded, we can say roughly $\sigma_{o,it}^2 \asymp \frac{1}{T} + \frac{|D_t|}{N}$. In Assumption (ii), we adjusted the scale by including $\sqrt{L}$ in the denominator to avoid divergence. Without difficulty, we can show that the variances $\sigma_{I,it}^2$, $\sigma_{\delta,qt}^2$, and $\Sigma_{xf,l}$ are bounded under our weak dependence assumption.

We are now in position to state the distributional properties of various parameters.

Asymptotic Distributions

We first derive the asymptotic distribution of the characteristic-loading matrix $\Gamma$. The spectral estimator is consistent but biased when $T$ is not small relative to $N$. The debiased estimator corrects this bias and enables valid inference.

theorem[Asymptotic Normality of $\Gamma$] Suppose that Assumptions (ref) -- (ref), (ref) (i) are satisfied.\\ (a) If $L/N \rightarrow 0$, $T/N \rightarrow 0$, and $T/\left( \frac{N}{L} \right)^{20} \rightarrow 0$, For each $1 \leq l \leq L$, $$ \sqrt{NT} \left( \tilde{\gamma}_l - H^\top_{\Gamma} \gamma_l \right) \to_d \mathcal{N} \left(0, \boldsymbol{H}^\top \Sigma_f^{-1} \Sigma_{xf,l} \Sigma_f^{-1}\boldsymbol{H} \right), $$ where $\boldsymbol{H}$ is the limit of $H_{\Gamma}$ and $H^{-1}_{F}$.\\ (b) If $L/N \rightarrow 0$, $T/N^3 \rightarrow 0$, and $\left(\frac{T}{N}\right) / \left( \frac{N}{L} \right)^{20} \rightarrow 0$, we have for each $1 \leq l \leq L$, $$ \sqrt{NT} \left( \hat{\gamma}_{l} - H^\top_{\Gamma} \gamma_l \right) \to_d \mathcal{N} \left(0, \boldsymbol{H}^\top \Sigma_f^{-1} \Sigma_{xf,l} \Sigma_f^{-1}\boldsymbol{H} \right). $$

Here, we present the asymptotic normality of each $\gamma_l$ rather than that of $\Gamma$ because the dimension of $\Gamma$ diverges when $L \rightarrow \infty$. The conditions for (b), $T/N^3 \rightarrow 0$ and $\left(\frac{T}{N}\right) / \left( \frac{N}{L} \right)^{20} \rightarrow 0$, are milder than the conditions for (a), $T/N \rightarrow 0$ and $T/\left( \frac{N}{L} \right)^{20} \rightarrow 0$. Hence, when $N$ is not much larger than $T$ (or smaller than $T$), the debiased estimator $\hat{\gamma}_{l}$ can be useful. Theorem (ref) shows that the debiased estimator is asymptotically normal even when the time dimension is moderately large relative to $N$. This permits standard inference on the relationship between firm characteristics and factor exposures in typical empirical panels.

We next consider the component of mispricing explained by firm characteristics but orthogonal to factors.

theorem[Asymptotic Normality of $\alpha_{I,it}$] Suppose that Assumptions (ref) -- (ref), (ref) (ii) are satisfied.\\ (a) If $L/N \rightarrow 0$, $T/N \rightarrow 0$ and $T/\left( \frac{N}{L} \right)^{20} \rightarrow 0$, we have $$ V_{I,it}^{-1/2} \left( \tilde{\alpha}_{I,it} - \alpha_{I,it} \right) \to_d \mathcal{N}(0,1), $$ where $V_{I,it} = \sigma_{I,it}^2L/NT$.\\ (b) If $L/N \rightarrow 0$, $T/N^3 \rightarrow 0$, $\left(\frac{T}{N}\right) / \left( \frac{N}{L} \right)^{20} \rightarrow 0$, then we have $$ V_{I,it}^{-1/2} \left( \hat{\alpha}_{I,it} - \alpha_{I,it} \right) \to_d \mathcal{N}(0,1). $$

Note that, because the convergence rate of $\hat{\alpha}_{I,it}$ is $\sqrt{L}/\sqrt{NT}$, the test using this estimator can have a higher power than that using $(P_{X_t} - P_{X_t \hat \Gamma}) R_{t+1}$ as an estimator. Similarly to Theorem (ref), the inferential theory based on $\hat{\alpha}_{I,it}$ requires milder conditions for $N$ and $T$ compared to that of $\tilde{\alpha}_{I,it}$. The convergence rate of $\hat{\alpha}_{I,it}$ is $\sqrt{L/(NT)}$, yielding high efficiency even in high-dimensional settings. This enables powerful tests for systematic pricing errors linked to observable fundamentals.

We now analyze the residual component $\alpha_O$, orthogonal to both factors and firm characteristics. The key intermediate parameter is the coefficient vector $\delta_{o,t}$.

theorem[Asymptotic Normality of $\delta_{o,t}$] Suppose that Assumptions (ref), (ref), and (ref) (iii) are satisfied. Then, we have $$ V_{\delta,tq}^{-1/2} \left( \tilde{\delta}_{o,t,q} - \delta_{o,t,q} \right) \to_d \mathcal{N}(0,1), \qquad \text{where } V_{\delta,tq} = \sigma_{\delta,qt}^2/N. $$

Importantly, this result is still valid without the assumptions regarding sub-Gaussianity and cross-sectionally independent noise as long as noise is weakly dependent across $i$. Moreover, it does not require the sparsity condition. This result can be utilized to conduct an outside alpha test whose null hypothesis is $H_o : \delta_{o,t} = 0$ for all $t$, because based on the asymptotic normality above, we can have $$ \mathbb{P} \left( \max_{t \leq T,q \leq N - L } \left\vert \hat{V}^{-1/2}_{\delta,tq} \left(\tilde{\delta}_{o,t,q} - \delta_{o,t,q} \right)\right\vert > \Phi^{-1} (1 - a /(2T(N-L)) ) \right) \leq a + o(1), $$ e.g., belloni2018high. Theorem (ref) allows testing for the existence of outside alphas via the null $H_0: \delta_{o,t} = 0$ for all $t$. The test can be implemented using extreme-value approximations as in belloni2018high, providing a way to detect residual anomalies beyond characteristic-based mispricing.

To extend inference from $\delta_{o,t}$ to $\alpha_{O,it}$, we impose mild regularity conditions controlling approximation bias.

assumption[Regularity for Outside-Alpha Bias Control] Conditional on $(x_{it})$, the following hold: \begin{itemize} • $\frac{|D_t|}{NT}\frac{1}{|D_t|}\sum_{q\in D_t}B_{o,t,iq}^2 \ll \sigma_{o,it}^2;$$\frac{1}{NT}\sum_{s=1}^T |D_s|\frac{1}{|D_s|}\sum_{q\in D_s\setminus D_t} B_{o,t,iq}^2 \ll \sigma_{o,it}^2;$$\frac{1}{T}\sum_{s=1}^T |D_s| \left(\frac{1}{|D_s|}\sum_{q\in D_s\setminus D_t} B_{o,t,iq}\bar{\xi}_q\right) \ll \sigma_{o,it}.$ \end{itemize}

Assumption (ref) is mild and automatically satisfied when the number of firms with nonzero transitory shocks is small relative to $N$ and $T$. It ensures that cross-sectional spillovers from temporary idiosyncratic shocks are asymptotically negligible.

In the case of the first relation, the order of the left side is roughly $\frac{|D_t|}{NT}$ while that of $\sigma_{o,it}^2$ is roughly $\frac{1}{T} + \frac{|D_t|}{N}$ as we noted in (ref). Hence, when $N,T \rightarrow \infty$, it would be satisfied. Similarly, because the order of the left side of the second relation is roughly $\frac{|\bar{D}_\star|}{NT}$ where $|\bar{D}_\star| = \frac{1}{T} \sum_{s=1}^T |D_s|$, the second condition would be satisfied. Lastly, the third relation would be satisfied by the sparsity of $\xi_{t}$. For instance, if $\{\xi_{t,q}\}$ is nonzero at a small number of time periods by the sparsity, the order of $\bar{\xi}_{q}$ would be roughly $\frac{1}{T}$. Hence, the order of the left side is roughly $\frac{|\bar{D}_\star|}{T}$ and less than $\frac{1}{\sqrt{T}} + \frac{\sqrt{|D_t|}}{\sqrt{N}}$, when $|\bar{D}_\star|$ is small due to the sparsity of $\xi$. Then, under the above conditions, we have the following asymptotic normality.

theorem[Asymptotic Normality of $\alpha_{O,it}$] Suppose that Assumptions (ref), (ref), (ref), (ref) (iv), (ref) are satisfied. Additionally, if $(\epsilon_{it})_{i\leq N , t \leq T}$ is dependent across $t$, assume that $$ \mathbb{E} \left[ \left\vert \frac{1}{\sqrt{NT}} \sum_{s=1}^T \sum_{j=1}^N B_{s,jq}^{o} \epsilon_{j,s+1} \right\vert^\alpha \right] \quad \text{is bounded} $$ for some integer $\alpha \geq 1$ where $N = O(T^{\alpha/2})$. Then, we have $$ V_{o,it}^{-1/2} \left( \hat{\alpha}_{O,it} - \alpha_{O,it} \right) \to_d \mathcal{N}(0,1), $$ where $V_{o,it} = \sigma_{o,it}^2 $ is in Assumption (ref).

Theorem (ref) completes the inferential theory by establishing Gaussian limits for the outside-alpha estimator. Together with Theorems (ref)--(ref), it provides a comprehensive inferential framework for both systematic and idiosyncratic components of mispricing.

Our inferential results provide the following empirical tools:

itemize• Testing characteristic relevance: Wald-type tests on each $\gamma_l$ identify which firm attributes significantly explain factor exposures. • Evaluating systematic mispricing: Tests on $\alpha_I$ detect whether pricing errors align with observable fundamentals. • Detecting residual anomalies: Tests on $\alpha_O$ assess whether idiosyncratic mispricing remains after accounting for all systematic sources.

These tools yield a unified econometric framework that is both theoretically grounded and empirically tractable, enabling rigorous inference in large-scale panels of asset returns with rich firm characteristics.

Application to U.S. Stock Data

We now illustrate the empirical relevance of our framework by applying it to U.S. equity returns. This section evaluates the magnitude, dynamics, and economic interpretation of both inside and outside alphas estimated using our methodology. The goal is to demonstrate how the inferential theory developed in Section (ref) translates into concrete insights about mispricing and factor structure in the cross-section of stock returns.

Data and Methods

\paragraph{Data.} We examine monthly excess returns on U.S.\ stocks from January 2000 through December 2019, yielding $T = 240$ time periods. Our data are drawn from the same sources as zhang2024testing, covering $N = 973$ continuously observed firms. We use the $36$ firm characteristics from kelly2019characteristics and chen2023semiparametric, augmented by a constant, as potential explanatory variables. These characteristics span size, value, profitability, investment, momentum, liquidity, and trading frictions, and are detailed in Appendix (ref).

Following standard practice, each characteristic $x_{i,t,l}$ is transformed into a rank-normalized variable across firms at time $t$: \[ x_{i,t,l} = -0.5 + \frac{z_{i,t,l}}{N}, \] where $z_{i,t,l}$ denotes the cross-sectional rank of firm $i$. This transformation mitigates the influence of outliers and ensures scale invariance.

\paragraph{Estimation.} We implement the debiased estimation procedure from Section (ref). Given that $N$ is of the same order of magnitude as $T$, we employ the debiased estimators $\widehat{\Gamma}$ and $\widehat{\alpha}_{I,it}$ to obtain valid inference under finite-sample bias. The rank of $\Gamma$ (the number of latent factors $K$) is selected using the eigenvalue-ratio criterion proposed by chen2023semiparametric. For the orthogonal complement $X_t^o$ in constructing $B_t^o$, we adopt the specification in Section (ref). The threshold parameter $\rho_t$ in the sparse outside-alpha estimation is set to \[ \rho_t = \widehat{\sigma}_{t+1} \frac{(\log NT)^{0.6}}{\sqrt{N}}, \] where $\widehat{\sigma}_t^2$ is the cross-sectional variance of residuals at time $t$. All variances used in inference are estimated under the assumption of independence and heteroskedasticity across time.

Empirical Findings

We now examine the estimated pricing errors and factor structure implied by the model. Throughout, we report results for $K = 1$ to $10$, highlighting $K = 5$ as the benchmark case selected by the data.

Testing for Outside Alphas

We first test whether the model admits a nontrivial outside-alpha component ($\alpha_O$) and whether these effects vary over time. The corresponding hypotheses are \[ H_0^{(1)}: \delta_{o,t} = 0 \quad \text{for all } t, \qquad H_0^{(2)}: \delta_{o,t} = \delta_o \quad \text{for all } t. \] The test statistics follow from Theorem (ref): \[ T\text{-stat}_1 = \max_{t \le T, q \le N-L} |\widehat{\tau}_{1,tq}|, \quad \widehat{\tau}_{1,tq} = \widehat{V}_{\delta,tq}^{-1/2}\,\tilde{\delta}_{o,t,q}, \] \[ T\text{-stat}_2 = \max_{t \le T, q \le N-L} |\widehat{\tau}_{2,tq}|, \quad \widehat{\tau}_{2,tq} = \widehat{V}_{\delta,tq}^{-1/2}\!\left( \tilde{\delta}_{o,t,q} - \frac{1}{T}\sum_{s=1}^T \tilde{\delta}_{o,s,q} \right). \] Table (ref) reports these statistics for $K = 1,\dots,10$. Under the null, the extreme-value bound from belloni2018high provides asymptotically valid $p$-values: \[ P\!\left( \max_{t,q} |\widehat{\tau}_{tq}| > \Phi^{-1}(1 - a/(2T(N-L))) \right) \le a + o(1). \]

table[table omitted — 1,173 chars of source]

As shown in Table (ref), both $T\text{-stat}_1$ and $T\text{-stat}_2$ exceed the 1% critical value (5.47) by a wide margin across all $K$. The associated $p$-values are below $10^{-10}$, decisively rejecting both null hypotheses. Hence, the data exhibit statistically and economically significant outside alphas, and these effects are time-varying. This finding underscores that idiosyncratic mispricing persists beyond the span of firm characteristics and evolves dynamically over time.

Testing for Inside and Outside Pricing Errors

Next, we test for the joint existence of both inside and outside alphas at the firm-month level using \[ T\text{-stat}_O = \max_{i,t} |\widehat{\tau}_{O,it}|, \quad \widehat{\tau}_{O,it} = \widehat{V}_{O,it}^{-1/2}\widehat{\alpha}_{O,it}, \] \[ T\text{-stat}_I = \max_{i,t} |\widehat{\tau}_{I,it}|, \quad \widehat{\tau}_{I,it} = \widehat{V}_{I,it}^{-1/2}\widehat{\alpha}_{I,it}. \] The null hypothesis is $H_0: \alpha_{\iota,it}=0$ for all $(i,t)$ and $\iota \in \{O,I\}$. Critical values are again obtained using the extreme-value approximation in belloni2018high. Table (ref) reports the resulting statistics and model $R^2$ values.

table[table omitted — 1,860 chars of source]

For all $K$, both $T\text{-stat}_O$ and $T\text{-stat}_I$ reject the null hypothesis at significance levels below $10^{-10}$. Hence, both inside and outside alphas are pervasive in the cross-section of returns. The explanatory power of the model increases with the number of factors, with $R^2$ rising from $6.3\%$ for $K=1$ to $27.2\%$ for $K=10$. At the empirically selected $K=5$, the model explains $14.4\%$ of total variation in returns, suggesting a balance between parsimony and explanatory strength. These results affirm the empirical relevance of decomposing mispricing into characteristic-driven and residual components.

Dynamics and Economic Interpretation of Inside Alphas

We now explore the temporal and cross-sectional behavior of the inside-alpha component $\widehat{\alpha}_I$, which captures systematic mispricing linked to firm characteristics but orthogonal to factor betas.

Figures (ref)-(ref) plot the estimated monthly inside alphas for representative firms and sector averages, together with 95% confidence intervals adjusted via the false discovery rate (FDR) control of benjamini2001control. In what follows, we discuss several representative patterns.

\paragraph{Technology Sector.} Figure (ref) depicts $\widehat{\alpha}_I$ for Apple and Microsoft. Both exhibit pronounced co-movement: alphas were low during the early 2000s following the dot-com crash, remained resilient through the 2008 financial crisis, and trended upward post-2010. The alignment of $\alpha_I$ across these firms suggests that inside alphas capture persistent industry-level fundamentals rather than firm-specific anomalies.

figure[figure omitted — 853 chars of source]

\paragraph{Financial Sector.} Figure (ref) plots $\widehat{\alpha}_I$ for J.P.\ Morgan Chase and Bank of America. Both series decline sharply during the 2007–2008 crisis, indicating that beyond the market-wide factor exposure, financial firms suffered deterioration in fundamentals not captured by standard betas. Post-crisis, their inside alphas recover gradually and move in tandem, again pointing to a strong sectoral component.

figure[figure omitted — 529 chars of source]

\paragraph{Energy and Consumer Sectors.} Figures (ref) display $\widehat{\alpha}_I$ for representative oil and consumer goods firms. Within-industry alphas exhibit substantial co-movement, most notably for ExxonMobil and Chevron, consistent with shared exposure to oil prices and global supply conditions.

figure[figure omitted — 553 chars of source]

\paragraph{Industry-Level Evidence.} Figure (ref) and Figure (ref) summarize sector-level average inside alphas based on NAICS classifications. Inside alphas display clear industry patterns: the IT sector shows sharp declines during the dot-com crash but little response to the financial crisis; the petrochemical and finance sectors experience simultaneous declines during 2008–2009; and the healthcare and consumer goods sectors maintain positive alphas during downturns, consistent with their resilience and inelastic demand. Overall, inside alphas track industry fundamentals and sectoral shocks rather than aggregate macroeconomic fluctuations, reinforcing their interpretation as characteristic-linked systematic mispricing.

figure[figure omitted — 1,577 chars of source]
figure[figure omitted — 1,581 chars of source]

Dynamics of Outside Alphas

We next examine the residual component $\widehat{\alpha}_O$, orthogonal to both characteristics and factors. Figures (ref) and (ref) plot representative firm-level and sector-averaged series. Unlike $\alpha_I$, the outside alphas exhibit no clear co-movement across firms or industries, suggesting that they primarily reflect idiosyncratic, transient deviations from fundamental value. This distinction between structured and residual mispricing provides new evidence on how inefficiencies manifest in the cross-section of returns.

figure[figure omitted — 523 chars of source]
figure[figure omitted — 516 chars of source]

Factor Loadings and Characteristic Relevance

Finally, we investigate the estimated $\widehat{\Gamma}$ matrix to assess which characteristics drive variation in factor exposures. We compute the Wald statistic \[ W_l = \widehat{\gamma}_l^\top \widehat{V}_{\gamma_l}^{-1} \widehat{\gamma}_l, \] which follows a $\chi^2(K)$ distribution under $H_0: \gamma_l = 0$. Table (ref) reports the results for $K = 1$$10$, with Bonferroni-adjusted critical values.

table[table omitted — 4,650 chars of source]

The number of statistically significant characteristics increases with $K$, as additional latent factors capture more structure in the cross-section. When $K = 10$, $30$ of $36$ characteristics significantly affect factor loadings. Variables such as book-to-market (LBM), Tobin’s $Q$, operating leverage (OL), market equity (LME), and capital turnover (CTO) consistently exhibit large test statistics, indicating that firm size, value, and operating efficiency are fundamental determinants of risk exposures. By contrast, investment (INV), leverage (LEV), and free cash flow (FCF) are generally insignificant.

Figure (ref) visualizes the estimated $\widehat{\Gamma}$ when $K = 5$. The first factor loads primarily on operating leverage and capital turnover, while the second is driven by cost ratios (SG&A-to-sales and fixed costs-to-sales), which together form a “cost” factor. The third factor contrasts market capitalization and book assets, resembling a value-like factor similar to the HML component in fama1993common and kelly2019characteristics. Later factors are less interpretable, reflecting more diffuse combinations of firm attributes.

figure[figure omitted — 464 chars of source]

Summary and Discussion

Taken together, our empirical findings confirm three key messages. First, both inside and outside alphas are statistically significant, highlighting that mispricing has distinct structured and idiosyncratic components. Second, inside alphas exhibit clear industry-level co-movement tied to fundamentals, while outside alphas capture transitory, firm-specific deviations. Third, characteristic-based factor loadings reveal economically interpretable dimensions of risk, including value, cost, and size components.

These results validate the inferential theory developed in Section (ref) and underscore the usefulness of our decomposition for understanding how firm fundamentals, latent factors, and residual mispricing jointly shape the cross-section of asset returns.

Concluding Remarks

This paper develops a unified econometric framework for modeling and inferring pricing errors in factor models that combine latent factors with firm characteristics. Our approach decomposes mispricing into two orthogonal components---inside alpha, which is systematically related to firm fundamentals but orthogonal to factor loadings, and outside alpha, which is orthogonal to both factors and characteristics. This decomposition reconciles the statistical efficiency of latent-factor approaches with the economic interpretability of characteristic-based models, thereby providing a coherent foundation for studying both systematic and idiosyncratic sources of mispricing.

Methodologically, we contribute a new class of low-rank estimators equipped with explicit debiasing and valid inferential theory. The resulting estimators admit closed-form expressions and Gaussian asymptotics even when the number of characteristics grows with the sample size, relaxing the restrictive conditions typically imposed in earlier work such as kelly2019characteristics and zhang2024testing. Our theoretical results establish the asymptotic normality of characteristic loadings, inside alphas, and outside alphas, allowing standard hypothesis tests on both factor structure and pricing errors. These inferential tools make it possible to distinguish between characteristic-driven and residual components of mispricing in a statistically rigorous way.

Empirically, applying the framework to U.S.\ equities from 2000--2019 reveals several new insights. Both inside and outside alphas are statistically significant, but they exhibit distinct economic patterns. Inside alphas display pronounced industry-level co-movement that aligns with persistent fundamentals such as technological change and sectoral shocks, while outside alphas behave as transient, firm-specific deviations that likely reflect liquidity frictions, behavioral biases, or short-term constraints. In addition, characteristic-based factor loadings highlight the importance of value, cost, and size dimensions in shaping cross-sectional risk exposures. Taken together, these results demonstrate that pricing errors in equity markets are structured, multi-layered phenomena rather than purely idiosyncratic residuals.

More broadly, our analysis bridges the gap between statistical and economic perspectives on asset pricing. By explicitly connecting latent factors to firm characteristics and by distinguishing between systematic and residual mispricing, the framework opens new avenues for understanding the sources and persistence of return anomalies. Future research could extend this setting to dynamic environments with time-varying characteristics, international markets, or alternative asset classes, as well as explore the interaction between inside and outside alphas in explaining cross-sectional risk premia. We hope that the theoretical tools and empirical evidence developed here will serve as a foundation for future studies at the intersection of econometrics, machine learning, and financial economics.