You will be provided with a reference and some statements. Please determine whether each statement is 'supported', 'unsupported', or 'unknown' with respect to the reference. Please note:
First, assess whether the reference contains any valid content. If the reference contains no valid information, such as a 'page not found' message, then all statements should be considered 'unknown'.
If the reference is valid, for a given statement: if the facts or data it contains can be found entirely or partially within the reference, it is considered 'supported' (data accepts rounding); if all facts and data in the statement cannot be found in the reference, it is considered 'unsupported'.

You should return the result in a JSON list format, where each item in the list contains the statement's index and the judgment result, for example:
[
    {
        "idx": 1,
        "result": "supported"
    },
    {
        "idx": 2,
        "result": "unsupported"
    }
]

Below are the reference and statements:
<reference>
Interpretable Deep Learning for Stock Returns:A Consensus-Bottleneck Asset Pricing Model



Report GitHub Issue

×

Title:

Content selection saved. Describe the issue below:

Description:

Submit without GitHub

Submit in GitHub

arXiv is now an independent nonprofit!

Learn more

×

Back to arXiv

Why HTML?

Report Issue

Back to Abstract

Download PDF

Abstract

1
Introduction

2
Model and Methodology

2.1
Return prediction model

2.2
Estimation

2.3
Model architecture

3
Data

3.1
Data description

3.2
Extracting macroeconomic state variables via autoencoder

3.3
Expanding window approach for model evaluation

4
Empirical Results

4.1
Cross-section of consensuses and stock returns

4.2
Portfolio-based pricing validation

4.2.1
Portfolio sorts on approximated consensuses and expected returns

4.2.2
Long-short portfolio performance

5
Dissecting Approximated Consensuses

5.1
Comparative regression analysis

5.2
Pricing error of test assets

6
Conclusion

References

A
Interpretable Artificial Intelligence

B
Data Preprocessing

B.1
Data lagging

B.2
Data Sampling

B.2.1
Firm screening

B.2.2
Variable selection

B.3
Data imputation

B.4
Data normalization

C
Implementing Neural Network for Asset Pricing

C.1
Activation functions

C.2
Stabilized learning

C.3
Model hyperparameters

D
Additional Results

D.1
Forecast horizon analysis

D.2
Properties of the joint optimization

D.3
Structure of the learned macroeconomic representations

D.4
Further robustness checks

D.4.1
Sensitivity of autoencoder performance to latent dimensionality

D.4.2
Comparison of state variables from principal components and autoencoder

D.4.3
Portfolio turnover and real-world implementability

D.5
Ablation studies

D.5.1
Effect of macroeconomic feature compression

D.5.2
Avoiding look-ahead of AI analysts

E
Detailed Data Description

E.1
Firm-level predictors

E.2
Macroeconomic predictors

E.3
Analysts’ consensus variables

License: CC Zero

arXiv:2512.16251v2 [q-fin.PR] 23 Dec 2025

This draft: December 23, 2025

First draft: March 5, 2025

Interpretable Deep Learning for Stock Returns:

A Consensus-Bottleneck Asset Pricing Model

We introduce the
Consensus-Bottleneck Asset Pricing Model
(CB-APM), a partially interpretable neural network that replicates the reasoning processes of sell-side analysts by capturing how dispersed investor beliefs are compressed into asset prices through a consensus formation process. By modeling this “bottleneck” to summarize firm- and macro-level information, CB-APM not only predicts future risk premiums of U.S. equities but also links belief aggregation to expected returns in a structurally interpretable manner. The model improves long-horizon return forecasts and outperforms standard deep learning approaches in both predictive accuracy and explanatory power. Comprehensive portfolio analyses show that CB-APM’s out-of-sample predictions translate into economically meaningful payoffs, with monotonic return differentials and stable long-short performance across regularization settings. Empirically, CB-APM leverages consensus as a regularizer to amplify long-horizon predictability and yields interpretable consensus-based components that clarify how information is priced in returns. Moreover, regression and Gibbons–Ross–Shanken (GRS)-based pricing diagnostics reveal that the learned consensus representations capture priced variation only partially spanned by traditional factor models, demonstrating that CB-APM uncovers belief-driven structure in expected returns beyond the canonical factor space. Overall, CB-APM provides an interpretable and empirically grounded framework for understanding belief-driven return dynamics.

Keywords:
Asset Pricing Model, Analysts’ Consensus, Neural Network, Interpretable Deep Learning, Cross-Section of Stock Returns

JEL Codes:
C45, C53, G0, G12, G17

1
Introduction

Empirical asset pricing has long relied on statistical modeling to explain stock returns, often within the framework of factor-based models such as those proposed by
Fama and French (1993)
;
Fama and French (2015)
and
Carhart (1997)
. These models aim to enhance explanatory power by identifying systematic risk factors that drive returns. However, despite decades of research, the ability of traditional models to predict future stock returns remains constrained, particularly in out-of-sample settings
(
Ang and Bekaert, 2007
;
Campbell and Thompson, 2008
;
Cochrane, 2008
)
. Moreover, it remains uncertain whether the results from the existing literature can be successfully reproduced and whether such predictors and econometric modeling methodologies can be generalized across a broader set of assets or diverse economic conditions. The proliferation of new factors—often referred to as the “factor zoo”
(
Cochrane, 2011
)
—has further complicated the landscape, raising concerns about robustness, data mining, and the true economic relevance of many proposed predictors.

To address these challenges, it is essential to explore deep inside the factor zoo to identify economically meaningful signals and evaluate their contribution to return prediction. Drawing from a number of studies on stock return predictors,
1
1

1

See
Welch and Goyal (2008)
,
Green et al. (2013)
,
Hou et al. (2015)
,
Harvey et al. (2016)
,
He et al. (2017)
,
Green et al. (2017)
,
Gu et al. (2020)
,
Feng et al. (2020)
,
Freyberger et al. (2020)
,
Bybee et al. (2023)
and
Jensen et al. (2023)
.
seminal work of
Gu et al. (2020)
proposes a “return prediction model” that integrates traditional asset pricing empirical frameworks and theories with the rapidly evolving field of machine learning. By utilizing a variety of machine learning algorithms including neural networks, and leveraging a high-dimensional set of predictive factors, their results significantly contribute to the literature by showing the effectiveness of nonlinear and complex modeling on empirical asset pricing. Several subsequent studies utilize the conceptual formulation of this study across diverse financial markets and assets such as bonds
(
Bianchi et al., 2021
)
, cryptocurrencies
(
Jaquart et al., 2021
;
Fang et al., 2024
)
and foreign stock exchanges
(
Leippold et al., 2022
)
. Theoretical studies have also emerged to justify the use of machine learning into empirical asset pricing. For instance,
Kelly et al. (2024)
illustrates how model complexity can be instrumental in achieving superior performance in cross-sectional return prediction, demonstrated through a simple example of penalized linear regression.

While return prediction models benefit from machine learning approaches due to their empirical flexibility, deep learning has also proven successful in approximating “asset pricing factor models”. Expanding on the research by
Kelly et al. (2019)
, which defines the covariance term
β
\beta
using the covariance of “characteristics”,
Feng et al. (2018)
and
Gu et al. (2021)
employ deep neural network architecture and the resulting latent factors to model the state variables of Intertemporal Capital Asset Pricing Model
(
Merton, 1973
, ICAPM, )
.
Chen et al. (2024)
introduce a novel architecture consisting of feedforward networks and LSTMs, that are trained via minimax optimization technique similar to that of Generative Adversarial Networks
(
Goodfellow et al., 2020
, GAN, )
. Based on arbitrage pricing theory (APT), the proposed model successfully approximates the stochastic discount factors (SDF) and corresponding risk loadings to formulate a highly predictive asset pricing model.

Despite the strong evidence that deep learning approaches illustrate evident potential in capturing the complex topology of predictor structures, critical limitation remains: Can the results from these models be considered trustworthy?
Rudin et al. (2022)
highlights such critical issue with machine learning black box models,

Black box models often predict the right answer for the wrong reason (the “Clever Hans” phenomenon), leading to excellent performance in training but poor performance in practice.

Recent studies in machine learning asset pricing frequently employ models that are not interpretable,
2
2

2

While the finance literature often uses the terms “explainability” and “interpretability” interchangeably, we maintain a strict conceptual distinction to ensure the reliability of our economic inferences. To keep the focus on the economic mechanisms of belief aggregation, we provide a detailed taxonomy of AI transparency in Internet Appendix
A
. This section clarifies why our framework prioritizes
interpretability-by-design
over the
post-hoc
explainability (XAI) common in recent black-box models. Such rigor is essential to verify that the CB-APM is capturing priced fundamental signals rather than suffering from the “Clever Hans” phenomenon, where models achieve high accuracy through spurious, non-economic correlations.
which raises concerns about relying on complex machine learning algorithms in empirical asset pricing without a clear understanding of why and how these models arrive at their conclusions. Furthermore, these papers often attempt to interpret the prediction results based on the learned models and derive economic implications. However,
Rudin (2019)
argues that such analyses are solely based on post-hoc explanations that should be considered as fitting narratives to the outcomes. These explanations are conveniently aligned with prevailing economic theories and tend to disregard contradictory evidence, which limits the scope and applicability of the findings. For these reasons,
Rudin (2019)
strives to rectify researchers and practitioners to use interpretable models over black-box algorithms. Despite growing interest in interpretable machine learning and trustworthy Artificial Intelligence (AI), a notable gap persists in applying and validating these approaches within asset pricing, beyond traditional regression or decision-tree models. In particular, existing machine-learning frameworks rarely achieve both strong predictive performance and economic interpretability. To address this gap, we propose the Consensus-Bottleneck Asset Pricing Model (CB-APM), a framework that employs a partially interpretable neural architecture to predict future stock returns while preserving clear economic structure.

Our approach builds upon two established pillars of financial economics, the rational expectations hypothesis and empirically documented relationships between analyst consensus information and asset prices. Rational expectations, proposed by
Muth (1961)
, posit that market participants form forecasts using all available historical information. Several research find evidence of such a hypothesis from the decisions of sell-side analysts, deriving the economic implications of analysts’ opinions and estimates. Subsequent empirical work demonstrates this principle in sell-side analysts’ behavior.
Lovell (1986)
shows economic agents systematically incorporate public information into earnings forecasts, while
Lim (2001)
establishes predictable patterns in analysts’ forecast revisions consistent with Bayesian updating. Crucially,
Jegadeesh et al. (2004)
identify specific style factors—including momentum, growth prospects, and trading volume—that systematically influence analysts’ stock recommendations, suggesting a quantifiable link between firm characteristics and consensus formation.
Barber et al. (2001)
further demonstrates the economic significance of consensus recommendations, showing that strategies based on the most and least favorable recommendations yield significant abnormal gross returns.

However, the efficacy of relying solely on these aggregated consensus measures is nuanced, as their predictive value is critically moderated by the underlying heterogeneity of beliefs and inherent institutional biases. For instance,
Palley et al. (2025)
demonstrate that the informativeness of consensus target prices depends on dispersion: while low dispersion yields positive return predictability, high dispersion—driven by incentive-driven staleness—results in a robust negative correlation. This behavioral contamination is further documented by studies using machine learning to construct unbiased benchmarks for expectations.
Van Binsbergen et al. (2023)
find that analysts’ conditional expectations are, on average, upwardly biased, which correlates with negative cross-sectional return predictability. While early evidence suggested that “AI analysts” could exploit these biases, recent critiques by
Zhang et al. (2025)
suggest that such outperformance may be sensitive to look-ahead biases, challenging the notion that black-box machine learning is a panacea for earnings forecasting. Furthermore,
Cao et al. (2024)
find that the perceived superiority of AI analysts stems primarily from the absence of directional human biases rather than superior information processing. This empirical complexity underscores the necessity of a framework like the CB-APM, which is specifically designed to disentangle the priced information from the behavioral noise that accompanies analyst belief aggregation.

Despite these complexities, the hypothesis that analyst consensus remains a critical mediator for future returns finds robust support. Recent evidence suggests a synergy between human and machine intelligence;
Cao et al. (2024)
show that combining AI’s computational power with the human capacity to synthesize “soft” institutional information yields the most accurate forecasts. This implies that analysts remain vital intermediaries whose inputs provide incremental value beyond what is captured by raw firm characteristics. Historically,
Diether et al. (2002)
documented that forecast dispersion affects risk premiums, while
Sorescu and Subrahmanyam (2006)
established pronounced price reactions to revisions in these estimates. More recently,
Van Binsbergen et al. (2023)
demonstrate that when machine learning is used to successfully isolate forecast biases, these signals are not only predictive of stock returns but also of corporate financing decisions, such as equity issuances. Taken together, these findings validate consensus information as a measurable economic construct that bridges the gap between high-dimensional firm characteristics and expected returns.

The CB-APM framework operationalizes these insights through a concept-bottleneck architecture inspired by
Koh et al. (2020)
, directly into the return prediction model. This architecture serves as a structural filter that disciplines the “factor zoo”, ensuring the model only utilizes characteristics that are salient enough to influence the expectations of market participants. By anchoring the latent states to observable analyst consensus, we effectively prevent the model from exploiting spurious correlations that lack a documented foundation in human belief formation. Building on the necessity to separate signal from noise, CB-APM is designed to recover the priced component of these expectations by explicitly filtering out the behavioral biases inherent in their aggregation. Its nonlinear “consensus formation” stage synthesizes firm characteristics and macroeconomic states into consensus-like latent expectations, reflecting the documented process through which analysts aggregate information. A subsequent linear “pricing” stage translates these learned expectations into expected returns, preserving interpretability through transparent economic loadings. By routing all predictive content through these latent expectations, the framework imposes an inherent information constraint that limits reliance on spurious high-dimensional patterns and anchors inference to economically interpretable drivers. In unifying rational-expectations principles with empirical evidence on analyst behavior, CB-APM achieves dual objectives: it delivers strong cross-sectional predictive accuracy while offering a tractable representation of how expectations are formed and translated into risk premiums.

Our contributions are threefold. First, we introduce a concept-bottleneck framework that synthesizes the high-dimensional predictor set into interpretable, consensus-style expectations, providing a structured economic link between characteristics, analysts’ beliefs, and expected returns. Second, we demonstrate that this architecture delivers economically large improvements in long-horizon return prediction across expanding-window evaluations. Third, we show that the learned consensus representations encode priced information that is only partially spanned by traditional factor models, offering new empirical insight into how belief heterogeneity and information aggregation shape risk premia. These contributions advance recent efforts to integrate interpretable machine learning with the core principles of empirical asset pricing.

To empirically validate the effectiveness of CB-APM, we assess its predictive performance and economic implications using a comprehensive dataset spanning from January 1994 to December 2023, consisting of 605,722 firm-month observations across 4,683 U.S. companies. The dataset integrates 114 firm-level predictors, 123 macroeconomic indicators, and 9 analysts’ consensus variables including EPS forecast revisions and forecast dispersions. To account for the time dynamics of return prediction, we employ an expanding window approach, where the training dataset grows over time while keeping validation and test sets fixed. This experimental setup allows us to assess the robustness of CB-APM under evolving market conditions.

Our empirical analysis demonstrates that CB-APM delivers substantial improvements in both predictive performance and economic interpretability. First, in the cross-section of consensus and stock returns, incorporating consensus learning markedly enhances long-horizon return forecasts: CB-APM attains an out-of-sample
R
2
R^{2}
of 10.46% for annual returns, representing a significant improvement over a standard deep learning benchmark (
R
2
=
7.63
%
R^{2}=7.63\%
), while simultaneously achieving an average
R
2
R^{2}
of 24.21% in approximating analyst consensus variables. These gains remain robust across expanding-window evaluations, indicating stable performance across different market regimes.

Second, portfolio-level analyses establish the model’s economic relevance. Portfolios formed on out-of-sample CB-APM predictions display strongly monotonic payoff structures, with high-minus-low spreads approaching 2.3% per month for regularized specifications (
λ
≥
0.3
\lambda\geq 0.3
). The double sorts on model-implied returns and analysts’ earnings forecasts further reveal that the model internalizes both the informational and behavioral components embedded in analyst expectations. In particular, the expected-return spreads are largest in states characterized by analyst pessimism—low analysts’ earnings forecasts levels—where expectation errors and mispricing are most pronounced, and they progressively shrink as analyst optimism increases. This state-dependent attenuation indicates that the CB-APM distills the priced component of forecasted earnings while appropriately adjusting for optimism-driven noise in analysts’ beliefs.

Finally, long-short portfolios derived from the model’s forecasts achieve economically significant and stable out-of-sample performance, with mean monthly log returns rising from 1.53% at
λ
=
0
\lambda=0
to 2.20% at
λ
=
0.3
\lambda=0.3
and the annualized Sharpe ratio improving from 1.10 to 1.44. These results establish a direct correspondence between predictive accuracy, cross-sectional return ordering, and risk-adjusted profitability, confirming that consensus regularization enhances not only statistical fit but also economic value.

Beyond predictive performance, we further examine whether the consensus-bottleneck captures economically meaningful pricing structure. A comparative regression analysis demonstrates that the CB-APM–implied consensuses deliver substantially stronger explanatory power for annual returns than raw analyst signals: pooled OLS regressions exhibit an order-of-magnitude improvement in adjusted
R
2
R^{2}
, together with economically interpretable shifts in coefficient signs and magnitudes. These gains arise because the consensus layer synthesizes information from firm characteristics and macroeconomic conditions into belief-like representations that are simultaneously close to observable analyst forecasts and tightly aligned with priced return variation. Variables that the model reconstructs with higher fidelity display more stable and economically intuitive return sensitivities, whereas poorly reconstructed dimensions exhibit weaker economic content or sign reversals—highlighting that economic interpretability depends jointly on approximation quality and return-pricing relevance. This evidence confirms that the consensus-bottleneck does not merely denoise analyst inputs but reorganizes information into latent expectations that better capture the priced component of belief dispersion.

We further evaluate the pricing relevance of these signals using Gibbons–Ross–Shanken (GRS) tests on benchmark portfolios and portfolios formed on model-implied returns and individual consensus dimensions. Consensus-based long–short factors span meaningful components of systematic return variation but do not fully replicate the benchmark factor structure, indicating that the learned expectations are economically relevant without collapsing onto the canonical dimensions of market, size, value, momentum, profitability, or investment. Conversely, traditional factor models increasingly fail to price portfolios formed on CB-APM’s predicted returns as the consensus-bottleneck tightens, suggesting that the model uncovers structured forms of nonlinear or interaction-based return heterogeneity that lie outside the linear span of standard factors. Portfolios sorted on individual consensus dimensions produce modest pricing errors, consistent with the view that belief-based signals reflect compressible yet economically meaningful combinations of characteristics. Taken together, these findings show that CB-APM extracts interpretable consensus representations that contain priced information only partially captured by existing factor models, positioning the framework as a complementary approach that links analysts’ heterogeneous beliefs to expected returns in a transparent and theoretically coherent manner.

Collectively, these results establish CB-APM as a novel and effective framework that integrates interpretable deep learning with foundational principles of financial economics. Unlike prior machine learning approaches that prioritize accuracy at the expense of transparency, CB-APM demonstrates that interpretable architectures can preserve theoretical grounding while achieving state-of-the-art empirical performance. By jointly modeling analysts’ expectations and stock returns, our framework provides a principled means of disentangling forward-looking information embedded in firm characteristics and macroeconomic variables, yielding insights into how such information is aggregated and priced. This dual capacity—enhancing return predictability while maintaining an economically interpretable structure—constitutes the central contribution of our paper and advances the emerging literature on interpretable machine learning in finance.

The remainder of the paper is organized as follows. Section
2
outlines the model, estimation procedure, and architecture. Section
3
describes the data and evaluation design, including the autoencoder for macroeconomic state extraction. Section
4
presents empirical results on predictive performance, macroeconomic embeddings, and portfolio-based pricing implications. Section
5
investigates the pricing content of the approximated consensuses using regression and GRS tests. Section
6
concludes. The Internet Appendix provides additional robustness analyses and supplementary results.

2
Model and Methodology

2.1
Return prediction model

Similar to the
Gu et al. (2020)
, the asset return prediction error model utilized in our work is formulated for the
h
h
-horizon forecasting problem as below,

R
i
,
t
+
h
=
𝔼
t
​
[
R
i
,
t
+
h
]
+
ε
i
,
t
+
h
,
R_{i,t+h}=\mathbb{E}_{t}\!\left[R_{i,t+h}\right]+\varepsilon_{i,t+h},

(1)

where
R
i
,
t
+
h
R_{i,t+h}
is the
h
h
-month return of asset
i
i
excess of the risk-free rate at time
t
+
h
t+h
, and
ε
i
,
t
+
h
\varepsilon_{i,t+h}
is an error term. In this context,
h
h
is used to assign the forecasting horizon, enabling the consideration of multi-horizon predictions, allowing CB-APM to model long-term dependencies. The expected excess return in equation (
1
) is defined as the expectation conditional on information sets,

𝔼
t
[
R
i
,
t
+
h
]
=
𝔼
[
R
i
,
t
+
h
∣
ℐ
i
,
t
f
,
ℐ
t
m
]
.
\mathbb{E}_{t}\!\left[R_{i,t+h}\right]=\mathbb{E}\!\left[R_{i,t+h}\mid\mathcal{I}^{f}_{i,t},\,\mathcal{I}^{m}_{t}\right].

Here,
ℐ
i
,
t
f
\mathcal{I}^{f}_{i,t}
and
ℐ
t
m
\mathcal{I}^{m}_{t}
are the sets of firm-specific characteristics and macroeconomic predictors at time
t
t
, respectively.
3
3

3

A detailed description of the predictors comprising the information sets is provided in Section
3
, and a complete list of variables is available in Internet Appendix
E
. The macroeconomic information set
ℐ
t
m
\mathcal{I}^{m}_{t}
is represented empirically by a latent vector extracted through an autoencoder trained on macroeconomic variables, as described in Section
3.2
.
It is important to note that consensus information is deliberately excluded from
ℐ
i
,
t
f
\mathcal{I}^{f}_{i,t}
.

This framework is further developed by defining the functional form of the conditional expectation as a composite function,

𝔼
[
R
i
,
t
+
h
∣
ℐ
i
,
t
f
,
ℐ
t
m
]
=
g
(
f
(
ℐ
i
,
t
f
,
ℐ
t
m
;
ϕ
)
;
θ
)
,
\mathbb{E}\!\left[R_{i,t+h}\mid\mathcal{I}^{f}_{i,t},\,\mathcal{I}^{m}_{t}\right]=g\!\left(f\!\left(\mathcal{I}^{f}_{i,t},\,\mathcal{I}^{m}_{t};\phi\right);\theta\right),

(2)

where the function
f
⁡
(
⋅
)
f(\cdot)
and
g
⁡
(
⋅
)
g(\cdot)
are smooth functions parameterized by learnable parameters
θ
\theta
and
ϕ
\phi
. The function
f
⁡
(
⋅
)
f(\cdot)
is specifically designed to model the conditional expectation of analyst consensus. Then, the function
g
⁡
(
⋅
)
g(\cdot)
models the expected return only using the features of approximated consensus from the previous step, creating a “concept-bottleneck” within the prediction model. This empirical design is predicated on the understanding that both researchers in empirical asset pricing and financial analysts share the objective of assessing a firm’s value and discerning the factors that influence these valuations. While analysts often have access to broader datasets, including some predictive signals that may not be publicly available or included in this article, asset pricing panel data can represent a information subset in a decent quality by providing a comprehensive and quantifiable measures of firm’s fundamentals and macroeconomic conditions that are crucial for the approximation of the consensus, as shown in the empirical results later on.

In mathematical form,
f
⁡
(
⋅
)
f(\cdot)
approximates the analyst consensus variables, denoted as
C
i
,
t
C_{i,t}
.

C
i
,
t
=
f
⁡
(
ℐ
i
,
t
f
,
ℐ
t
m
,
ϕ
)
.
C_{i,t}=f\!\left(\mathcal{I}^{f}_{i,t},\,\mathcal{I}^{m}_{t}\,;\,\phi\right).

Let the approximated value of
C
i
,
t
C_{i,t}
and the parameter
ϕ
\phi
be
C
^
i
,
t
\hat{C}_{i,t}
and
ϕ
^
\hat{\phi}
, respectively, then,

C
^
i
,
t
=
f
⁡
(
ℐ
i
,
t
f
,
ℐ
t
m
,
ϕ
^
)
.
\hat{C}_{i,t}=f\!\left(\mathcal{I}^{f}_{i,t},\,\mathcal{I}^{m}_{t}\,;\,\hat{\phi}\right).

Finally, the expected excess return is defined with function
g
⁡
(
⋅
)
g(\cdot)
and the approximated
C
^
i
,
t
\hat{C}_{i,t}
as below,

𝔼
t
​
[
R
i
,
t
+
h
]
=
g
⁡
(
C
^
i
,
t
,
θ
)
.
\mathbb{E}_{t}\!\left[R_{i,t+h}\right]=g\!\left(\hat{C}_{i,t}\,;\,\theta\right).

(3)

As discussed in
Daniel and Titman (1997)
, the main limitation of the return prediction error model is the absence of economic constraints. For instance, the fundamental theorem of asset pricing constrains the arbitrage opportunity, which implies that the difference between the price of an identical asset is improbable. This condition is referred to as “the law of one price” in asset pricing theory. In the cases without such condition, two different assets can have identical price despite disparate fundamental values.

However, CB-APM diverges from the approach of return prediction modeling for several reasons. Firstly, it offers greater flexibility, accommodating diverse scenarios involving analyst estimates and future returns. Unlike factor models, which do not differentiate prices of identical risk factors, CB-APM acknowledges that similar analyst opinions across firms may yield distinct future returns. While we assume rational decision-making by analysts, as discussed in subsequent sections, it is prudent not to constrain such scenarios initially. Secondly, CB-APM facilitates a range of optimization approaches in approximating the asset pricing model. Unlike factor models, where the estimation process is mostly the extension of Fama–MacBeth regression
(
Fama and MacBeth, 1973
)
restricting the integration of the entire expected return modeling process, CB-APM allows for a more holistic training process, avoiding multiple optimization procedures. Overall, given that the consensus-bottleneck represents a novel approach in asset pricing research, we aimed to maintain the underlying framework as simple and flexible as possible.

Although it is designed as intended, given that neural networks are well-recognized as “universal approximators”, the model can allow any scenarios and consequences as outcomes, that don’t necessarily align with the economic theories. To overcome the limitation of the proposed prediction model, we apply stabilized optimization approaches proposed in the machine learning literature, such as regularization and scheduling. Such techniques are expected to function as “universal constraints”, achieving both practical performances and theoretical rigor. See Internet Appendix
C.3
for detailed discussions and experimental settings.

2.2
Estimation

In this section, we provide the loss function of the model that simultaneously estimates the parameters of function
f
⁡
(
⋅
)
f(\cdot)
and
g
⁡
(
⋅
)
g(\cdot)
from equation (
2
) in a single optimization step.

Given
λ
>
0
\lambda>0
, the model’s loss function is structured as a joint optimization task, represented by a weighted sum of two distinct loss functions:

ℒ
=
ℒ
R
+
λ
⁡
⟨
1
,
ℒ
C
⟩
,
\mathcal{L}=\mathcal{L}_{R}+\lambda\,\langle\textbf{1},\mathcal{L}_{C}\rangle,

(4)

where the “return loss”
ℒ
R
\mathcal{L}_{R}
is formulated as,

ℒ
R
​
(
ϕ
,
θ
)
=
1
N
​
T
​
∑
i
=
1
N
∑
t
=
1
T
(
R
i
,
t
+
h
−
g
⁡
(
f
⁡
(
ℐ
i
,
t
f
,
ℐ
t
m
,
ϕ
)
,
θ
)
)
2
,
\mathcal{L}_{R}(\phi,\theta)=\frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}\Bigl(R_{i,t+h}-g\!\left(f\!\left(\mathcal{I}^{f}_{i,t},\,\mathcal{I}^{m}_{t}\,;\,\phi\right)\,;\,\theta\right)\Bigr)^{2},

(5)

and the “consensus loss”
ℒ
C
\mathcal{L}_{C}
is formulated as,

ℒ
C
​
(
ϕ
)
=
1
N
​
T
​
∑
i
=
1
N
∑
t
=
1
T
(
C
i
,
t
−
f
⁡
(
ℐ
i
,
t
f
,
ℐ
t
m
,
ϕ
)
)
2
.
\mathcal{L}_{C}(\phi)=\frac{1}{NT}\sum_{i=1}^{N}\sum_{t=1}^{T}\Bigl(C_{i,t}-f\!\left(\mathcal{I}^{f}_{i,t},\,\mathcal{I}^{m}_{t}\,;\,\phi\right)\Bigr)^{2}.

(6)

ℒ
R
\mathcal{L}_{R}
and
ℒ
C
\mathcal{L}_{C}
are cross-sectional mean squared errors (MSE) from a standard pooled OLS estimator, where
λ
\lambda
is a hyperparameter that assigns weight to the consensus loss, providing additional flexibility into the empirical design of the model.

We also estimate a benchmark model taking
λ
=
0
\lambda=0
from equation (
4
), which ignores learning analyst opinions by removing the consensus loss term, making the model identical to the naïve return prediction model. Although
f
⁡
(
⋅
)
f(\cdot)
and
g
⁡
(
⋅
)
g(\cdot)
are the model defined with a separate set of learnable parameters
θ
\theta
and
ϕ
\phi
, they can be considered as a single neural network when
λ
=
0
\lambda=0
since the optimization procedures for each networks are not independent.

The strategy of jointly learning the consensus and excess return offers several advantages. Firstly, it tends to yield higher performance metrics due to the synergistic learning of interconnected variables. An alternative method might involve independent optimization, where the
f
⁡
(
⋅
)
f(\cdot)
and
g
⁡
(
⋅
)
g(\cdot)
are trained independently in two separate steps. However, this segmented approach often fails to capture the potential inter-dependencies between the consensus estimates and the resulting excess returns. Furthermore, since the information set
ℐ
i
,
t
f
\mathcal{I}^{f}_{i,t}
and
ℐ
t
m
\mathcal{I}^{m}_{t}
are not included in equation (
5
), the training of
g
⁡
(
⋅
)
g(\cdot)
entirely depends on the quality of the extracted signals in approximated consensus, which makes training with the loss function
ℒ
R
\mathcal{L}_{R}
extremely challenging.

Secondly, it provides deeper insights and more intuitive understanding of the underlying financial dynamics. Independently learning
f
⁡
(
⋅
)
f(\cdot)
using the equation (
6
) is not a novel concept and aligns with the existing literature supporting the evidence of the rational expectations hypothesis. As discussed in previous sections, the set of predictor signals used in this study is regarded to contain a significant amount of information sufficient to make ‘‘rational” expectations,
4
4

4

Appendix
D.5.2
formally evaluates this claim by examining the consensus-only specification corresponding to
λ
→
∞
\lambda\to\infty
in Equation (
4
), showing that the model learns analysts’ consensus variables remarkably well (out-of-sample
R
2
=
30.30
%
R^{2}=30.30\%
) even without any return-prediction objective. This validates the architectural design of the consensus-bottleneck and provides empirical support for the rational expectations interpretation underlying the model.
which simplifies the problem of approximating the opinions of individuals, compared to predicting future returns of assets. However, since analysts perform their analysis as their job, they must think and act beyond being merely “rational”; they must be “professional”. Therefore, we posit that professional and successful analysts strive to make their estimates predict not only “macroeconomic” consequences but also “firm-specific” outcomes. More specifically, proficient analysts will make decisions that better predict the future returns of a firm’s stocks.

2.3
Model architecture

In this section, we provide detailed explanations the model architecture. The overall framework of CB-APM is described in Figure
1
. The model consists of two main components; the consensus module and the prediction module. Each of these modules corresponds to the function
f
⁡
(
⋅
)
f(\cdot)
and
g
⁡
(
⋅
)
g(\cdot)
in equation (
2
).

[Insert Figure
1
here]

In the proposed model, the consensus module is designed as an arbitrary feedforward network, while the prediction module is restricted to a simple linear regression that receives consensus variables as inputs and yields the expected excess return. This design choice is critical for enhancing interpretability in particular. When both modules are complex feedforward networks with multiple hidden layers, the advantage of using a consensus-based approach diminishes since it creates two separate black-box models from a single black-box model.

The loss functions, as defined in equations (
5
) and (
6
), are computed using the outputs from the respective modules. Once we get the return loss from the return module, the final loss function is calculated via weighted sum of these two loss functions as described in equation (
4
). The backpropagation in the CB-APM is conducted in a single step, utilizing the composite loss function in equation (
4
), which simultaneously adjusts the weights in both the consensus and prediction modules.

For an activation function, we utilize Gaussian Error Linear Units function (GELU) as non-linearity of the neural network. The mathematical formulation of GELU is given as below.

GELU
⁡
(
x
)
=
x
⋅
ℙ
⁡
(
X
≤
x
)
=
x
⋅
Φ
⁡
(
x
)
,
\mathrm{GELU}(x)=x\cdot\mathbb{P}(X\leq x)=x\cdot\Phi(x),

where
Φ
⁡
(
x
)
\Phi(x)
the cumulative distribution function for Gaussian distribution
X
∼
N
⁡
(
0
,
σ
2
)
X\sim N(0,\sigma^{2})
. GELU was first introduced by
Hendrycks and Gimpel (2016)
as an alternative of Rectified linear units (ReLU)
(
Nair and Hinton, 2010
)
. Figure
2
shows that GELU permits some small interval for negative inputs to propagate through subsequent layers. See Internet Appendix
C
for a review of the literature on activation functions and justification for the selection of GELU.

[Insert Figure
2
here]

The mathematical form of the model architecture is given as follows. First, let
X
X
denote the input layer, and
H
(
1
)
,
H
(
2
)
,
…
,
H
(
n
)
H^{(1)},H^{(2)},\ldots,H^{(n)}
represent the hidden layers. The weight matrices connecting the layers are denoted as
W
0
,
W
1
,
…
,
W
n
c
,
W
n
r
W_{0},W_{1},\ldots,W_{n}^{c},W_{n}^{r}
, where
W
0
W_{0}
connects the input layer to the first hidden layer,
W
1
W_{1}
connects the first hidden layer to the second hidden layer, and so forth, up to
W
n
c
W_{n}^{c}
connecting the
n
n
-th hidden layer to the output layer of the consensus module, and
W
n
r
W_{n}^{r}
connecting the output layer of the consensus module to the output layer of the return module. Similarly, the bias vectors are represented as
b
0
,
b
1
,
…
,
b
n
c
,
b
n
r
b_{0},b_{1},\ldots,b_{n}^{c},b_{n}^{r}
. The computations for the hidden layers are as follows,

H
(
1
)
\displaystyle H^{(1)}

=
GELU
⁡
(
W
0
​
(
ℐ
f
⊕
ℐ
m
)
+
b
0
)
,
\displaystyle=\mathrm{GELU}\!\left(W_{0}\!\left(\mathcal{I}^{f}\oplus\mathcal{I}^{m}\right)+b_{0}\right),

H
(
2
)
\displaystyle H^{(2)}

=
GELU
⁡
(
W
1
​
H
(
1
)
+
b
1
)
,
\displaystyle=\mathrm{GELU}\!\left(W_{1}H^{(1)}+b_{1}\right),

⋮
\displaystyle\vdots

H
(
n
)
\displaystyle H^{(n)}

=
GELU
⁡
(
W
n
−
1
​
H
(
n
−
1
)
+
b
n
−
1
)
.
\displaystyle=\mathrm{GELU}\!\left(W_{n-1}H^{(n-1)}+b_{\,n-1}\right).

Then the output layer computation of the consensus module is given by,

f
⁡
(
ℐ
f
,
ℐ
m
,
ϕ
)
=
W
n
c
​
H
(
n
)
+
b
n
c
.
f\!\left(\mathcal{I}^{f},\,\mathcal{I}^{m}\,;\,\phi\right)=W^{c}_{n}\,H^{(n)}+b^{c}_{n}.

and the output layer computation of the return module is given by,

g
⁡
(
f
⁡
(
ℐ
f
,
ℐ
m
,
ϕ
)
,
θ
)
=
W
n
r
​
f
​
(
ℐ
f
,
ℐ
m
,
ϕ
)
+
b
n
r
.
g\!\left(f\!\left(\mathcal{I}^{f},\,\mathcal{I}^{m}\,;\,\phi\right)\,;\,\theta\right)=W^{r}_{n}\,f\!\left(\mathcal{I}^{f},\,\mathcal{I}^{m}\,;\,\phi\right)+b^{r}_{n}.

Note that there are no activation layers between the consensus and return modules for interpretability. Therefore, when
λ
=
0
\lambda=0
, the CB-APM functions as a simple feedforward network, with the number of hidden layers matching that of the consensus module. The learnable weights of CB-APM are initialized by adopting the He initialization proposed by
He et al. (2015)
.

3
Data

3.1
Data description

In this section, we provide the brief explanations on the dataset and the sampling splitting scheme employed for empirical studies. The dataset comes from four distinct sources, which are all publicly available at the moment. Firstly, we obtain open-source asset pricing panel data from
Chen and Zimmermann (2022)
, available to download on their website (
https://www.openassetpricing.com/
).
5
5

5

Data from
Chen and Zimmermann (2022)
undergoes several preprocessing steps including lagging, data sampling, data imputation, and rank normalization, as detailed in Internet Appendix
B
.
It comprises 114 firm-level predictors consisting of diverse financial metrics such as accounting figures, 13F filings, trading activities, and derivatives data.

Chen and Zimmermann (2022)
also features 9 analysts’ consensus variables including EPS forecast revision (
AnalystRevision
), Change in recommendation (
ChangeInRecommendation
), Change in Forecast and Accrual (
ChForecastAccrual
), Long-vs-short EPS forecasts (
EarningsForecastDisparity
), Analyst earnings per share (
FEPS
), EPS Forecast Dispersion (
ForecastDispersion
), Earnings forecast revisions (
REV6
), Analyst Value (
AnalystValue
), and Analyst Optimism (
AOP
).

Secondly, stock prices and firm sizes data are sourced from CRSP (Center for Research in Security Prices)
6
6

6

Accessible via WRDS (Wharton Research Data Services).
, companies listed on the NYSE, Amex, and Nasdaq. This dataset is synchronized with the firm list from the panel data provided by
Chen and Zimmermann (2022)
.

Lastly, the macroeconomic variables are obtained from FRED-MD database
(
McCracken and Ng, 2016
)
and
Welch and Goyal (2008)
. FRED-MD consists of 115 monthly predictors that includes macroeconomic indicators reflecting the U.S. labor markets, consumption rates, monetary policies, etc. An additional set of 8 macroeconomic variables is constructed from the database maintained by
Welch and Goyal (2008)
on Goyal’s website (
https://sites.google.com/view/agoyal145
), following
Gu et al. (2020)
. T-bill rate is also obtained from this dataset, which is used for calculating risk premiums.

The final merged dataset consists of samples spanning from January 1994 to December 2023, with total 605,722 samples from 4,683 U.S. companies. Detailed descriptions of the dataset components and their respective sources are provided in Internet Appendix
E
.

3.2
Extracting macroeconomic state variables via autoencoder

To incorporate macroeconomic dynamics into the conditional expectation function
𝔼
t
​
[
R
i
,
t
+
h
]
\mathbb{E}_{t}[R_{i,t+h}]
defined in equation (
2
), we encode the aggregate information set
ℐ
t
m
\mathcal{I}_{t}^{m}
through an autoencoder-based representation. While macroeconomic variables are often dismissed in cross-sectional asset pricing due to their perceived homogeneity across firms, we argue that macro context exerts differentiated influence through sectoral dynamics, capital structure sensitivity, and behavioral channeling. However, the sheer volume and redundancy of macroeconomic indicators, particularly those sourced from databases such as FRED-MD, pose significant challenges for model training. Including hundreds of highly correlated variables not only increases the risk of overfitting but also dilutes the learning signal by overwhelming the model with noise and irrelevant information.

Moreover, many macroeconomic series track similar phenomena at varying lags, granularities, or levels of transformation (e.g., growth rates, differences, log-levels), creating unnecessary dimensionality without proportional gains in explanatory power. This redundancy hinders both the stability and interpretability of predictive models, especially those trained on firm-level data where macro variables are shared across the entire cross-section. Reducing this high-dimensional input into a compact, informative representation is thus not only computationally efficient but also essential for isolating the latent economic regimes that meaningfully affect asset returns.

Dimensionality reduction techniques have long been used in financial modeling to address such issues. Principal Component Analysis (PCA) has served as a standard tool for extracting latent factors from large panels of macroeconomic variables
(
Ludvigson and Ng, 2007
)
, while extensions such as Sparse PCA and Independent Component Analysis (ICA) have been applied to improve factor interpretability and reduce multicollinearity
(
Fan et al., 2016
;
Erichson et al., 2020
)
. More recently, deep learning approaches—particularly autoencoders—have gained traction in the asset pricing literature for their ability to capture nonlinear interactions and extract economically meaningful latent structures from noisy, high-dimensional data
(
Chen et al., 2024
;
Gu et al., 2021
)
. These methods have proven effective in modeling complex macro-financial dynamics that traditional linear techniques may fail to uncover.
7
7

7

Appendix
D.4.2
demonstrates that replacing the autoencoder with a 32-factor PCA markedly weakens out-of-sample return predictability, despite both approaches delivering similar accuracy in reconstructing analysts’ consensus. This divergence highlights the advantage of nonlinear compression in capturing macroeconomic structure relevant for pricing.

To address these challenges, we enhance model performance by encoding the macroeconomic regime using an autoencoder, thereby providing structured, compact, and economically interpretable representations that condition firm-level predictions. Instead of feeding all 115 raw macroeconomic variables from the FRED-MD database directly into the model, we train an autoencoder to learn a lower-dimensional latent representation of the macroeconomic environment at each time step. Figure
3
illustrates this process, where the encoder compresses high-dimensional macroeconomic inputs into a latent macroeconomic state vector
𝐳
t
\mathbf{z}_{t}
, which is subsequently concatenated with firm-level features and passed into the CB-APM architecture. During training, the decoder reconstructs the input variables, and the network is optimized to minimize the mean squared reconstruction error. After training, only the encoder is retained to generate macroeconomic embeddings for prediction.

Formally, let the macroeconomic input at time
t
t
be
𝐱
t
∈
ℝ
D
\mathbf{x}_{t}\in\mathbb{R}^{D}
, where
D
=
123
D=123
. The encoder
ℰ
ϕ
​
(
⋅
)
\mathcal{E}_{\phi}(\cdot)
maps this input to a latent representation
𝐳
t
∈
ℝ
d
\mathbf{z}_{t}\in\mathbb{R}^{d}
:
8
8

8

Empirically, setting the latent dimension to
d
=
32
d=32
yields the best out-of-sample performance (see Internet Appendix
D.4.1
).

𝐳
t
=
ℰ
ϕ
​
(
𝐱
t
)
.
\mathbf{z}_{t}=\mathcal{E}_{\phi}\!\left(\mathbf{x}_{t}\right).

The decoder
𝒟
θ
​
(
⋅
)
\mathcal{D}_{\theta}(\cdot)
reconstructs the input:

𝐱
^
t
=
𝒟
θ
​
(
𝐳
t
)
,
\hat{\mathbf{x}}_{t}=\mathcal{D}_{\theta}\!\left(\mathbf{z}_{t}\right),

and the model is trained to minimize the reconstruction loss:

ℒ
AE
​
(
θ
,
ϕ
)
=
1
T
​
∑
t
=
1
T
‖
𝐱
t
−
𝒟
θ
​
(
ℰ
ϕ
​
(
𝐱
t
)
)
‖
2
2
.
\mathcal{L}_{\text{AE}}(\theta,\phi)=\frac{1}{T}\sum_{t=1}^{T}\left\|\mathbf{x}_{t}-\mathcal{D}_{\theta}\!\bigl(\mathcal{E}_{\phi}(\mathbf{x}_{t})\bigr)\right\|_{2}^{2}.

After training, the encoder output
𝐳
t
\mathbf{z}_{t}
is concatenated with each firm’s feature vector
𝐱
i
,
t
firm
\mathbf{x}_{i,t}^{\text{firm}}
to form the model input:

𝐱
i
,
t
input
=
[
𝐱
i
,
t
firm
;
𝐳
t
]
.
\mathbf{x}_{i,t}^{\text{input}}=\left[\mathbf{x}_{i,t}^{\text{firm}}\,;\,\mathbf{z}_{t}\right].

Formally, the latent representation
𝐳
t
\mathbf{z}_{t}
learned through the autoencoder serves as an empirical proxy for the macroeconomic information set
ℐ
t
m
\mathcal{I}_{t}^{m}
introduced in equation (
2
). In this context,
𝐳
t
\mathbf{z}_{t}
functions as a compressed, data-driven approximation of the macroeconomic state observable to investors at time
t
t
. This design allows the CB-APM to integrate the high-dimensional macroeconomic information set into a tractable latent representation, ensuring that firm-level forecasts remain conditioned on a parsimonious yet informative depiction of the aggregate economic environment. By mapping
ℐ
t
m
\mathcal{I}_{t}^{m}
into
𝐳
t
\mathbf{z}_{t}
, the model effectively operationalizes the theoretical information set within a learnable structure, thereby linking the empirical implementation of the macro encoder to the conditional expectation framework defined in equation (
2
).

As illustrated in Figure
3
, this framework visualizes the overall data pipeline of the CB-APM, depicting how macroeconomic inputs are encoded, compressed, and subsequently integrated with firm-level characteristics for return prediction. The figure serves as a conceptual representation clarifying the interaction between the macro autoencoder and the return-prediction module. The full architecture details of the autoencoder, including hidden-layer configurations and activation functions, are provided in Internet Appendix
C.3
. The empirical findings underscore the importance of representing macroeconomic regimes in shaping cross-sectional return dynamics and highlight the utility of neural representation learning in extracting economically salient signals from high-dimensional macro data. At each expanding-window step, the autoencoder is trained only on macro data available up to the window end date, and the encoder is then used to compute
𝐳
t
\mathbf{z}_{t}
for that window’s validation and test months, thereby preventing look-ahead bias.

[Insert Figure
3
here]

An ablation study, presented in Internet Appendix
D.5.1
, further confirms the contribution of this component. Removing the autoencoder from CB-APM leads to a pronounced deterioration in predictive performance—particularly under higher
λ
\lambda
values—demonstrating that the learned macroeconomic embedding is essential to preserving both interpretability and accuracy. These results underscore that macroeconomic state conditioning is not a redundant extension but a core mechanism that stabilizes learning and improves out-of-sample generalization.

Finally, Appendix
D.3
provides direct empirical evidence that the learned macroeconomic embeddings are economically revelatory. A two-dimensional projection of the 32-dimensional latent vectors
9
9

9

We apply PCA to reduce the 32-dimensional latent state vectors to two dimensions only for visualization purpose.
reveals a smooth temporal trajectory that coherently tracks major macroeconomic transitions, including distinct clusters corresponding to National Bureau of Economic Research(NBER)-defined recessions such as the 2001 and 2008 downturns. Beyond these discrete regime shifts, the latent trajectory captures gradual cyclical and structural evolutions in the U.S. economy, reflecting shifts in growth, inflation, and monetary policy regimes. Collectively, these findings validate that the autoencoder encodes meaningful macro-financial state dynamics rather than statistical artifacts, yielding a compact and economically coherent representation that conditions firm-level return predictions within the CB-APM framework.

3.3
Expanding window approach for model evaluation

To evaluate model performance under realistic and evolving market conditions, this study employs an expanding window as a sample splitting scheme. Unlike static train-validation-test splits, the expanding window approach incrementally grows the training dataset over time while keeping the validation and testing sets fixed in size. This dynamic design mirrors the constraints of real-world applications, where future regimes are unknown and models must generalize across economic environments without the benefit of hindsight. By gradually shifting the end point of the training set forward, the expanding window simulates a time-consistent learning process that naturally adapts to structural changes in the data. As a result, this framework offers both methodological rigor and practical relevance, allowing the model to be evaluated not only on statistical metrics but also on its robustness across different economic cycles.

Figure
4
illustrates the expanding window approach for dataset partitioning, with the arrow along the bottom denoting the timeline of the window. The validation dataset spans two years, while the testing dataset spans a single year. Starting from the training set from January 1994 to December 2010, each training window ends at December of a given year, and subsequently expands by one year for the next window. This process continues sequentially, ensuring that the testing datasets do not overlap in any window. Consequently, the complete testing set spans from January, 2013 to December, 2022.
10
10

10

The final year of the dataset (January–December 2023) is reserved solely for constructing annual stock returns, as computing these returns requires at least one full year of subsequent observations.

[Insert Figure
4
here]

4
Empirical Results

4.1
Cross-section of consensuses and stock returns

This section presents empirical results on the cross-sectional prediction of stock returns and consensus variables. We evaluate predictive performance under varying forecast horizons
h
h
from equation (
3
) to assess the effectiveness of the consensus-bottleneck in asset pricing. Out-of-sample
R
2
R^{2}
is used as the primary evaluation metric and is defined as:

R
return
2
=
1
−
∑
i
=
1
N
∑
t
=
1
T
(
R
i
,
t
+
h
−
R
^
i
,
t
+
h
)
2
∑
i
=
1
N
∑
t
=
1
T
R
i
,
t
+
h
2
,
R^{2}_{\text{return}}=1-\frac{\sum_{i=1}^{N}\sum_{t=1}^{T}\bigl(R_{i,t+h}-\hat{R}_{i,t+h}\bigr)^{2}}{\sum_{i=1}^{N}\sum_{t=1}^{T}R_{i,t+h}^{2}},

for return prediction and,

R
consensus
2
=
1
−
∑
i
=
1
N
∑
t
=
1
T
(
C
i
,
t
−
C
^
i
,
t
)
2
∑
i
=
1
N
∑
t
=
1
T
C
i
,
t
2
,
R^{2}_{\text{consensus}}=1-\frac{\sum_{i=1}^{N}\sum_{t=1}^{T}\bigl(C_{i,t}-\hat{C}_{i,t}\bigr)^{2}}{\sum_{i=1}^{N}\sum_{t=1}^{T}C_{i,t}^{2}},

for consensus approximation, where
N
N
and
T
T
denote the number of firms and time periods, respectively.

While much of the asset pricing literature emphasizes short-horizon return forecasts, sell-side analysts typically issue multi-quarter to annual forecasts. Consensus measures therefore reflect longer-term expectations about fundamentals and risk premia rather than short-term price fluctuations. Evaluating the consensus-bottleneck over horizons that align with analysts’ forecast horizons is more economically relevant than using noisy short-term intervals. Accordingly, we focus on annual return prediction, consistent with prior studies on long-horizon predictability
(
Gu et al., 2020
;
Leippold et al., 2022
)
.
11
11

11

The results for other forecasting horizons are provided in Internet Appendix
D.1

Table
1
reports the monthly out-of-sample
R
2
R^{2}
values (in percentage) for both annual stock return prediction (
R
t
+
12
R_{t+12}
) and the approximation of analysts’ consensus variables (
C
t
C_{t}
) across different values of the regularization parameter
λ
\lambda
. The benchmark case (
λ
=
0
\lambda=0
), which excludes consensus learning, yields an annual return
R
2
R^{2}
of 7.63%, serving as a baseline for evaluating the incremental benefits of integrating consensus prediction into the CB-APM framework.

[Insert Table
1
here]

Introducing consensus learning via
λ
>
0
\lambda>0
leads to a pronounced improvement in return predictability. The out-of-sample
R
2
R^{2}
for annual returns rises steadily, peaking at 10.46% when
λ
=
0.3
\lambda=0.3
, a 37% increase relative to the benchmark. While larger
λ
\lambda
values beyond
0.3
0.3
result in a gradual decline in
R
2
R^{2}
, it is noteworthy that even at
λ
=
1.0
\lambda=1.0
, the return forecasting accuracy remains above the benchmark case (9.37% versus 7.63%), demonstrating that the integration of consensus information provides robust predictive gains across all tested settings.

The consensus approximation results provide further insight into this regularization effect. Among the nine consensus variables,
Analyst Earnings per Share
dominates, achieving an
R
2
R^{2}
of 71.43% at
λ
=
1.0
\lambda=1.0
, followed by strong performance in
EPS Forecast Dispersion
and
Analyst Optimism
. These results corroborate empirical findings that earnings estimates and their associated dispersion contain salient information about future returns
(
Diether et al., 2002
;
Jegadeesh et al., 2004
)
. By contrast,
Change in Recommendation
exhibits persistently negative
R
2
R^{2}
, consistent with prior evidence of limited incremental predictive content in recommendation changes once earnings revisions are accounted for.

The consensus average
R
2
R^{2}
increases monotonically from 7.33% at
λ
=
0.1
\lambda=0.1
to 24.21% at
λ
=
1.0
\lambda=1.0
, indicating that the model becomes progressively better at reconstructing analyst consensus as
λ
\lambda
grows. However, the modest decline in return
R
2
R^{2}
beyond
λ
=
0.3
\lambda=0.3
reflects the trade-off inherent in joint optimization; while higher
λ
\lambda
emphasizes consensus approximation, return forecasting benefits most when consensus serves as an auxiliary concept rather than the dominant objective.

Figure
5
complements Table
1
by visualizing these trends. The left panel shows how return predictability improves sharply with the introduction of consensus learning, peaks around
λ
=
0.3
\lambda=0.3
-
0.4
0.4
, and then tapers slightly while remaining above the benchmark even at
λ
=
1.0
\lambda=1.0
. The right panel demonstrates the monotonic improvement in consensus approximation with increasing
λ
\lambda
, eventually plateauing near 24%. Together, these panels illustrate the trade-off, where moderate
λ
\lambda
balances return prediction and consensus learning most effectively, while larger
λ
\lambda
values shift focus toward consensus reconstruction.

[Insert Figure
5
here]

Collectively, these results validate the core design of CB-APM that by incorporating consensus learning as a concept-bottleneck enhances return prediction while retaining interpretability. The model’s ability to achieve robust gains across different market environments underscores both its practical relevance under realistic, expanding-window evaluation and its theoretical grounding in analyst-driven information aggregation.

While the out-of-sample
R
2
R^{2}
metrics directly capture forecasting accuracy, they do not reveal how the joint loss function in equation (
4
) balances return prediction and consensus approximation during training. To address this, Internet Appendix
D.2
provides additional evidence on the optimization dynamics of CB-APM by reporting the in-sample MSE, which demonstrates that, at short horizons, increasing
λ
\lambda
introduces the expected trade-off between predictive accuracy and consensus reconstruction, whereas at longer horizons the two objectives reinforce each other, yielding what we term an interpretability-accuracy amplification effect.

4.2
Portfolio-based pricing validation

We conduct further empirical analysis of the CB-APM by examining its economic implications through portfolio-level tests. While the preceding sections evaluated the model’s predictive and explanatory power using out-of-sample
R
2
R^{2}
metrics, these statistical measures alone do not reveal whether the predicted returns merely reflect transitory noise. Portfolio-based analyses provide a more direct and economically interpretable assessment of model performance by linking cross-sectional predictions to realized investment payoffs. In particular, if the CB-APM successfully extracts a priced component of expected returns from the consensus structure, portfolios formed on its predictions should yield monotonic and persistent return differentials across quantiles.

Our portfolio analysis proceeds in three steps. First, we perform single-sort tests that rank stocks by CB-APM-predicted annual returns to evaluate the model’s raw cross-sectional discriminating power. These tests quantify whether higher model-implied expected returns translate into higher realized payoffs and whether the strength of this relationship varies with the degree of consensus regularization. Second, we conduct double-sort analyses that jointly sort stocks by both predicted returns and consensus variables to examine how the model’s inferred expectations interact with, and potentially refine, traditional analyst forecasts. Finally, we form long-short portfolios based on out-of-sample CB-APM predictions to evaluate their risk-adjusted performance relative to benchmark strategies and to assess the model’s practical value from an asset-management perspective.

These portfolio-level analyses allow us to connect the statistical accuracy of the CB-APM to its economic relevance. By translating predictive signals into realized return differentials, we can determine whether the consensus-bottleneck representation captures genuinely priced information—consistent with rational risk compensation—or reflects transitory deviations unrelated to systematic risk premia. The following subsections detail the construction of these portfolio tests and discuss their empirical results.

4.2.1
Portfolio sorts on approximated consensuses and expected returns

For each month in the out-of-sample evaluation period, the CB-APM produces annual return forecasts for all stocks. Based on these out-of-sample predictions, stocks are ranked by their expected returns and assigned to ten value-weighted decile portfolios, ranging from the lowest (decile 1) to the highest (decile 10) predicted-return group. Portfolio constituents and weights are updated monthly as new forecasts become available, ensuring that portfolio formation relies exclusively on information observable at the prediction date. The realized monthly returns of each decile are then computed over the subsequent month, thereby evaluating the model’s ex-ante forecasts in a strictly out-of-sample setting.

[Insert Table
2
here]

The single-sort portfolio results in Table
2
reinforce the predictive validity of the CB-APM framework in the cross-section of returns. Average realized returns increase monotonically from the lowest to the highest predicted-return decile, with the bottom portfolios consistently yielding negative returns and the top portfolios earning approximately 1.3% per month. The resulting high-minus-low (H–L) spreads range from 1.64% for the naïve neural network (
λ
=
0
\lambda=0
) to around 2.3% for regularized CB-APM specifications (
λ
≥
0.3
\lambda\geq 0.3
). This progressive widening of the return differential highlights the model’s ability to produce more economically meaningful and stable return rankings as the degree of consensus regularization increases. Beyond the level effects, the distribution of decile returns also becomes smoother and more monotonic as
λ
\lambda
rises, suggesting that the bottleneck constraint mitigates noise in the model-implied expected returns.

The patterns in portfolio payoffs align closely with the out-of-sample performance metrics reported in Table
1
. While the predictive
R
2
R^{2}
for stock returns peaks around 10% and remains relatively stable across higher
λ
\lambda
values, the
R
2
R^{2}
for consensus variable approximation improves dramatically—from roughly 7% at
λ
=
0.1
\lambda=0.1
to over 24% at
λ
=
1.0
\lambda=1.0
. This joint evidence implies that better recovery of analysts’ consensus structure translates into more reliable expected-return forecasts. In other words, the improvement in cross-sectional pricing performance—as captured by the H–L spread—parallels the enhanced interpretability and generalization observed in the consensus approximation task. Together, the results indicate that the consensus-bottleneck regularization enables the model to balance flexibility and economic discipline, yielding forecasts that are both interpretable and empirically potent in explaining the cross-section of returns.

To further examine the pricing content embedded in CB-APM forecasts, we conduct a double-sorting exercise based on the model-implied expected returns and the analyst-based measure
FEPS
. At each month in the out-of-sample period, all stocks are first assigned to quintiles using their CB-APM-approximated
FEPS
levels. Within each
FEPS
group, stocks are then independently sorted into quintiles by their CB-APM-predicted annual returns. This procedure yields 5×5 portfolios rebalanced monthly, ensuring that both the sorting signal and subsequent return evaluation rely strictly on information available at the prediction date. For each panel, the bottom and rightmost rows report high-minus-low (H–L) spreads along the predicted-return and consensus dimensions, measuring the incremental ordering power of CB-APM forecasts conditional on
FEPS
.

[Insert Table
3
here]

Table
3
shows that the CB-APM generates economically meaningful spreads across both sorting dimensions, highlighting an interaction between model-implied expected returns and analysts’ earnings expectations. The
FEPS
variable—the most recent I/B/E/S consensus forecast of next-fiscal-year earnings per share—is widely used as a standardized proxy for expected profitability. Prior evidence from
Cen (2006)
demonstrates that
FEPS
predicts future returns even after controlling for common risk factors, with the premium concentrated among small and neglected firms and persisting without reversal. These patterns suggest that
FEPS
embeds both valuable information about firm fundamentals and systematic expectation errors.

The double-sort design provides a natural setting to assess how the CB-APM processes this dual nature of analyst expectations. By construction, the model’s consensus-bottleneck is designed to extract the priced component of forecasted earnings while mitigating noise arising from optimism-driven biases. This mechanism is consistent with recent evidence such as
Palley et al. (2025)
, who document that consensus signals become unreliable when analyst dispersion is high, a condition strongly associated with stale or incentive-driven optimism. The state-dependent attenuation visible in Table
3
—where CB-APM’s expected-return differentiation is largest in low-
FEPS
states and diminishes as optimism rises—is precisely the pattern one would expect if behavioral components contaminate raw analyst forecasts while the model selectively filters them.

Across all regularization levels
λ
\lambda
, mean realized returns increase monotonically from the lower-left (low
FEPS
, low predicted return) to the upper-right (high
FEPS
, high predicted return), confirming strong joint ordering power. Within each
FEPS
quintile, the predicted-return portfolios exhibit clear monotonicity, with H–L spreads ranging from roughly 0.9% to 2.5% per month. These spreads peak at intermediate regularization strengths (
λ
=
0.3
\lambda=0.3
–
0.6
0.6
), consistent with the interpretation that moderate consensus constraints balance flexibility with economic discipline, whereas very small
λ
\lambda
introduces noise and very large
λ
\lambda
(
>
0.8
>0.8
) leads to over-regularization.

More revealing is the cross-sectional pattern along the
FEPS
dimension. The H–L spreads for
FEPS
are positive among stocks with low model-predicted returns but turn negative among those with high predicted returns. This inversion indicates that firms with high analyst-forecasted earnings outperform in segments where the model sees limited return potential but underperform where the model projects high returns. Simultaneously, the magnitude of the expected-return H–L spread declines systematically from low to high
FEPS
quintiles. Taken together, these findings imply that the CB-APM’s return signal is most potent precisely where analyst optimism is weakest, reinforcing the idea that the model distinguishes fundamental information from optimism-induced distortions.

These results extend the regularities documented by
Cen (2006)
. Although
FEPS
generally predicts higher future returns, the largest expectation errors occur where forecasts are pessimistic, allowing the CB-APM to retain their predictive content while tempering the behavioral component. The observed reversals in the double-sort tables thus reflect not contradictions but adjustments: the CB-APM internalizes the asymmetric way markets react to forecasted earnings, preserving the informative component of
FEPS
while reweighting it in states where optimism clouds the signal.

Overall, the evidence indicates that CB-APM forecasts complement rather than replicate the information in
FEPS
. The consensus-bottleneck extracts the priced, risk-aligned component of analysts’ expectations while filtering optimism-related noise. The resulting reversal and attenuation patterns provide direct support for the interpretation that the CB-APM transforms raw forecasted earnings into a state-dependent pricing signal that refines, rather than contradicts, the analysts’ consensus view.

4.2.2
Long-short portfolio performance

We construct the long-short portfolio as follows. The first step involves generating monthly predicted annual returns for each stock within the universe from CB-APM. These predicted returns are then ranked from highest to lowest and sorted into deciles based on their values. Subsequently, a long portfolio is formed by purchasing the top 10% of stocks with the highest predicted returns, while concurrently establishing a short portfolio by selling the bottom 10% of stocks with the lowest predicted returns. Weighting of the stocks within each portfolio is executed based on the size of the firm, ensuring that larger firms are assigned higher weights. Then the long-short portfolio is rebalanced every month to uphold the desired exposure and maintain alignment with the initial strategy.

The long-short construction directly operationalizes the cross-sectional ordering evidence reported in Tables
2
. The monotonic increase in realized returns across predicted-return deciles translates naturally into economically significant long-short spreads.

To evaluate the risk-adjusted performance of the CB-APM portfolio, we compute seven portfolio metrics: monthly mean log return, standard deviation, cumulative log return, annualized Sharpe ratio, maximum one-month loss, maximum drawdown, and turnover rate. Monthly mean and cumulative returns quantify the overall profitability of the model, while the Sharpe ratio measures risk-adjusted performance by relating expected excess returns to return volatility. Maximum one-month loss and maximum drawdown capture downside risk by quantifying the worst historical losses, both in single periods and cumulatively. Finally, portfolio turnover measures the degree of portfolio rebalancing activity, which is directly linked to transaction costs and practical implementability.

Maximum drawdown (Max DD) is defined as the largest cumulative loss from a historical peak in portfolio wealth:

Max DD
=
max
t
∈
T
⁡
(
1
−
W
t
max
τ
≤
t
⁡
W
τ
)
,
W
t
=
∏
τ
=
1
t
(
1
+
R
τ
)
,
\text{Max DD}=\max_{t\in T}\left(1-\frac{W_{t}}{\max_{\tau\leq t}W_{\tau}}\right),\qquad W_{t}=\prod_{\tau=1}^{t}(1+R_{\tau}),

where
W
t
W_{t}
denotes cumulative portfolio wealth at time
t
t
.
This measure captures the worst peak-to-trough decline experienced over the sample period.

Portfolio turnover is calculated as,

Turnover
=
1
T
r
​
∑
t
∈
T
r
(
∑
i
=
1
N
|
w
i
,
t
+
1
−
w
i
,
t
​
(
1
+
R
i
,
t
)
1
+
∑
j
=
1
N
w
j
,
t
​
R
j
,
t
|
)
,
\text{Turnover}=\frac{1}{T_{r}}\sum_{t\in T_{r}}\left(\sum_{i=1}^{N}\left|w_{i,t+1}-\frac{w_{i,t}\left(1+R_{i,t}\right)}{1+\sum_{j=1}^{N}w_{j,t}R_{j,t}}\right|\right),

(7)

where
w
i
,
t
w_{i,t}
denotes the portfolio weight of asset
i
i
at time
t
t
,

R
i
,
t
R_{i,t}
is its arithmetic monthly return,
and
T
r
⊂
T
T_{r}\subset T
denotes the set of rebalancing dates.

Portfolio positions are formed using CB-APM’s out-of-sample return forecasts, allowing the portfolio tests to evaluate genuine real-time predictability over a long-horizon target.

[Insert Table
4
here]

The portfolio performance results in Table
4
mirror the statistical improvements in predictive and explanatory performance documented in Table
1
. As the hyperparameter
λ
\lambda
increases to moderate values around 0.3–0.4, both out-of-sample return
R
2
R^{2}
and consensus-approximation accuracy rise sharply, and this improvement translates directly into superior realized portfolio returns. Mean monthly log returns climb from 1.53% at
λ
=
0
\lambda=0
to 2.20% at
λ
=
0.3
\lambda=0.3
, while the annualized Sharpe ratio concurrently increases from 1.10 to 1.44. This near one-to-one correspondence between predictive power and portfolio profitability substantiates the economic value of the consensus-bottleneck: the same mechanism that refines predictive signal extraction in-sample also enhances risk-adjusted returns out-of-sample.

Beyond moderate
λ
\lambda
values, both predictive and portfolio metrics exhibit mild flattening, as excessive weighting on consensus reconstruction (
λ
>
0.4
\lambda>0.4
) marginally reduces return
R
2
R^{2}
and diminishes economic gains. This pattern implies a practical upper bound to interpretability regularization, beyond which the model overemphasizes consensus consistency at the expense of direct return optimization. Nonetheless, even at high
λ
\lambda
values, performance remains consistently above the benchmark, confirming that consensus learning contributes persistently to economically meaningful predictability rather than statistical overfitting.

Risk profiles exhibit a moderate but economically intuitive trade-off between profitability and downside exposure. As
λ
\lambda
increases to 0.3–0.4, maximum one-month losses rise slightly relative to the naïve network (
λ
=
0
\lambda=0
), while remaining of similar magnitude at
λ
=
0.3
\lambda=0.3
, which yields the highest Sharpe ratio. Maximum drawdowns, by contrast, are consistently lower than those of the S&P 500 benchmark—staying below 21% versus the market’s 25%—indicating that CB-APM’s consensus-regularized predictions generate smoother long-term wealth trajectories. The modest increase in short-horizon losses is more than compensated by the substantial improvement in mean return and Sharpe ratio, implying enhanced efficiency on a risk-adjusted basis. Overall, the co-movement of predictive
R
2
R^{2}
, Sharpe ratios, and drawdown behavior captures an economically meaningful balance between return amplification and risk containment, reflecting the emergence of stable, consensus-aligned risk premia rather than transient noise-fitting effects.

Portfolio turnover remains high---approximately 60% per month---which is consistent with the characteristics of complex nonlinear architectures.
12
12

12

A formal transaction‐cost analysis based on the turnover definition in Equation (
7
) is provided in Internet Appendix
D.4.3
. The results show that the main economic conclusions are robust to proportional trading costs.
This observation aligns with the findings of
Gu et al. (2020)
, suggesting that neural-network-based return predictors typically produce higher turnover than linear or tree-based models due to their greater sensitivity to small shifts in cross-sectional signals. While
Kelly et al. (2024)
argue that out-of-sample predictive
R
2
R^{2}
and Sharpe ratios of characteristics-sorted portfolios may not always constitute decisive evidence of pricing relevance, the convergence of both statistical and economic measures in CB-APM suggests that its latent consensus components capture systematically priced information that conventional deep learning frameworks fail to isolate. Together, these results affirm that CB-APM’s consensus-bottleneck not only improves explanatory power but also yields tangible, risk-adjusted portfolio benefits, linking interpretability and profitability within a unified empirical asset pricing framework.

[Insert Figure
7
here]

Figure
7
visualizes the cumulative out-of-sample performance of CB-APM long-short portfolios across different regularization strengths
λ
\lambda
. All neural-network portfolios substantially outperform the S&P 500 buy-and-hold benchmark (black dashed line), demonstrating that the model’s predictive signals translate into economically meaningful excess returns. The naïve network (
λ
=
0
\lambda=0
, purple line) already yields notable outperformance relative to the market, yet introducing the consensus-bottleneck regularization (
λ
>
0
\lambda>0
) substantially elevates cumulative returns. Portfolio performance improves sharply up to
λ
≈
0.3
\lambda\approx 0.3
, after which cumulative returns remain at a comparably high level with minor oscillations across subsequent
λ
\lambda
values. The best-performing specification at
λ
=
1.0
\lambda=1.0
represents a continuation of this high-return plateau rather than a strict monotonic gain, highlighting the robustness of CB-APM’s economic performance across a wide range of regularization intensities. This stability suggests that consensus regularization consistently enhances the model’s predictive and economic relevance without overfitting to a narrow hyperparameter regime.

The figure further highlights the temporal robustness of CB-APM’s performance. Even during adverse market conditions—notably the 2020 downturn—consensus-regularized portfolios experience smaller and more rapidly recovered drawdowns relative to both the market and the unregularized model, reflecting smoother wealth accumulation and improved resilience to macro shocks. The consistent separation between the consensus-based portfolios and the S&P 500 benchmark indicates that the learned consensus representations capture priced information that is both persistent and broadly exploitable.

5
Dissecting Approximated Consensuses

The CB-APM framework is designed not only to forecast risk premia but also to provide a transparent interpretation of how firm- and macro-level information maps into priced return variation.
Unlike most machine-learning predictors—which typically compress characteristics into opaque nonlinear transformations—the CB-APM architecture explicitly separates two economic mechanisms:
(i) a nonlinear mapping that synthesizes the high-dimensional information set into consensus-like latent expectations, and (ii) a final linear stage that maps these expectations into forecasts of future returns. This structural decomposition allows the consensus layer to be interpreted as a set of economically meaningful conditional expectations, while the final linear layer mirrors the role of factor loadings in a traditional cross-sectional model.

[Insert Figure
6
here]

Figure
6
visualizes the estimated prediction-layer coefficients at
(
λ
=
1
)
(\lambda=1)
, computed using expanding training windows.
13
13

13

We focus on
λ
=
1
\lambda=1
because it delivers the highest out-of-sample
R
2
R^{2}
for consensus approximation. Analyzing the most accurate consensus-reconstruction specification provides the clearest window into how CB-APM translates analyst information into interpretable pricing components.

Each coefficient reflects the model’s inferred sensitivity of expected returns to a given consensus element, while the color shading indicates the corresponding out-of-sample
R
2
R^{2}
for consensus approximation. Because the prediction module is linear, these coefficients admit a familiar interpretation: they reveal the direction and magnitude with which each consensus dimension influences expected returns, analogous to factor loadings in conventional asset pricing regressions.

Several patterns emerge. First, sentiment-oriented variables such as
Analyst Optimism
load positively and persistently, indicating that firms with stronger analyst sentiment are assigned higher expected-return forecasts. In contrast, variables reflecting recommendation changes or forecast revisions often load negatively, suggesting that optimistic updates embed short-lived overreaction that subsequently reverses. Second, consensus dimensions that the model reconstructs more accurately—particularly dispersion- and accrual-related variables—tend to receive larger-magnitude coefficients.
This alignment between approximation quality and economic relevance implies that the CB-APM’s interpretive layer concentrates information in the dimensions where analyst signals are both reliably reconstructable and strongly predictive of return heterogeneity.

These observations underscore an important conceptual feature: the interpretable consensus layer can be evaluated independently of the model’s nonlinear feature-extraction stage.
Once the consensus mapping is estimated, the subsequent linear relation

R
^
i
,
t
+
h
=
a
+
𝒃
⊤
​
C
^
i
,
t
\hat{R}_{i,t+h}=a+\bm{b}^{\top}\hat{C}_{i,t}

can be analyzed using the same tools employed to study traditional factor models.
This allows us to examine, in a transparent and economically interpretable manner, whether the learned consensus dimensions behave like priced sources of return variation or simply capture information-based heterogeneity unrelated to systematic risk.

To formalize this connection, we align our empirical strategy with standard asset pricing methodology and implement two complementary tests. First, we estimate pooled panel OLS regressions of annual stock returns on either raw analyst consensus variables or their CB-APM–inferred counterparts, thereby quantifying the incremental explanatory content gained through the consensus-bottleneck transformation. Second, we examine whether factor-mimicking portfolios—constructed from the consensus dimensions via decile sorts—span the SDF by applying the GRS test for mean–variance efficiency
(
Gibbons et al., 1989
)
. These analyses serve a dual purpose: they link the interpretability of CB-APM’s consensus layer to established empirical asset pricing tools, and they enable a direct assessment of whether the machine-inferred beliefs embody priced economic content beyond what is observable from raw analyst forecasts.

5.1
Comparative regression analysis

Having established that CB-APM’s interpretable layer produces economically meaningful consensus coefficients, we examine whether variations in the CB-APM–implied consensus translate into priced differences in expected annual returns, following the empirical design of standard asset pricing regressions. These regressions are not intended as structural pricing tests; rather, they serve as diagnostic tools that evaluate whether the model’s consensus representations capture priced variation more effectively than the raw analyst signals.

Table
5
reports pooled OLS regressions that relate future annual stock returns to either raw analyst consensus variables or CB-APM-implied consensus at
(
λ
=
1
)
(\lambda=1)
. We estimate the following pooled panel regression:

R
i
,
t
+
h
=
a
(
C
)
+
𝒃
(
C
)
⊤
​
C
i
,
t
+
ε
i
,
t
+
h
(
C
)
or
R
i
,
t
+
h
=
a
(
C
^
)
+
𝒃
(
C
^
)
⊤
​
C
^
i
,
t
+
ε
i
,
t
+
h
(
C
^
)
.
R_{i,t+h}=a^{(C)}+\bm{b}^{(C)\top}C_{i,t}+\varepsilon^{(C)}_{i,t+h}\quad\text{or}\quad R_{i,t+h}=a^{(\hat{C})}+\bm{b}^{(\hat{C})\top}\hat{C}_{i,t}+\varepsilon^{(\hat{C})}_{i,t+h}.

(8)

For the analysis, we set
h
=
12
h=12
to focus on annual return predictability, thereby aligning the return horizon with prior empirical studies discussed in the preceding sections. The model is estimated on the stacked cross-section of firm-month
(
i
,
t
)
(i,t)
observations to obtain a time-invariant coefficient vector
𝒃
^
\widehat{\bm{b}}
. Inference is based on heteroskedasticity-robust covariance estimation tailored to the dependent variable’s structure. To handle an overlapping long-horizon return, we compute Driscoll–Kraay (kernel HAC) standard errors with a Bartlett kernel and an eleven-month bandwidth, which are robust to heteroskedasticity, cross-sectional dependence, and the serial correlation induced by overlapping observations
(
Driscoll and Kraay, 1998
;
Newey and West, 1986
;
Hodrick, 1992
)
.

Panel A reports coefficient estimates,
t
t
-statistics, and a variable-level fit measure; Panel B summarizes the intercept and overall adjusted
R
2
R^{2}
. The comparison isolates the incremental explanatory content obtained when analyst information is first synthesized by CB-APM’s consensus-bottleneck and then mapped linearly into expected returns.

[Insert Table
5
here]

The pooled OLS regression using raw analyst consensus variables yields limited explanatory power, with an adjusted
R
2
R^{2}
of just 0.40%. Most predictors exhibit weak statistical significance; for example,
EPS forecast revision
and
Earnings forecast revisions
produce
t
t
-statistics of
−
1.31
-1.31
and
−
0.62
-0.62
, respectively, with neither variable exhibiting meaningful predictive content. Although
Change in recommendation
(
t
=
11.54
t=11.54
) and
Change in Forecast and Accrual
(
t
=
8.57
t=8.57
) are statistically significant at the 1% level, their estimated effects are modest in magnitude, and the overall model fit remains poor.

By contrast, the regression using CB-APM-inferred consensus achieves a substantially higher adjusted
R
2
R^{2}
of 8.35%, representing more than a twentyfold improvement in explanatory power. A few key coefficients also reverse in sign relative to their raw counterparts. For instance, the coefficient on
Change in recommendation
shifts from
+
0.0307
+0.0307
(
t
=
11.54
t=11.54
) to
−
3.9080
-3.9080
(
t
=
−
6.67
t=-6.67
), while
EPS Forecast Dispersion
turns from
+
0.0076
+0.0076
(
t
=
0.59
t=0.59
) to
−
0.6263
-0.6263
(
t
=
−
4.87
t=-4.87
). In addition,
Analyst earnings per share
becomes strongly positive and significant (
t
=
4.46
t=4.46
), whereas
Analyst Value
becomes significantly negative (
t
=
−
2.75
t=-2.75
). These shifts suggest that CB-APM extracts transformed representations that encode economically distinct pricing content beyond the raw analyst signals.

While the CB-APM-inferred consensus variables deliver substantial improvements in explanatory power, caution is warranted in interpreting their individual coefficients. As shown in Table
5
, consensus variables with low predictor-level approximation
R
2
R^{2}
values often exhibit coefficient patterns that diverge from those estimated using raw consensus inputs. For example,
Change in recommendation
, which has one of the lowest approximation
R
2
R^{2}
values, exhibits a pronounced sign reversal, while
Change in Forecast and Accrual
weakens substantially in magnitude. This pattern highlights that the CB-APM approximations are not one-to-one reconstructions of analyst beliefs but rather encode transformed features with distinct pricing implications.

Consequently, interpretability must be grounded in a dual-lens framework. The consensus-level approximation
R
2
R^{2}
, previously reported in Table
1
, reflects the degree to which a model-inferred variable aligns with its human-interpretable counterpart, whereas the coefficient estimate from the return regression captures the economic relevance of that signal. Coefficients associated with well-approximated variables are more directly interpretable as refinements of analyst expectations, whereas those tied to poorly reconstructed signals likely reflect alternative representations or re-weightings learned by the model. Thus, proper interpretation requires joint consideration of both approximation fidelity and return sensitivity rather than treating coefficients in isolation.

Importantly, this improvement follows directly from the design of the framework, which trains the approximated consensus layer under a joint objective that simultaneously targets return prediction and consensus reconstruction. By doing so, the model synthesizes information from a wide set of firm-level characteristics and macroeconomic variables into consensus features that retain risk-relevant content while reducing noise. Although the reported
t
t
-statistics primarily capture in-sample explanatory strength and are not intended for direct investment use due to inherent look-ahead bias, they underscore that the approximated consensus simultaneously explains both realized analyst consensus and future returns. This dual property makes the learned consensus features a rich source of information that merits closer examination beyond forecasting alone.

Taken together, the results indicate that CB-APM’s consensus module extracts signals that are both more informative and more economically meaningful than raw analyst inputs. Although the reported regressions are estimated in-sample and do not directly measure out-of-sample predictive accuracy, they nonetheless support the model’s central objective: to learn interpretable latent representations that jointly capture analyst expectations and priced return variation. Proper interpretation of these results requires concurrent evaluation of (i) approximation
R
2
R^{2}
, which gauges the alignment between model-inferred signals and observable analyst variables, and (ii) coefficient sign and magnitude, which reflect the economic relevance of each signal for cross-sectional return prediction.

5.2
Pricing error of test assets

The preceding analysis establishes that the CB-APM’s interpretable consensus layer captures return-relevant structure in the cross-section of individual stocks. We now examine whether these signals possess asset pricing content when evaluated through the lens of linear factor models. Specifically, we assess (i) whether the latent representations learned by CB-APM
14
14

14

It is important to point out that the consensus representations learned by the CB-APM are not designed to approximate the span of the SDF. Rather, the architecture learns a set of conditional expectation operators that map firm characteristics and macroeconomic conditions into consensus-like forecasts of future fundamentals. These latent expectations summarize belief-based or information-based heterogeneity, not compensated sources of systematic risk. Consequently, the CB-APM should not be expected to replicate the factor structure implicit in linear SDF models; instead, its consensus layer provides an economically interpretable decomposition of expected returns that is complementary to—rather than a substitute for—the traditional factor space.
can serve as risk factors capable of
pricing standard benchmark portfolios, and (ii) whether traditional factor models can price the return patterns implied by the CB-APM’s predictions and consensus-based characteristics. To do so, we employ the multivariate GRS test
(
Gibbons et al., 1989
)
, which jointly evaluates whether the intercepts (
𝜶
\bm{\alpha}
) in time-series regressions are statistically different from zero.

We consider three sets of standard test portfolios widely used in empirical asset pricing: the Fama–French 25 portfolios sorted on size and book-to-market ratio (
5
×
5
5\times 5
), the 25 portfolios sorted on size and momentum, and the 30 value-weighted industry portfolios. These portfolios span well-known sources of cross-sectional variation linked to value, momentum, and industry structure, and thus provide a benchmark for evaluating alternative factor models. As reference models, we estimate the CAPM, the Fama–French three-factor model (FF3), the Carhart four-factor model, the Fama–French five-factor model (FF5), and the Fama–French six-factor model (FF6). All models are estimated using monthly excess returns, and all results are reported in-sample to maintain comparability with the standard evaluation framework in the factor-pricing literature.

A central element of the empirical design is the construction of tradable portfolios that proxy for the consensus signals extracted by the CB-APM. Because the model produces firm-level consensus measures rather than aggregate time-series factors, we translate each consensus dimension into a value-weighted long–short portfolio by sorting firms into deciles and taking the return spread between the highest and lowest deciles. This approach parallels the construction of empirically traded factors such as HML or UMD and yields a set of zero-investment portfolios whose returns reflect cross-sectional variation in the corresponding consensus dimension. These portfolios provide a tractable representation of the CB-APM signals within a traditional factor-pricing framework and permit a direct comparison with benchmark linear factor models using GRS tests.

Importantly, the CB-APM portfolios used in these tests are constructed from the same model configuration employed in the empirical return-forecasting exercise. That is, we apply the trained CB-APM—optimized to forecast annual excess returns—to generate firm-level predicted returns and consensus representations, which are then used to form sorted portfolios and factor-mimicking returns. This design ensures coherence across empirical sections: the factor-pricing analysis evaluates the economic content of the very signals that the CB-APM learns to use for long-horizon prediction.

We conduct three complementary GRS exercises. First, we evaluate whether the CB-APM factor-mimicking portfolios can jointly price the benchmark 25– and 30–portfolio test assets. Successful pricing performance would indicate that the consensus-based signals span systematic risks similar to those captured by traditional factors. Second, we form decile portfolios based on CB-APM predicted returns and test whether standard factor models can explain their realized returns. This analysis assesses whether the return patterns generated by the model are incremental to the span of existing factors. Third, we construct decile portfolios sorted on each individual consensus dimension and examine whether traditional models can price these portfolios. This final exercise isolates which consensus channels are most and least aligned with traditional factor structures.

Each specification is evaluated using the GRS
F
F
-statistic, its associated
p
p
-value, and mean absolute and root-mean-squared pricing errors. All results are computed in-sample, consistent with empirical asset pricing conventions in which factor-pricing tests focus on explaining cross-sectional return patterns rather than forecasting performance. Together, these exercises provide a comprehensive assessment of whether the consensus representations learned by the CB-APM contain distinct factor-pricing information or whether their explanatory power is largely captured by established benchmark models.

Tables
6
–
8
present a comprehensive set of in-sample GRS tests evaluating the pricing performance of the CB-APM relative to conventional factor models. Across all tests, the GRS
F
F
-statistic assesses the joint null hypothesis that all pricing errors (
𝜶
\bm{\alpha}
) are zero, such that lower
F
F
-statistics and higher
p
p
-values indicate superior mean–variance efficiency. The accompanying mean absolute and root-mean-squared alphas summarize the magnitude of mispricing across the
corresponding test assets.

[Insert Table
6
here]

Panels A–C of Table
6
evaluate whether the CB-APM’s consensus-based factor-mimicking portfolios can price the returns of the Fama–French 25 size–book-to-market portfolios, the 25 size–momentum portfolios, and the 30 industry portfolios. Across these benchmarks, the CB-APM factors deliver GRS statistics that are broadly comparable to those of standard models, but they remain somewhat higher than the Fama–French five-factor model and Fama–French six-factor model specifications. Mean and RMS pricing errors are likewise modest yet consistently larger than those generated by traditional factor structures. These results indicate that the consensus-based factors span meaningful components of systematic return variation, but not the full set captured by benchmark style factors. This is consistent with evidence that only a limited number of characteristic-based factors are strongly priced in the cross-section, while many signals are redundant or weakly informative
(
Kozak et al., 2020
)
. Within this environment, the CB-APM factors behave as an additional block of characteristic-sorted portfolios that contributes incremental explanatory variation without supplanting the canonical factor structure.

[Insert Table
7
here]

Table
7
examines whether traditional factor models can jointly price decile portfolios formed on the CB-APM’s predicted return scores. When the consensus-bottleneck is weak (small
λ
\lambda
), conventional factor models achieve moderate GRS statistics and economically small pricing errors, suggesting that a substantial portion of the model’s predictive content overlaps with standard style factors. As
λ
\lambda
increases, however, the GRS statistics rise sharply and the joint null of zero pricing errors is rejected uniformly. This monotonic deterioration indicates that stronger reliance on the consensus-bottleneck induces expected-return patterns that increasingly depart from the linear span of market, size, value, momentum, and profitability/investment factors. Conceptually, this aligns with evidence that machine-learning models often extract nonlinear or interaction-based transformations of firm characteristics that extend beyond linear factor structures
(
Freyberger et al., 2020
;
Gu et al., 2020
)
. In particular, the CB-APM with a tight consensus constraint appears to generate forecasts that incorporate structured forms of return heterogeneity that are difficult to reconcile with the standard factor space.

[Insert Table
8
here]

Table
8
evaluates portfolios formed on the individual consensus signals at
λ
=
1.0
\lambda=1.0
. Several dimensions—most prominently
Analyst Value
,
Analyst Optimism
, and
Analyst Earnings per Share
—produce relatively low GRS statistics and economically small pricing errors, suggesting strong alignment between these inferred consensus measures and established factor structures. Forecast-based and dispersion-based dimensions (such as
EPS forecast dispersion
and related revisions) exhibit somewhat larger pricing errors, but even here the magnitudes remain concentrated in the range of a few basis points per month. These patterns reinforce the idea that much of the predictive information contained in analyst-derived consensus measures can be represented through low-dimensional combinations of characteristics, often with sparse or localized influence
(
Chinco et al., 2019
)
, while still accommodating nonlinear interactions and heterogeneous partitions
(
Bryzgalova et al., 2025
)
. The CB-APM’s consensus variables therefore fit naturally within the broader empirical finding that return-relevant structure can be extracted by compressing high-dimensional characteristics into well-organized representations.

The three sets of GRS tests reveal how the CB-APM relates to the traditional factor space. First, consensus-based factor-mimicking portfolios do not fully price the classic benchmark portfolios, which indicates that the latent consensus dimensions do not function as close substitutes for the core priced factors. Second, the ability of traditional factor models to price CB-APM–generated portfolios deteriorates as the consensus-bottleneck becomes more stringent, implying that the model’s predictive signals progressively move outside the span of standard linear characteristics. This behavior is consistent with the broader view that, while the priced dimension of the SDF is relatively low, flexible methods can uncover structured forms of return heterogeneity that improve portfolio efficiency
(
Cong et al., 2025
)
without reproducing the canonical factors directly. Third, portfolios sorted on individual consensus dimensions exhibit moderate but non-negligible pricing errors, suggesting that the learned signals contain meaningful information about expected returns but do not themselves constitute a new standalone factor system.

This finding crucially aligns with the emerging methodological consensus that traditional characteristic–based sorting procedures fundamentally fail to capture the full mean–variance efficient (MVE) frontier due to their neglect of nonlinearity and asymmetric characteristic interactions. Recent goal-oriented machine learning approaches—most notably the Panel Tree (P-Tree) framework
(
Cong et al., 2025
)
and the Asset Pricing Tree (AP-Tree) framework
(
Bryzgalova et al., 2025
)
—demonstrate that test assets constructed by explicitly optimizing for SDF spanning or mean–variance efficiency are substantially harder to price with conventional factor models, often yielding extremely high GRS statistics. In this sense, the behavior observed in Table
7
echoes the insight that once test assets begin to reflect structured, state-dependent return heterogeneity, linear factor structures fail sharply.

The CB-APM achieves a conceptually parallel outcome, but through an economically structured consensus-bottleneck rather than recursive partitioning rules. By restricting predictive content to pass through interpretable consensus dimensions, the model induces return patterns that resemble the “goal-oriented” test assets emphasized in the tree-based literature—namely, portfolios that expose deficiencies in the linear factor span precisely because they encode higher-order interactions and conditional pricing structure. This makes the resulting portfolios harder to price not as a flaw, but as evidence that the model recovers meaningful variation in expected returns that traditional factor models systematically miss.

Overall, the evidence positions the CB-APM as a complementary asset pricing framework: it enhances cross-sectional return prediction by compressing analysts’ heterogeneous beliefs into interpretable consensus signals that partially overlap with—but do not collapse onto—the priced dimensions emphasized in modern work on characteristic-based factor representations
(
Cochrane, 2011
, e.g.,)
. At the same time, the results indicate that the CB-APM does not merely denoise or reweight analyst inputs. Instead, it isolates structured and economically relevant components of analyst-derived information that are priced in the cross-section. The model therefore reveals that analyst-based information contains priced elements that conventional factor models only partially span, and that the consensus-bottleneck organizes these elements into an interpretable and economically coherent representation. This places the CB-APM within a growing line of research showing that belief-based or characteristic-based signals can be synthesized into low-dimensional, economically meaningful components without reproducing the canonical factor structure directly.

6
Conclusion

This study introduces the CB-APM, a novel framework that integrates interpretable deep learning with empirical asset pricing. By embedding a concept-bottleneck architecture into a neural network, CB-APM not only achieves state-of-the-art predictive accuracy in cross-sectional stock return forecasts but also provides transparent insights into the role of analysts’ consensus in shaping risk premiums. Our empirical results demonstrate that interpretability and performance are not inherently conflicting. CB-APM outperforms conventional deep learning benchmarks in long-horizon forecasts while preserving a clear, economically grounded structure. By linking machine learning’s predictive capabilities with the theoretical underpinnings of financial economics, and by demonstrating that interpretable deep learning can yield both statistical and economic validity, this work offers a blueprint for building models that are both high-performing and aligned with established asset pricing principles.

The success of CB-APM highlights three key implications for empirical finance. First, interpretable neural architectures can reconcile the flexibility of machine learning with economic reasoning, enabling researchers to assess whether models capture meaningful risk factors rather than spurious correlations. Second, embedding interpretability directly within model design fosters transparency and trust, addressing the skepticism that often surrounds “black-box” methods in high-stakes financial applications. Third, by explicitly modeling analysts’ consensus as a latent mediator between firm characteristics and returns, CB-APM sheds new light on how information aggregation mechanisms influence asset prices, aligning closely with rational expectations theory and empirical evidence on analyst behavior.

Future research can extend this framework in several promising directions. Incorporating additional economically meaningful bottlenecks—such as investor sentiment or narrative-driven pricing component
(
Bybee et al., 2023
)
—could further disentangle the sources of risk premiums and strengthen the theoretical interpretability of model outputs. Addressing practical constraints, such as data latency in analyst consensus measures or improving computational efficiency for large-scale implementation, would enhance CB-APM’s applicability in real-world investment contexts. More broadly, as the “factor zoo” continues to grow, interpretable frameworks like CB-APM will be instrumental in bridging data-driven discovery with economic theory, offering a structured approach to understanding how high-dimensional predictors translate into priced information. By demonstrating that interpretable AI can achieve both predictive accuracy and theoretical coherence, this study lays the groundwork for a new generation of financially grounded machine learning models, advancing the study of asset pricing in both academic research and practical decision-making.

References

Ang and Bekaert (2007)

A. Ang and G. Bekaert

Stock return predictability: is it there?
.

The Review of Financial Studies

20
(
3
),
pp. 651–707
.

Cited by:
§1
.

Ba
et al.
(2016)

J. L. Ba, J. R. Kiros, and G. E. Hinton

Layer normalization
.

External Links:
1607.06450

Cited by:
§C.2
.

Barber
et al.
(2001)

B. Barber, R. Lehavy, M. McNichols, and B. Trueman

Can investors profit from the prophets? security analyst recommendations and stock returns
.

The Journal of finance

56
(
2
),
pp. 531–563
.

Cited by:
§1
.

Barron (2017)

J. T. Barron

Continuously differentiable exponential linear units
.

External Links:
1704.07483

Cited by:
§C.1
.

Benítez
et al.
(1997)

J. M. Benítez, J. L. Castro, and I. Requena

Are artificial neural networks black boxes?
.

IEEE Transactions on neural networks

8
(
5
),
pp. 1156–1164
.

Cited by:
Appendix A
.

Bessembinder (2003)

H. Bessembinder

Trade execution costs and market quality after decimalization
.

Journal of Financial and Quantitative Analysis

38
(
4
),
pp. 747–777
.

Cited by:
§D.4.3
.

Bianchi
et al.
(2021)

D. Bianchi, M. Büchner, and A. Tamoni

Bond risk premiums with machine learning
.

The Review of Financial Studies

34
(
2
),
pp. 1046–1089
.

Cited by:
§1
.

Bryzgalova
et al.
(2025)

S. Bryzgalova, M. Pelger, and J. Zhu

Forest through the trees: building cross-sections of stock returns
.

The Journal of Finance

80
(
5
),
pp. 2447–2506
.

Cited by:
§5.2
,

§5.2
.

Bybee
et al.
(2023)

L. Bybee, B. Kelly, and Y. Su

Narrative asset pricing: interpretable systematic risk factors from news text
.

The Review of Financial Studies

36
(
12
),
pp. 4759–4787
.

Cited by:
§6
,

footnote 1
.

Campbell and Thompson (2008)

J. Y. Campbell and S. B. Thompson

Predicting excess stock returns out of sample: can anything beat the historical average?
.

The Review of Financial Studies

21
(
4
),
pp. 1509–1531
.

Cited by:
§1
.

Cao
et al.
(2024)

S. Cao, W. Jiang, J. Wang, and B. Yang

From man vs. machine to man+ machine: the art and ai of stock analyses
.

Journal of Financial Economics

160
,
pp. 103910
.

Cited by:
§1
,

§1
.

Carhart (1997)

M. M. Carhart

On persistence in mutual fund performance
.

The Journal of finance

52
(
1
),
pp. 57–82
.

Cited by:
§1
.

Cen (2006)

L. Cen

Forecasted earnings per share and the cross section of expected stock returns
.

External Links:
Link

Cited by:
§4.2.1
,

§4.2.1
.

Chen and Zimmermann (2022)

A. Y. Chen and T. Zimmermann

Open source cross-sectional asset pricing
.

Critical Finance Review

27
(
2
),
pp. 207–264
.

Cited by:
§B.1
,

§B.2.1
,

§B.2.2
,

Table E.1
,

Table E.4
,

§3.1
,

§3.1
,

§3.1
,

footnote 5
.

Chen
et al.
(2024)

L. Chen, M. Pelger, and J. Zhu

Deep learning in asset pricing
.

Management Science

70
(
2
),
pp. 714–750
.

Cited by:
§B.2.1
,

§C.1
,

§D.3
,

§1
,

§3.2
.

Chen
et al.
(2016)

X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel

Infogan: interpretable representation learning by information maximizing generative adversarial nets
.

Advances in neural information processing systems

29
.

Cited by:
footnote 16
.

Chen
et al.
(2020)

Z. Chen, Y. Bei, and C. Rudin

Concept whitening for interpretable image recognition
.

Nature Machine Intelligence

2
(
12
),
pp. 772–782
.

Cited by:
footnote 16
.

Chinco
et al.
(2019)

A. Chinco, A. D. Clark-Joseph, and M. Ye

Sparse signals in the cross-section of returns
.

The Journal of Finance

74
(
1
),
pp. 449–492
.

Cited by:
§5.2
.

Clevert
et al.
(2016)

D. Clevert, T. Unterthiner, and S. Hochreiter

Fast and accurate deep network learning by exponential linear units (elus)
.

External Links:
1511.07289

Cited by:
§C.1
.

Cochrane (2008)

J. H. Cochrane

The dog that did not bark: a defense of return predictability
.

The Review of Financial Studies

21
(
4
),
pp. 1533–1575
.

Cited by:
§1
.

Cochrane (2011)

J. H. Cochrane

Presidential address: discount rates
.

The Journal of finance

66
(
4
),
pp. 1047–1108
.

Cited by:
§1
,

§5.2
.

Cong
et al.
(2025)

L. W. Cong, G. Feng, J. He, and X. He

Growing the efficient frontier on panel trees
.

Journal of Financial Economics

167
,
pp. 104024
.

Cited by:
§5.2
,

§5.2
.

Daniel and Titman (1997)

K. Daniel and S. Titman

Evidence on the characteristics of cross sectional variation in stock returns
.

the Journal of Finance

52
(
1
),
pp. 1–33
.

Cited by:
§2.1
.

Devlin
et al.
(2019)

J. Devlin, M. Chang, K. Lee, and K. Toutanova

Bert: pre-training of deep bidirectional transformers for language understanding
.

In
Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers)
,

pp. 4171–4186
.

Cited by:
footnote 19
.

Diether
et al.
(2002)

K. B. Diether, C. J. Malloy, and A. Scherbina

Differences of opinion and the cross section of stock returns
.

The journal of finance

57
(
5
),
pp. 2113–2141
.

Cited by:
§1
,

§4.1
.

Driscoll and Kraay (1998)

J. C. Driscoll and A. C. Kraay

Consistent covariance matrix estimation with spatially dependent panel data
.

Review of economics and statistics

80
(
4
),
pp. 549–560
.

Cited by:
§5.1
.

Erichson
et al.
(2020)

N. B. Erichson, P. Zheng, K. Manohar, S. L. Brunton, J. N. Kutz, and A. Y. Aravkin

Sparse principal component analysis via variable projection
.

SIAM Journal on Applied Mathematics

80
(
2
),
pp. 977–1002
.

Cited by:
§3.2
.

Fama and French (1993)

E. F. Fama and K. R. French

Common risk factors in the returns on stocks and bonds
.

Journal of financial economics

33
(
1
),
pp. 3–56
.

Cited by:
§1
.

Fama and French (2015)

E. F. Fama and K. R. French

A five-factor asset pricing model
.

Journal of financial economics

116
(
1
),
pp. 1–22
.

Cited by:
§1
.

Fama and MacBeth (1973)

E. F. Fama and J. D. MacBeth

Risk, return, and equilibrium: empirical tests
.

Journal of political economy

81
(
3
),
pp. 607–636
.

Cited by:
§2.1
.

Fan
et al.
(2016)

J. Fan, Y. Liao, and W. Wang

Projected principal component analysis in factor models
.

Annals of statistics

44
(
1
),
pp. 219
.

Cited by:
§3.2
.

Fang
et al.
(2024)

F. Fang, W. Chung, C. Ventre, M. Basios, L. Kanthan, L. Li, and F. Wu

Ascertaining price formation in cryptocurrency markets with machine learning
.

The European Journal of Finance

30
(
1
),
pp. 78–100
.

Cited by:
§1
.

Feng
et al.
(2020)

G. Feng, S. Giglio, and D. Xiu

Taming the factor zoo: a test of new factors
.

The Journal of Finance

75
(
3
),
pp. 1327–1370
.

Cited by:
footnote 1
.

Feng
et al.
(2018)

G. Feng, J. He, N. G. Polson, and J. Xu

Deep learning in characteristics-sorted factor models
.

Journal of Financial and Quantitative Analysis
,
pp. 1–36
.

Cited by:
§D.3
,

§1
.

Feurer and Hutter (2019)

M. Feurer and F. Hutter

Hyperparameter optimization
.

Automated machine learning: Methods, systems, challenges
,
pp. 3–33
.

Cited by:
§C.3
.

Frazzini
et al.
(2012)

A. Frazzini, R. Israel, and T. J. Moskowitz

Trading costs of asset pricing anomalies
.

Fama-Miller Working Paper, Chicago Booth Research Paper
(
14-05
).

Cited by:
§D.4.3
.

Freyberger
et al.
(2020)

J. Freyberger, A. Neuhierl, and M. Weber

Dissecting characteristics nonparametrically
.

The Review of Financial Studies

33
(
5
),
pp. 2326–2377
.

Cited by:
§B.4
,

§5.2
,

footnote 1
.

Gao
et al.
(2019)

S. Gao, M. Cheng, K. Zhao, X. Zhang, M. Yang, and P. Torr

Res2net: a new multi-scale backbone architecture
.

IEEE transactions on pattern analysis and machine intelligence

43
(
2
),
pp. 652–662
.

Cited by:
footnote 19
.

Gibbons
et al.
(1989)

M. R. Gibbons, S. A. Ross, and J. Shanken

A test of the efficiency of a given portfolio
.

Econometrica: Journal of the Econometric Society
,
pp. 1121–1152
.

Cited by:
§5.2
,

§5
.

Goodfellow
et al.
(2020)

I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio

Generative adversarial networks
.

Communications of the ACM

63
(
11
),
pp. 139–144
.

Cited by:
§1
.

Green
et al.
(2013)

J. Green, J. R. Hand, and X. F. Zhang

The supraview of return predictive signals
.

Review of Accounting Studies

18
,
pp. 692–730
.

Cited by:
footnote 1
.

Green
et al.
(2017)

J. Green, J. R. Hand, and X. F. Zhang

The characteristics that provide independent information about average us monthly stock returns
.

The Review of Financial Studies

30
(
12
),
pp. 4389–4436
.

Cited by:
§B.3
,

footnote 1
.

Gu
et al.
(2020)

S. Gu, B. Kelly, and D. Xiu

Empirical asset pricing via machine learning
.

The Review of Financial Studies

33
(
5
),
pp. 2223–2273
.

Cited by:
§B.2.1
,

§B.3
,

§B.4
,

§C.1
,

§1
,

§2.1
,

§3.1
,

§4.1
,

§4.2.2
,

§5.2
,

footnote 1
.

Gu
et al.
(2021)

S. Gu, B. Kelly, and D. Xiu

Autoencoder asset pricing models
.

Journal of Econometrics

222
(
1
),
pp. 429–450
.

Cited by:
§B.4
,

§D.3
,

§1
,

§3.2
.

Harvey
et al.
(2016)

C. R. Harvey, Y. Liu, and H. Zhu

… And the cross-section of expected returns
.

The Review of Financial Studies

29
(
1
),
pp. 5–68
.

Cited by:
footnote 1
.

He
et al.
(2015)

K. He, X. Zhang, S. Ren, and J. Sun

Delving deep into rectifiers: surpassing human-level performance on imagenet classification
.

In
Proceedings of the IEEE international conference on computer vision
,

pp. 1026–1034
.

Cited by:
§2.3
.

He
et al.
(2017)

Z. He, B. Kelly, and A. Manela

Intermediary asset pricing: new evidence from many asset classes
.

Journal of Financial Economics

126
(
1
),
pp. 1–35
.

Cited by:
footnote 1
.

Hendrycks and Gimpel (2016)

D. Hendrycks and K. Gimpel

Gaussian error linear units (gelus)
.

arXiv preprint arXiv:1606.08415
.

Cited by:
§2.3
.

Higgins
et al.
(2018)

I. Higgins, D. Amos, D. Pfau, S. Racaniere, L. Matthey, D. Rezende, and A. Lerchner

Towards a definition of disentangled representations
.

arXiv preprint arXiv:1812.02230
.

Cited by:
Appendix A
.

Higgins
et al.
(2017)

I. Higgins, L. Matthey, A. Pal, C. P. Burgess, X. Glorot, M. M. Botvinick, S. Mohamed, and A. Lerchner

Beta-vae: learning basic visual concepts with a constrained variational framework.
.

ICLR (Poster)

3
.

Cited by:
footnote 16
.

Hinton and Salakhutdinov (2006)

G. E. Hinton and R. R. Salakhutdinov

Reducing the dimensionality of data with neural networks
.

science

313
(
5786
),
pp. 504–507
.

Cited by:
§D.4.1
,

§D.4.1
.

Hinton
et al.
(2012a)

G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. R. Salakhutdinov

Improving neural networks by preventing co-adaptation of feature detectors
.

External Links:
1207.0580

Cited by:
§C.2
.

Hinton
et al.
(2012b)

G. Hinton, N. Srivastava, and K. Swersky

Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
.

Cited on

14
(
8
),
pp. 2
.

Cited by:
§C.2
.

Hodrick (1992)

R. J. Hodrick

Dividend yields and expected stock returns: alternative procedures for inference and measurement
.

The Review of Financial Studies

5
(
3
),
pp. 357–386
.

Cited by:
§5.1
.

Hou
et al.
(2015)

K. Hou, C. Xue, and L. Zhang

Digesting anomalies: an investment approach
.

The Review of Financial Studies

28
(
3
),
pp. 650–705
.

Cited by:
footnote 1
.

Ioffe and Szegedy (2015)

S. Ioffe and C. Szegedy

Batch normalization: accelerating deep network training by reducing internal covariate shift
.

In
International conference on machine learning
,

pp. 448–456
.

Cited by:
§C.2
.

Jaquart
et al.
(2021)

P. Jaquart, D. Dann, and C. Weinhardt

Short-term bitcoin market prediction via machine learning
.

The journal of finance and data science

7
,
pp. 45–66
.

Cited by:
§1
.

Jegadeesh
et al.
(2004)

N. Jegadeesh, J. Kim, S. D. Krische, and C. M. Lee

Analyzing the analysts: when do recommendations add value?
.

The journal of finance

59
(
3
),
pp. 1083–1124
.

Cited by:
§1
,

§4.1
.

Jensen
et al.
(2023)

T. I. Jensen, B. Kelly, and L. H. Pedersen

Is there a replication crisis in finance?
.

The Journal of Finance

78
(
5
),
pp. 2465–2518
.

Cited by:
footnote 1
.

Kelly
et al.
(2024)

B. Kelly, S. Malamud, and K. Zhou

The virtue of complexity in return prediction
.

The Journal of Finance

79
(
1
),
pp. 459–503
.

Cited by:
§C.3
,

§1
,

§4.2.2
.

Kelly
et al.
(2019)

B. T. Kelly, S. Pruitt, and Y. Su

Characteristics are covariances: a unified model of risk and return
.

Journal of Financial Economics

134
(
3
),
pp. 501–524
.

Cited by:
§B.4
,

§1
.

Kingma and Ba (2017)

D. P. Kingma and J. Ba

Adam: a method for stochastic optimization
.

External Links:
1412.6980

Cited by:
§C.2
.

Koh
et al.
(2020)

P. W. Koh, T. Nguyen, Y. S. Tang, S. Mussmann, E. Pierson, B. Kim, and P. Liang

Concept bottleneck models
.

In
International conference on machine learning
,

pp. 5338–5348
.

Cited by:
§D.2
,

§D.5
,

§1
,

footnote 16
.

Kozak
et al.
(2020)

S. Kozak, S. Nagel, and S. Santosh

Shrinking the cross-section
.

Journal of Financial Economics

135
(
2
),
pp. 271–292
.

Cited by:
§5.2
.

Leippold
et al.
(2022)

M. Leippold, Q. Wang, and W. Zhou

Machine learning in the chinese stock market
.

Journal of Financial Economics

145
(
2
),
pp. 64–82
.

Cited by:
§1
,

§4.1
.

Lim (2001)

T. Lim

Rationality and analysts’ forecast bias
.

The journal of Finance

56
(
1
),
pp. 369–385
.

Cited by:
§1
.

Litterman (1991)

R. Litterman

Common factors affecting bond returns
.

Journal of fixed income
,
pp. 54–61
.

Cited by:
footnote 15
.

Locatello
et al.
(2019a)

F. Locatello, G. Abbati, T. Rainforth, S. Bauer, B. Schölkopf, and O. Bachem

On the fairness of disentangled representations
.

Advances in neural information processing systems

32
.

Cited by:
Appendix A
.

Locatello
et al.
(2019b)

F. Locatello, S. Bauer, M. Lucic, G. Raetsch, S. Gelly, B. Schölkopf, and O. Bachem

Challenging common assumptions in the unsupervised learning of disentangled representations
.

In
international conference on machine learning
,

pp. 4114–4124
.

Cited by:
Appendix A
.

Lovell (1986)

M. C. Lovell

Tests of the rational expectations hypothesis
.

The American Economic Review

76
(
1
),
pp. 110–124
.

Cited by:
§1
.

Ludvigson and Ng (2007)

S. C. Ludvigson and S. Ng

The empirical risk–return relation: a factor analysis approach
.

Journal of financial economics

83
(
1
),
pp. 171–222
.

Cited by:
§3.2
.

McCracken and Ng (2016)

M. W. McCracken and S. Ng

FRED-md: a monthly database for macroeconomic research
.

Journal of Business & Economic Statistics

34
(
4
),
pp. 574–589
.

Cited by:
Table E.3
,

§3.1
.

Merton (1973)

R. C. Merton

An intertemporal capital asset pricing model
.

Econometrica: Journal of the Econometric Society
,
pp. 867–887
.

Cited by:
§1
.

Muth (1961)

J. F. Muth

Rational expectations and the theory of price movements
.

Econometrica: journal of the Econometric Society
,
pp. 315–335
.

Cited by:
§1
.

Nair and Hinton (2010)

V. Nair and G. E. Hinton

Rectified linear units improve restricted boltzmann machines
.

In
Proceedings of the 27th international conference on machine learning (ICML-10)
,

pp. 807–814
.

Cited by:
§2.3
.

Nelson and Siegel (1987)

C. R. Nelson and A. F. Siegel

Parsimonious modeling of yield curves
.

Journal of business
,
pp. 473–489
.

Cited by:
F

[The evaluation harness truncated this reference: showing the first 120000 of 249886 characters.]
</reference>

<statements>
1. Interpretable architectures: Concept-bottleneck or partially interpretable neural networks (e.g. CB-APM) embed economically meaningful bottlenecks (analyst consensus, factors) that preserve structure while delivering state-of-the-art performance.
2. Use interpretable or semi-interpretable ML architectures (e.g. concept-bottleneck models like CB-APM, or structured deep nets with economic constraints) to predict conditional expected returns or risk premia.
3. This yields a traceable decision record that can be explained to stakeholders, addressing the regulatory and practical interpretability concerns around black-box models.
4. Recent work on interpretable deep asset pricing (e.g. CB‑APM) shows that interpretability constraints can sometimes **improve** performance, not just preserve it, by acting as regularizers aligned with economic structure.
</statements>

Begin the assessment now. Output only the JSON list, without any conversational text or explanations.