SearcharxivSearch

arXiv · 2508.19006

Is attention truly all we need? An empirical study of asset pricing in pretrained RNN sparse and global attention models

Abstract

This study investigates the pre-trained RNN attention models with the mainstream attention mechanisms, such as additive attention, Luong's three attentions, global self-attention and sliding window sparse attention, for the empirical asset pricing research on the top 420 large-cap US stocks. This is the first paper on the large-scale state-of-the-art (SOTA) attention mechanisms applied in the asset pricing context. They overcome the limitations of the traditional machine learning-based asset pricing, such as mis-capturing the temporal dependency and short memory. Moreover, the enforced causal masks in the attention mechanisms address the future data leaking issue ignored by the more advanced attention-based models, such as the classic Transformer. The proposed attention models also consider the temporal sparsity characteristic of asset pricing data and mitigate potential overfitting issues by deploying the simplified model structures. This provides some insights for future empirical economic research. All models are examined in three periods, which cover pre-COVID-19, COVID-19 and one year post-COVID-19, for testing the stability of these models under extreme market conditions. The study finds that in value-weighted portfolio back testing, the global self-attention model and the sliding window sparse attention model exhibit excellent capabilities in deriving the absolute returns and hedging downside risks, while they achieve an annualized Sortino ratio of 2.0 and 1.80 respectively in the period with COVID-19 in the static transaction cost scenario. Moreover, the sliding window sparse attention model performs more stably than the global self-attention model from the perspective of absolute portfolio returns with respect to the size of stocks' market capitalization.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Shanyan Lai. 2025-08-26. Is attention truly all we need? An empirical study of asset pricing in pretrained RNN sparse and global attention models. https://arxiv.org/abs/2508.19006

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Pricing and Hedging of Discretely Monitored Asian Options in the Volterra-Heston Model

We develop semi-closed pricing formulas and lifted-model hedging methods for discretely monitored geometric and arithmetic Asian options in the Volterra-Heston stochastic volatility model. Exploiting the affine Volterra structure, we derive a tractable transform for the joint law of the terminal log-price and the discretely monitored geometric average. This transform yields semi-closed pricing formulas for geometric Asian options, which in turn provide effective control variates for Monte Carlo valuation of arithmetic Asian options. Under the stated real-moment and affine-transform hypotheses, we also derive the Galtchouk-Kunita-Watanabe decomposition for Fourier-representable payoffs and obtain a variance-optimal hedge in terms of the Riccati-Volterra equation and the forward-variance curve. Using N-factor Markovian approximations, we obtain a finite-dimensional numerical implementation for hedging Asian options. Our numerical experiments document factor convergence for a regular non-Markovian kernel and the effect of rebalancing frequency on hedging error. In the Heston benchmark, geometric Asian controls substantially reduce the variance of arithmetic-Asian price estimates and improve the finite-sample stability of regression-based hedging relative to direct regression.

q-fin.PR

Beyond Lognormal Sums: A Four-Moment Probability Framework for Basket and Spread Option Pricing

Basket options are difficult to value under correlated lognormal dynamics because weighted sums and differences of lognormal variables have no tractable distribution. This paper develops a probability-based four-moment framework that separates the exact pricing representation from the distributional approximation. A change of measure first writes a basket price as a linear combination of probabilities. For a standard basket with one positive weight, these probabilities become CDF values of positive correlated lognormal sums. Each sum is approximated by a shifted lognormal variance mixture matched to its first four moments. For an unrestricted mixed-sign basket, a signed shifted lognormal proxy gives an analytical call-price formula. We state admissibility conditions, provide a practical root-selection rule, establish the main strike-based financial properties of the direct proxy, and derive exact pricing-error identities in terms of cumulative distribution function (CDF) discrepancies. The numerical analysis combines standard-basket benchmarks with an empirical application to a normalized $3{:}2{:}1$ crack spread constructed from RBOB gasoline, ULSD or heating oil, and WTI futures. The results show that the probability reformulation and the fourth-moment condition improve the distributional fit and pricing accuracy, particularly when maturity and tail asymmetry increase. The framework remains analytical, transparent, and suitable for repeated valuation across strikes and maturities.

q-fin.PR

When to Sell an Asset? - A Distribution Builder Approach

We consider the question of the optimal timing of the sale of an asset with stochastic dynamics. Our analysis is based on the method of the distribution builder introduced by Sharpe, Goldstein and Blythe [SGB00] for the purpose of optimal portfolio selection. Instead of specifying a utility function or risk aversion coefficient, this tool directly elicits the target distribution of the investor. We show how the problem of an optimal asset sale is in this setting linked to the problem of finding a Skorokhod embedding of a distribution into a diffusion process. In the case where the asset process follows a geometric Brownian motion and a specific family of distributions is targeted, one can observe a risk-return tradeoff.

q-fin.PR