SearcharxivSearch

arXiv · 2511.08608

When Reasoning Fails: Evaluating 'Thinking' LLMs for Stock Prediction

Abstract

Problem. "Thinking" LLMs (TLLMs) expose explicit or hidden reasoning traces and are widely believed to generalize better on complex tasks than direct LLMs. Whether this promise carries to noisy, heavy-tailed and regime-switching financial data remains unclear. Approach. Using Indian equities (NIFTY constituents), we run a rolling 48m/1m walk-forward evaluation at horizon k = 1 day and dial cross-sectional complexity via the universe size U in {5, 11, 21, 36} while keeping the reasoning budget fixed (B = 512 tokens) for the TLLM. We compare a direct LLM (gpt-4o-mini), a TLLM (gpt-5), and classical learners (ridge, random forest) on cross-sectional ranking loss 1 - IC, MSE, and long/short backtests with realistic costs. Statistical confidence is measured with Diebold-Mariano, Pesaran-Timmermann, and SPA tests. Main findings. (i) As U grows under a fixed budget B, the TLLM's ranking quality deteriorates, whereas the direct LLM remains flat and classical baselines are stable. (ii) TLLM variance is higher, requiring ex-post calibration (winsorization and blending) for stability. (iii) Portfolio results under transaction costs do not support a net advantage for the TLLM. Hypotheses. Our results are consistent with the following testable hypotheses: H1 (Capacity-Complexity Mismatch): for fixed B, TLLM accuracy degrades superlinearly in cross-sectional complexity. H2 (Reasoning Variance): TLLM outputs exhibit higher dispersion date-by-date than direct LLMs, increasing error bars and turnover. H3 (Domain Misfit): next-token prediction objectives and token-budgeted inference are poorly aligned with heavy-tailed, weakly predictable stock returns. Implication. In our setting, "thinking" LLMs are not yet ready to replace classical or direct methods for short-horizon stock ranking; scaling the reasoning budget and/or re-aligning objectives appears necessary.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rakeshkumar H Sodha. 2025-11-05. When Reasoning Fails: Evaluating 'Thinking' LLMs for Stock Prediction. https://arxiv.org/abs/2511.08608

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Log S-fBM model: Statistical analysis

The Log S-fBM model, introduced by Wu et al., is a stochastic volatility model whose log volatility is a stationary fractional Brownian motion (S-fBM): a stationary Gaussian process with power-decaying autocovariance driven by the Hurst exponent $H$, and variance scaled by an intermittency coefficient. A key property is that it reconciles rough volatility, where $H$ is typically near $0.1$ (see Gatheral et al.), with multifractal volatility, where $H$ is close to $0$ as in Bacry, Muzy et al.: the model's volatility measure converges to a multifractal random measure as $H\to0$. Numerical findings in Wu et al. show intermittency of order $0.02$ across financial assets, motivating a small intermittency approximation of log volatility moments for calibration via the general method of moments (GMM). In this work, we conduct a statistical analysis of the Log S-fBM model. We derive scaling properties of the S-fBM process and the Log S-fBM integrated volatility measure, present deviation inequalities with tail distributions sensitive to $H$ and intermittency, and develop a hypothesis test for the null Hurst exponent, i.e.\ rough versus multifractal dynamics. Finally, we revisit scale invariance of the log volatility increment process via explicit small-intermittency formulas, reproducing analogous properties in both regimes.

q-fin.ST

Asymmetric Long-Memory GARCH: Sign-Dependent Kernel Injection in a Two-Dimensional Markov Chain

We introduce ALM-GARCH, an asymmetric long-memory GARCH model in which positive and negative innovations enter conditional variance with different injection amplitudes and kernel offsets. These departures define testable level and memory channels relative to a nested symmetric benchmark. Positive Harris recurrence holds for interior configurations under a Foster-Lyapunov condition. Across five equity indices and Bitcoin, joint symmetry is rejected throughout, driven primarily by the level channel. The memory channel is supported for the Nikkei 225, KOSPI, and Bitcoin but is weakly identified when the positive branch is nearly inactive. Out-of-sample performance is broadly comparable to standard benchmarks.

q-fin.ST

Modeling Trade Durations under Temporal Granularity Effects in Forex Markets

Trade durations in high-frequency foreign exchange data exhibit increased occurrence near integer values. To address this empirical phenomenon, we propose the granularity-adjusted autoregressive conditional duration (GA-ACD) model. It is based on a novel two-component mixture distribution consisting of a standard generalized gamma component for regular durations and a second component that locally redistributes probability mass around integer values to capture heaping. Conditional dynamics are modeled within a score-driven framework, allowing the scale parameter to vary over time in response to past durations, and enabling maximum likelihood estimation of all model parameters. A simulation study shows that ignoring heaping leads to biased parameter estimates and distorted inference regarding both the distribution and the dynamics of durations. An empirical analysis demonstrates that integer-duration clustering is pervasive across major currency pairs and that the GA-ACD model outperforms the standard generalized gamma ACD model.

q-fin.ST