SearcharxivSearch

arXiv · 2606.27100

Pretrained Time-Series Foundation Models for Financial Return Forecasting

Abstract

Financial return forecasting is a difficult test case for time-series foundation models (TSFMs) due to low signal-to-noise ratios, structural breaks, heavy tails, and weak persistence. This paper benchmarks pretrained TSFMs against train-from-scratch neural baselines in a deliberately conservative financial setting. We evaluate TimeGPT/TimeGPT-LH, TimesFM-2.5, Moirai-2.0, Chronos, and Chronos-2 against NBEATS, NHITS, PatchTST, iTransformer, and KAN on five liquid U.S. equities (AAPL, AMZN, GOOG, JPM, META) using linear and log returns. Models are compared under an equalized context budget, a rolling-origin protocol, and against random-walk benchmarks. We provide a theoretical framing of pretraining as an inductive prior, linking PAC-Bayes transfer intuition, information-theoretic predictability limits, and attention geometry. This clarifies why strong model rankings need not imply economically meaningful predictability in noisy markets. Pragmatically, pretrained TSFMs dominate the ranking distribution, accounting for 8 of 10 task-level wins. Moirai-2.0 and TimesFM-2.5 achieve the strongest average ranks, leading tasks for AAPL, JPM, GOOG, and AMZN, while Chronos wins the remaining AMZN task. However, the iTransformer baseline wins both META tasks, showing local supervised learning can still outperform generic pretraining for specific assets. Crucially, gains over the random-walk benchmark are small and sparse. A one-sided Diebold-Mariano test rejects equal or inferior predictive accuracy only for Chronos on AMZN and Moirai-2.0 on GOOG. We conclude that TSFMs serve as useful practical priors that reduce model-development costs in low-data financial forecasting, but are not universal engines for statistically reliable alpha generation in realistic empirical deployment.

Explore related subjects

Keep this discovery

BibTeXRIS

Miquel Noguer I Alonso, Rodolfo Pereira Franklin. 2026-06-25. Pretrained Time-Series Foundation Models for Financial Return Forecasting. https://arxiv.org/abs/2606.27100

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Variance-Optimal Hedging in the Rough Hawkes--Heston Model

We study variance-optimal stock hedging and the convergence of approximate strategies in the rough Hawkes--Heston model. Starting from the model's affine conditional transform and the affine Volterra jump framework, we obtain semi-explicit hedges for European calls and a representation of the minimum quadratic error through the Galtchouk--Kunita--Watanabe projection. Our main approximation result keeps the original stock, variance driver, and information flow fixed while regularizing the kernel used to evaluate the hedge. To handle singular memory and common marked jumps, we construct the approximate holdings from histories available before trading and preserve the conditional transform's random modulus envelope. Riccati--Volterra stability and weighted truncation then yield convergence in the original stock's trading norm on compact Fourier intervals. For calls, a joint choice of kernel regularization and Fourier cutoff gives convergence of the initial capitals and strategies, uniform-in-time square-mean convergence of continuous-time gains, and convergence of the terminal mean-square error to the variance-optimal value. A numerical experiment with shifted fractional kernels illustrates the construction on common original-market paths.

q-fin.MF

Numeraire Invariance of Entropy-Projected Martingale Measures

Let \(P\) be a fixed physical law and let \(Q\) be an equivalent martingale measure selected from the martingale-measure set associated with a chosen numeraire. A change of numeraire maps \(Q\) to \(T_LQ\), where \(d(T_LQ)=L\,dQ\) and \(L\) is the terminal likelihood ratio. The forward relative-entropy projection minimizing \(D_{\mathrm{KL}}(P\Vert Q)\) commutes with this transform because its objective changes only by the constant \(-E_P\log L\). The minimal entropy martingale measure (MEMM) orientation \(D_{\mathrm{KL}}(Q\Vert P)\) does not have this property, and a trinomial counterexample shows that independently recomputed MEMMs need not be likelihood compatible. We make two economic consequences explicit. First, the two entropy orientations are precisely the \(Q\)-dependent terms in the classical convex-dual objectives for logarithmic and exponential utility, respectively. Second, likelihood compatibility is equivalent to equality of the pricing functionals obtained in the two numeraires. Hence the forward selectors value every integrable claim consistently across numeraires, whereas the two MEMMs in the counterexample assign different prices to a nonreplicable digital claim. We also prove a finite-state class-level characterization: uniform invariance over the elementary one-period likelihood-ratio families forces a smooth convex \(f\)-divergence to be logarithmic, up to scaling and affine equivalence. Finally, in finite-state markets, the forward projection exists under the usual strictly positive feasible-point condition; its density \(dP/dQ^*\) is attainable log-optimal terminal wealth, and the minimum forward entropy equals maximal expected log growth.

q-fin.MF

The Delta of a Variance Swap

We define the variance swap delta as the sensitivity of the price of variance to a change in underlying price. We use Carr-Madan spanning formulas to analyze this sensitivity when the implied volatility smile curve may depend on the underlying price. We show that the variance swap total delta is zero for the class of smile curves that are pure functions of (log) moneyness, which goes against the empirical observation that variance is up when the market is down. We propose a simple modification of the smile to correct this issue.

q-fin.MF