Searcharxiv⌕ Search

arXiv · 2610.00589

Foundation or Formula? A Simulation-Based Comparison of Google's TimesFM-3 and Classical Time-Series Models

Abstract

Time-series foundation models such as Google's TimesFM-3 forecast series they were never trained on, but whether they should replace classical methods such as ARIMA and exponential smoothing is hard to settle on public benchmarks, because few benchmark datasets are absent from every model's pre-training corpus. This exploratory study compares TimesFM-3 with classical methods on data generated from nine known processes, which the model cannot have seen. Across 7200 series and a pre-registered protocol with a 12-step horizon, neither family dominates. TimesFM-3's error is never more than 1.31 times that of the best method in a scenario, whereas every classical method's is at least 2.2 times somewhere; the largest classical failures occur with only two seasonal cycles, where the automatic methods fall back to non-seasonal models. With four or more cycles automatic ARIMA is 21-24% more accurate than TimesFM-3 on seasonal ARIMA data. Among five foundation models this robustness is specific to TimesFM-3, and it holds only at short horizons: at steps 25 to 48 the three foundation models tested there forecast a decline on a saturating process whose level stays flat, and automatic ARIMA has the smaller worst case elsewhere. On intermittent demand TimesFM-3 shows no advantage over simple benchmarks built from the history. Randomised parameters, heavy-tailed noise and outliers leave these patterns qualitatively unchanged; on the official test period of 1000 M4 monthly series TimesFM-3 is level with the fifth- to seventh-ranked M4 entries, behind the four best, and on 101 macroeconomic series observed after the models' documented training data it is statistically level with the classical methods. The study characterises the behaviour of a black-box model, not the reasons for it, and its findings are conditional on the processes and horizons examined.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ebrahim Khaled Ebrahim, Somaia Mohamed Ali, Ahmed El-Kotory. 2026-09-30. Foundation or Formula? A Simulation-Based Comparison of Google's TimesFM-3 and Classical Time-Series Models. https://arxiv.org/abs/2610.00589

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Projection-based significance test for longitudinally collected functional data under a crossover design

Wearable devices for continuous electronic health monitoring often capture data at frequent intervals under a dense functional design. The focal point is the analysis of longitudinal functional data, wherein functional trajectories are observed repeatedly over time. This work is motivated by the interest in assessing the efficacy of a noninflammatory medication, meloxicam, on the daily activity levels of household cats with a pre-existing condition of osteoarthritis under a crossover design. An accelerometer records these activity profiles at a minute level over the entire study period. To this aspect, we propose an orthogonal projection-based pseudo-generalized F test to determine the significance of the functional treatment effect under a functional additive crossover model while accounting for carryover effects and baseline covariates. Under mild conditions, we derive the asymptotic null distribution of the test statistic and the theoretical power function when the projection function is estimated from the data. Numerical studies demonstrate that the proposed test maintains size, proves powerful in detecting the smooth effect of meloxicam, and exhibits high efficiency compared to bootstrap-based alternatives. Application of the test on the accelerometric activity profiles reveals a strong positive effect of meloxicam on the joint pain score of the cats after adjusting for other baseline covariates.

stat.ME↗

Statistical Methods for Crossover Trials: A Review of Classical and Recent Developments

A comprehensive review of the literature on crossover design is needed to highlight its evolution, applications, and methodological advancements across various fields. Given its widespread use in clinical trials and other research domains, understanding this design's challenges, assumptions, and innovations is essential for optimizing its implementation and ensuring accurate, unbiased results. This article extensively reviews the history and statistical inference methods for crossover designs. A primary focus is given to the AB-BA design as it is the most widely used design in literature. Extension from two periods to higher-order designs is discussed, and a general inference procedure for continuous response is studied. Analysis of multivariate and categorical responses is also reviewed in this context. Recent developments, including causal (potential outcomes) formulations of carryover, efficient use of period-specific baselines, generalized estimating equations for repeated measurements within periods, and estimands for incomplete crossover trials, are also discussed. Several open problems in this area are shortlisted.

stat.ME↗

A robust regression approach to synthetic control with interference

Synthetic control methods are widely used for policy evaluation, but most existing approaches rule out interference among units, compromising validity when such effects are present. We develop a framework that accommodates contaminated donor pools and unknown interference patterns through two stages: factor-model adjustment for unobserved confounding, followed by robust regression in which direct and interference effects appear as a sparse outlier component. We study two asymptotic regimes. When the number of units is fixed and at least half are unaffected by interference, high-breakdown robust regression yields consistent identification of valid controls and asymptotically normal inference. When the number of units diverges, we allow for sparse large and dense weak interference, with robust M-estimation remaining valid even when the post-intervention period is short. Unlike existing approaches requiring prespecification of valid controls or parametric modeling of interference, our framework relies only on coarse sparsity information and enables formal inference on both direct and interference effects. We assess the proposed methods through simulations and two empirical applications. An analysis of the US embassy relocation to Jerusalem reveals significant interference effects on conflict outcomes in Jordan, and an analysis of Beijing's air pollution policy uncovers spatial interference patterns consistent with prevailing wind directions.

stat.ME↗