SearcharxivSearch

arXiv subjects

Kieran Wood

Publications and source records attributed to Kieran Wood.

14 recordsLinked to original sources

Macro-aware time series forecasting via hierarchical mixed-frequency attention models

Deep learning models show promise in financial forecasting, yet their generalization is often undermined by small datasets, noisy signals, and non-stationarity. While meta-learning and related techniques mitigate some of these issues, they typically do not account for a core limitation in macro-financial prediction: the scarcity of distinct macroeconomic regimes that drive asset returns. We introduce HANET (Hierarchical Attention Network), a hybrid LSTM-based architecture that integrates macroeconomic domain knowledge through attention over long-run macro contexts while preserving high-frequency market dynamics. HANET organizes information in a hierarchical mixed-frequency structure, with daily asset-return signals nested within monthly macroeconomic windows, and introduces a Hierarchical Cross-Attention mechanism that reconciles low-frequency macro signals with high-frequency returns without discarding granular daily information. By framing regime selection as attention over macroeconomic contexts, the model adapts to scarce and shifting regimes. Empirically, across 55 liquid futures spanning multiple asset classes, HANET consistently outperforms neural forecasters that ignore macroeconomic information, particularly during turbulent periods, improving risk-adjusted returns and mitigating losses. Ablation studies show that these gains rely on structured macro conditioning rather than naive feature augmentation: an LSTM with the same macro representation performs poorly, and shuffling macro contexts substantially degrades performance. Finally, HANET provides interpretability through attention weights, highlighting which historical regimes are most influential for each forecast and linking macro conditions to portfolio outcomes. These results establish HANET as a systematic approach to integrating macroeconomic information into attention-based deep learning for financial forecasting.

q-fin.ST

DeRegiME: Deep Regime Mixtures for Probabilistic Forecasting under Distribution Shift

We introduce DeRegiME -- Deep Regime Mixture of Experts -- a direct multi-horizon probabilistic forecaster that separates latent uncertainty regimes from the underlying signal and softly assigns each forecast location to learned recurring regimes using a sparse variational Gaussian process (GP) whose nonstationary regime-mixing kernel and Student-t likelihood combine per-regime sub-kernels and noise processes via a shared gate. This yields a single sparse-GP posterior, not a mixture of GP experts. DeRegiME addresses a key limitation of neural forecasters: point forecasts discard residual uncertainty, and probabilistic heads -- whether single marginals, uninterpreted mixtures, quantile sets, or diffusion samples -- rarely expose the regime structure of the residual. Yet distribution shift in noisy heteroskedastic time series may be abrupt, gradual, or horizon-dependent and often appears in residual uncertainty rather than the conditional mean. DeRegiME yields an interpretable mean-residual-noise decomposition with a direct-sum feature-space representation that anchors regimes as clusters of residual similarity whose transitions surface as implicit changepoints. The effective number of regimes is pruned by the stick-breaking gate. We prove kernel validity and predictive-density propriety, and across ten benchmarks and three encoder grids DeRegiME improves negative log predictive density (NLPD) by 20.3% over the strongest encoder-matched baseline, a DeepAR/GluonTS-style dynamic Student-t head, with parallel gains on CRPS (3.0%) and MSE (4.7%). Improvements are consistent across all datasets, which span abrupt, gradual, and seasonal shifts.

cs.LG

Detecting Changes in Causal Dependence with Kernels and Copulas

We propose a framework for determining whether the causal dependence of an outcome $Y$ on a covariate $X$ changes at a given time point, given confounders $\boldsymbol{Z}$. For instance, in financial markets, the effect of a market indicator on asset returns may causally change over time. While many existing measures of association can be used to detect changes in joint and marginal distributions, in the absence of strong assumptions on the data generating process none are suitable for detecting changes in the causal mechanism or in the strength of causal relationship. In this work we approach the problem from a fully non-parametric perspective, and treat the causal mechanism as well as the distribution of the data as unknown. We introduce a quantity based on the integrated difference between kernel mean embeddings of certain conditionals copula, which is provably equal to zero if the causal dependence does not change and strictly positive else. A near-linear time estimator for the quantity is proposed, with rates of convergence explicitly spelled out. Extensive experiments demonstrate that the proposed statistic achieves high accuracy on multiple synthetic and real-world datasets. We additionally show how the proposed statistic can be used for change point detection when the goal is to detect changes in causal dependence occurring at an unknown times.

stat.ME

Deep Learning for Financial Time Series: A Large-Scale Benchmark of Risk-Adjusted Performance

We present a large scale benchmark of modern deep learning architectures for a financial time series prediction and position sizing task, with a primary focus on Sharpe ratio optimization. Evaluating linear models, recurrent networks, transformer based architectures, state space models, and recent sequence representation approaches, we assess out of sample performance on a daily futures dataset spanning commodities, equity indices, bonds, and FX spanning 2010 to 2025. Our evaluation goes beyond average returns and includes statistical significance, downside and tail risk measures, breakeven transaction cost analysis, robustness to random seed selection, and computational efficiency. We find that models explicitly designed to learn rich temporal representations consistently outperform linear benchmarks and generic deep learning models, which often lead the ranking in standard time series benchmarks. Hybrid models such as VSN with LSTM, a combination of Variable Selection Networks (VSN) and LSTMs, achieves the highest overall Sharpe ratio, while VSN with xLSTM and LSTM with PatchTST exhibit superior downside adjusted characteristics. xLSTM demonstrates the largest breakeven transaction cost buffer, indicating improved robustness to trading frictions.

q-fin.TR

DeePM: Regime-Robust Deep Learning for Systematic Macro Portfolio Management

We propose DeePM (Deep Portfolio Manager), a structured deep-learning macro portfolio manager trained end-to-end to maximize a robust, risk-adjusted utility. DeePM addresses three fundamental challenges in financial learning: (1) it resolves the asynchronous "ragged filtration" problem via a Directed Delay (Causal Sieve) mechanism that prioritizes causal impulse-response learning over information freshness; (2) it combats low signal-to-noise ratios via a Macroeconomic Graph Prior, regularizing cross-asset dependence according to economic first principles; and (3) it optimizes a distributionally robust objective where a smooth worst-window penalty serves as a differentiable proxy for Entropic Value-at-Risk (EVaR) - a window-robust utility encouraging strong performance in the most adverse historical subperiods. In large-scale backtests from 2010-2025 on 50 diversified futures with highly realistic transaction costs, DeePM attains net risk-adjusted returns that are roughly twice those of classical trend-following strategies and passive benchmarks, solely using daily closing prices. Furthermore, DeePM improves upon the state-of-the-art Momentum Transformer architecture by roughly fifty percent. The model demonstrates structural resilience across the 2010s "CTA (Commodity Trading Advisor) Winter" and the post-2020 volatility regime shift, maintaining consistent performance through the pandemic, inflation shocks, and the subsequent higher-for-longer environment. Ablation studies confirm that strictly lagged cross-sectional attention, graph prior, principled treatment of transaction costs, and robust minimax optimization are the primary drivers of this generalization capability.

q-fin.TR

New formalism for perturbations of massive gravity theories around arbitrary background spacetimes

We develop a new technique for studying the perturbations of dRGT-type massive gravity theories around arbitrary background spacetimes. Built initially from the vielbein formulation of the theory, but switching back to the metric formulation afterwards, our approach bypasses many of the complications that arise in previous metric formulation approaches to linearising massive gravity around generic backgrounds, naturally elucidates the ghost-free structure of the interactions, and readily generalises to higher orders in perturbation theory, as well as to multiple interacting metric tensor fields. To demonstrate the power of our technique, we apply our formalism to a number of commonly occurring example backgrounds - proportional, cosmological, and black hole - recovering and extending many known results from the literature at linear order. Lastly, we provide, for the first time, the cubic order multi-gravity potential around a generic background spacetime.

hep-th

A new look at multi-gravity and dimensional deconstruction

It has long been understood that certain theories of ghost free massive gravity and their multi-graviton extensions can be thought of as arising from a higher dimensional theory of gravity, upon discretising the extra dimension. However, this correspondence between standard multi-gravity and extra dimensional gravity holds only when one discretises the extra dimension after gauge fixing the lapse function associated to the various lower dimensional hypersurfaces. The lapse provides crucial structure to the extra dimensional theory: in pure general relativity (GR), it ensures full diffeomorphism invariance of the theory, and enforces its Hamiltonian constraint. Thus, upon deconstruction, important information related to the extra dimension is missing in the resulting multi-gravity theory; as a result one could never hope to recover higher dimensional GR in its entirety upon taking the appropriate continuum limit. Here, we develop an improved deconstruction procedure that maintains the free lapse, and show that the resulting deconstructed theory is essentially multi-gravity equipped with additional dynamical scalar fields, whose field equations encode the Hamiltonian constraint in the extra dimensional theory. As an example, we explicitly demonstrate that - with an FLRW ansatz for the metrics in this new theory - one may recover all of the equations and constraints of 5-dimensional brane cosmology upon taking the continuum limit. We then treat the deconstructed theory as an entity in its own right, and generalise it to arbitrary dimension and interaction structures beyond those admitting a well-defined continuum limit. We dub this theory `scalar-tensor multi-gravity', and show that the new scalar equations change the structure of some simple solutions that were previously allowed in standard multi-gravity, in a manner that exactly mirrors what we expect from higher dimensional GR.

hep-th

Black Holes in Multi-Metric Gravity II: Hairy Solutions and Linear Stability of the Non- and Partially Proportional Branches

Owing to our work in part I of this series of papers, it is understood that the analytically known black hole solutions in the theory of ghost free multi-metric gravity can be split into three distinct classes, and that one of these classes - the proportional branch - exhibits the Gregory-Laflamme instability at linear level in the metric perturbations, whenever the black hole horizon size is smaller than (roughly) the Compton wavelength of the theory's lightest massive graviton. In this first of two sequels, we determine the linear stability of the two remaining classes of black hole solutions - the non-proportional and partially proportional branches - and discuss how our results likely differ at nonlinear level. We also give a general prescription to construct multi-metric solutions describing black holes endowed with massive graviton hair, which may constitute the end state of the instability in the proportional branch. We utilise a tractable example model involving 3 metrics to see how this works in practice, and determine the asymptotic form of its corresponding hairy solutions at infinity, where one can clearly see the individual contributions from each of the graviton mass modes.

gr-qc

Black Holes in Multi-Metric Gravity

We construct a wide class of black hole solutions to the general theory of ghost free multi-metric gravity in arbitrary spacetime dimension, extending and generalising the known results in 4-dimensional dRGT massive gravity and bigravity. The solutions are split into three generic classes based on whether the metrics can be simultaneously diagonalised - one of which does not exist in dRGT massive gravity nor bigravity, and is only possible when one has more than two interacting metric fields. We also linearise the general multi-metric theory to determine the dynamics of the massive spin-2 modes, including examples where this can be done analytically, and use the linear theory to discuss the stability of the 4-dimensional multi-Schwarzchild and multi-Kerr solutions. We explain how the instabilities that plague these solutions in dRGT massive gravity and bigravity carry across to the general multi-metric theory, touching upon ideas of dimensional deconstruction to make sense of the results.

gr-qc

Few-Shot Learning Patterns in Financial Time-Series for Trend-Following Strategies

Forecasting models for systematic trading strategies do not adapt quickly when financial market conditions rapidly change, as was seen in the advent of the COVID-19 pandemic in 2020, causing many forecasting models to take loss-making positions. To deal with such situations, we propose a novel time-series trend-following forecaster that can quickly adapt to new market conditions, referred to as regimes. We leverage recent developments from the deep learning community and use few-shot learning. We propose the Cross Attentive Time-Series Trend Network -- X-Trend -- which takes positions attending over a context set of financial time-series regimes. X-Trend transfers trends from similar patterns in the context set to make forecasts, then subsequently takes positions for a new distinct target regime. By quickly adapting to new financial regimes, X-Trend increases Sharpe ratio by 18.9% over a neural forecaster and 10-fold over a conventional Time-series Momentum strategy during the turbulent market period from 2018 to 2023. Our strategy recovers twice as quickly from the COVID-19 drawdown compared to the neural-forecaster. X-Trend can also take zero-shot positions on novel unseen financial assets obtaining a 5-fold Sharpe ratio increase versus a neural time-series trend forecaster over the same period. Furthermore, the cross-attention mechanism allows us to interpret the relationship between forecasts and patterns in the context set.

q-fin.TR

Clockwork Cosmology

The higher order generalisation of the clockwork mechanism to gravitational interactions provides a means to generate an exponentially suppressed coupling to matter from a fundamental theory of multiple interacting gravitons, without introducing large hierarchies in the underlying potential and without the need for a dilaton, suggesting a possible application to the hierarchy problem. We work in the framework of ghost free multi-gravity with "nearest-neighbour" interactions, and present a formalism by which one is able to construct potentials such that the theory will always exhibit this clockwork effect. We also consider cosmological solutions to the general theory, where all metrics are of FRW form, with site-dependent scale factors/lapses. We demonstrate the existence of multiple deSitter vacua where all metrics share the same Hubble parameter, and we solve the modified Einstein equations numerically for an example clockwork model constructed using our formalism, finding that the evolution of the metric that matter couples to is essentially equivalent to that of general relativity at the modified Planck scale. It is important to stress that while we focus on the application to clockwork theories, our work is entirely general and facilitates finding cosmological solutions to any ghost free multi-gravity theory with "nearest-neighbour" interactions. Moreover, we clarify previous work on the continuum limit of the theory, which is generically a scalar-tensor braneworld, using the Randall-Sundrum model as a special case and showing how the discrete-clockwork cosmological results map to the continuum results in the appropriate limit.

hep-th

Trading with the Momentum Transformer: An Intelligent and Interpretable Architecture

We introduce the Momentum Transformer, an attention-based deep-learning architecture, which outperforms benchmark time-series momentum and mean-reversion trading strategies. Unlike state-of-the-art Long Short-Term Memory (LSTM) architectures, which are sequential in nature and tailored to local processing, an attention mechanism provides our architecture with a direct connection to all previous time-steps. Our architecture, an attention-LSTM hybrid, enables us to learn longer-term dependencies, improves performance when considering returns net of transaction costs and naturally adapts to new market regimes, such as during the SARS-CoV-2 crisis. Via the introduction of multiple attention heads, we can capture concurrent regimes, or temporal dynamics, which are occurring at different timescales. The Momentum Transformer is inherently interpretable, providing us with greater insights into our deep-learning momentum trading strategy, including the importance of different factors over time and the past time-steps which are of the greatest significance to the model.

cs.LG

Slow Momentum with Fast Reversion: A Trading Strategy Using Deep Learning and Changepoint Detection

Momentum strategies are an important part of alternative investments and are at the heart of commodity trading advisors (CTAs). These strategies have, however, been found to have difficulties adjusting to rapid changes in market conditions, such as during the 2020 market crash. In particular, immediately after momentum turning points, where a trend reverses from an uptrend (downtrend) to a downtrend (uptrend), time-series momentum (TSMOM) strategies are prone to making bad bets. To improve the response to regime change, we introduce a novel approach, where we insert an online changepoint detection (CPD) module into a Deep Momentum Network (DMN) [1904.04912] pipeline, which uses an LSTM deep-learning architecture to simultaneously learn both trend estimation and position sizing. Furthermore, our model is able to optimise the way in which it balances 1) a slow momentum strategy which exploits persisting trends, but does not overreact to localised price moves, and 2) a fast mean-reversion strategy regime by quickly flipping its position, then swapping it back again to exploit localised price moves. Our CPD module outputs a changepoint location and severity score, allowing our model to learn to respond to varying degrees of disequilibrium, or smaller and more localised changepoints, in a data driven manner. Back-testing our model over the period 1995-2020, the addition of the CPD module leads to an improvement in Sharpe ratio of one-third. The module is especially beneficial in periods of significant nonstationarity, and in particular, over the most recent years tested (2015-2020) the performance boost is approximately two-thirds. This is interesting as traditional momentum strategies have been underperforming in this period.

stat.ML

Josephson Photonics with Simultaneous Resonances

Inelastic Cooper pair tunneling across a voltage-biased Josephson junction in series with one or more microwave cavities can generate photons via resonant processes in which the energy lost by the Cooper pair matches that of the photon(s) produced. We generalise previous theoretical treatments of such systems to analyse cases where two or more different photon generation processes are resonant simultaneously. We also explore in detail a specific case where generation of a single photon in one cavity mode is simultaneously resonant with the generation of two photons in a second mode. We find that the coexistence of the two resonances leads to effective couplings between the modes which in turn generate entanglement.

cond-mat.mes-hall