SearcharxivSearch

arXiv subjects

Terry Lyons

Publications and source records attributed to Terry Lyons.

At least 19 recordsLinked to original sources

Concise $(\varepsilon,r)$-representations of a path

Paths $X \colon [0,T] \to \mathbb R^d$ are traditionally stored in finite memory as time series. Recent research has underscored the benefits of instead representing them as collections of iterated integrals $\{\int_{0 < u_1 < \ldots < u_n < T} \mathrm{d} X_{u_1} \otimes \cdots \otimes \mathrm{d} X_{u_n}\}_{n = 0}^N$. These two encodings can be viewed as the extrema on a two-parameter spectrum of representations of the path as degree-$N$ signatures on $m$ intervals in a partition of $[0,T]$. We ask the question of which such representation takes up the least amount of memory, measured as number of real values needed to store the truncated log-signature, subject to the constraint of it being able to approximate solutions to linear controlled differential equations (CDEs) $\mathrm{d} Y = AY \mathrm{d} X$ with $|A| \leq r$ at accuracy at least $\varepsilon$. Estimating the error in terms of the length of $X$, we find that the optimal representation generally lies strictly in between the two naive choices $N = 1$ or $m = 1$, and derive its asymptotics as $r \to \infty$ and $\varepsilon \to 0^+$. Similar considerations can be made when estimating the error in terms of the $p$-variation norm of $X$: in this regime we prove an error bound of the degree-$N$ Euler scheme for linear CDEs with decay in both $m$ and (factorially) in $N$ with the other arbitrarily fixed. We conclude by setting up the analogous problem for SDEs, with the error measured in $L^2$, and derive a similar $L^2$-Euler error estimate for It\^o SDEs with drift. We include an empirical study of the optimisation problem, which we demonstrate for toy examples of $p$-rough paths and for fractional Brownian motion.

math.NA

Chess-World-Model: A 10M-Game Benchmark for Exact State Tracking from Chess Move Sequences

World models require state tracking, which is the ability to maintain a correct latent state across action sequences. Existing benchmarks are often synthetic or language-based, limiting their value as tests of structured state updates in realistic domains. We introduce Chess-World-Model, a large-scale state-tracking benchmark built from 10 million real chess games, where models predict the exact board state reached after a sequence of legal moves. Alongside a held-out real-game split, we include an out-of-distribution split from uniformly random legal play, which tests whether models learn the transition rules rather than shortcuts from common human positions. Prior theoretical and empirical work has shown that Transformers struggle to state-track, while input-dependent linear RNNs require expressive state-transition matrices to do so. We therefore benchmark a causal Transformer, block-diagonal SLiCE, Mamba-3, and Gated DeltaNet with negative eigenvalues under a matched interface and training protocol. The recurrent models strongly outperform the Transformer at 3 and 8 million parameters. Real-game performance saturates above 18 million parameters, but the random-uniform split remains discriminative up to 40 million, exposing failures otherwise hidden by scale. Additionally, ablations show that less expressive state-transition mechanisms reduce performance on the out-of-distribution split for all three recurrent models. Together, these results establish Chess-World-Model as a practical large-scale benchmark for state tracking that exposes failures model scale would otherwise conceal.

cs.LG

Faithful Embeddings of Irregular and Asynchronous Data for Online Log-NCDEs

Continuous-time models are a natural choice for irregular and asynchronous data. A central design choice is how to embed discrete observations into continuous time. Interpolation- and imputation-based embeddings reconstruct a continuous observation path, making the model sensitive to the choice of reconstruction. We show that this reconstruction step is unnecessary; under mild conditions, compact-set universality on the model input space transfers to the data space whenever the embedding from data to input is continuous and injective. Guided by this result, and building on the rectilinear control path for Neural Controlled Differential Equations (NCDEs), we introduce a continuous and injective embedding for Log-NCDEs, a universal class of continuous-time models. Our approach records observations as increments and composes them over arbitrary query intervals to directly form log-signatures. This provides interval-level summaries without first interpolating the observed variables, while supporting online computation. Experiments on synthetic controlled dynamics and real-world time-series datasets show that the representation is accurate, efficient, and robust to irregular, asynchronous, and sparse observations.

cs.LG

The Exponentially Weighted Signature

The signature is a canonical representation of a multidimensional path over an interval. However, it treats all historical information uniformly, offering no intrinsic mechanism for contextualising the relevance of the past. To address this, we introduce the Exponentially Weighted Signature (EWS), generalising the Exponentially Fading Memory (EFM) signature from diagonal to general bounded linear operators. These operators enable cross-channel coupling at the level of temporal weighting together with richer memory dynamics including oscillatory, growth, and regime-dependent behaviour, while preserving the algebraic strengths of the classical signature. We show that the EWS is the unique solution to a linear controlled differential equation on the tensor algebra, and that it generalises both state-space models and the Laplace and Fourier transforms of the path. The group-like structure of the EWS enables efficient computation and makes the framework amenable to gradient-based learning, with the full semigroup action parametrised by and learned through its generator. We use this framework to empirically demonstrate the expressivity gap between the EWS and both the signature and EFM on two SDE-based regression tasks.

stat.ML

Seeking SOTA: Time-Series Forecasting Must Adopt Taxonomy-Specific Evaluation to Dispel Illusory Gains

We argue that the current practice of evaluating AI/ML time-series forecasting models, predominantly on benchmarks characterized by strong, persistent periodicities and seasonalities, obscures real progress by overlooking the performance of efficient classical methods. We demonstrate that these "standard" datasets often exhibit dominant autocorrelation patterns and seasonal cycles that can be effectively captured by simpler linear or statistical models, rendering complex deep learning architectures frequently no more performant than their classical counterparts for these specific data characteristics, and raising questions as to whether any marginal improvements justify the significant increase in computational overhead and model complexity. We call on the community to (I) retire or substantially augment current benchmarks with datasets exhibiting a wider spectrum of non-stationarities, such as structural breaks, time-varying volatility, and concept drift, and less predictable dynamics drawn from diverse real-world domains, and (II) require every deep learning submission to include robust classical and simple baselines, appropriately chosen for the specific characteristics of the downstream tasks' time series. By doing so, we will help ensure that reported gains reflect genuine scientific methodological advances rather than artifacts of benchmark selection favoring models adept at learning repetitive patterns.

cs.LG

Orthogonal polynomials on path-space

We consider the orthogonalisation of the signature of a stochastic process as the analogue of orthogonal polynomials on path-space. Under an infinite radius of convergence assumption, we prove density of linear functions on the signature in $L^p$ functions on grouplike elements, making it possible to represent a square-integrable function on (rough) paths as an $L^2$-convergent series. By viewing the shuffle algebra as commutative polynomials on the free Lie algebra, we revisit much of the theory of classical orthogonal polynomials in several variables, such as the recurrence relation and Favard's theorem. Finally, we restrict our attention to the case of Brownian motion with and without drift, and prove that dimension-independent orthogonal signature exists with drift but not without. We end with numerical examples of how orthogonal signature polynomials of Brownian motion can be applied for the approximation of functions on paths sampled from the Wiener measure.

math.PR

The Geometry of Rough Path Space

We describe $H^p(V)$, a subset of $p$-rough path space $\Omega_p(V)$ which is a vector space under an addition operation $\boxplus$ and a scalar multiplication $\odot$. We show that the domain of $\boxplus$ can be extended to $\Omega_p(V)\times H^p(V)$, allowing any $p$-rough path $X$ to be additively perturbed by an $H\in H^p(V)$. We prove associativity $(X\boxplus H)\boxplus \tilde H = X\boxplus (H\boxplus \tilde H)$ and trivial kernel $X\boxplus H = X \Leftrightarrow H = 1$, where $1$ is the additive zero in $(H^p(V),\boxplus,\odot)$. Finally, we show that enlarging $H^p(V)$ to almost rough paths $H^{am,p}(V)$ does not enlarge the set of displacements of a given $X$, i.e. $\{X\boxplus H: H\in H^p(V)\}=\{X\boxplus H: H\in H^{am,p}(V)\}$.

math.CA

Novelty detection on path space

We frame novelty detection on path space as a hypothesis testing problem with signature-based test statistics. Using transportation-cost inequalities of Gasteratos and Jacquier (2023), we obtain tail bounds for false positive rates that extend beyond Gaussian measures to laws of RDE solutions with smooth bounded vector fields, yielding estimates of quantiles and p-values. Exploiting the shuffle product, we derive exact formulae for smooth surrogates of conditional value-at-risk (CVaR) in terms of expected signatures, leading to new one-class SVM algorithms optimising smooth CVaR objectives. We then establish lower bounds on type-$\mathrm{II}$ error for alternatives with finite first moment, giving general power bounds when the reference measure and the alternative are absolutely continuous with respect to each other. Finally, we evaluate numerically the type-$\mathrm{I}$ error and statistical power of signature-based test statistic, using synthetic anomalous diffusion data and real-world molecular biology data.

stat.ML

Path Signatures Enable Model-Free Mapping of RNA Modifications

Detecting chemical modifications on RNA molecules remains a key challenge in epitranscriptomics. Traditional reverse transcription-based sequencing methods introduce enzyme- and sequence-dependent biases and fragment RNA molecules, confounding the accurate mapping of modifications across the transcriptome. Nanopore direct RNA sequencing offers a powerful alternative by preserving native RNA molecules, enabling the detection of modifications at single-molecule resolution. However, current computational tools can identify only a limited subset of modification types within well-characterized sequence contexts for which ample training data exists. Here, we introduce a model-free computational method that reframes modification detection as an anomaly detection problem, requiring only canonical (unmodified) RNA reads without any other annotated data. For each nanopore read, our approach extracts robust, modification-sensitive features from the raw ionic current signal at a site using the signature transform, then computes an anomaly score by comparing the resulting feature vector to its nearest neighbors in an unmodified reference dataset. We convert anomaly scores into statistical p-values to enable anomaly detection at both individual read and site levels. Validation on densely-modified \textit{E. coli} rRNA demonstrates that our approach detects known sites harboring diverse modification types, without prior training on these modifications. We further applyied this framework to dengue virus (DENV) transcripts and mammalian mRNAs. For DENV sfRNA, it led to revealing a novel 2'-O-methylated site, which we validate orthogonally by qRT-PCR assays. These results demonstrate that our model-free approach operates robustly across different types of RNAs and datasets generated with different nanopore sequencing chemistries.

q-bio.GN

Structured Linear CDEs: Maximally Expressive and Parallel-in-Time Sequence Models

This work introduces Structured Linear Controlled Differential Equations (SLiCEs), a unifying framework for sequence models with structured, input-dependent state-transition matrices that retain the maximal expressivity of dense matrices whilst being cheaper to compute. The framework encompasses existing architectures, such as input-dependent block-diagonal linear recurrent neural networks and DeltaNet's diagonal-plus-low-rank structure, as well as two novel variants based on sparsity and the Walsh-Hadamard transform. We prove that, unlike the diagonal state-transition matrices of S4D and Mamba, SLiCEs employing block-diagonal, sparse, or Walsh-Hadamard matrices match the maximal expressivity of dense matrices. Empirically, SLiCEs solve the $A_5$ state-tracking benchmark with a single layer, achieve best-in-class length generalisation on regular language tasks among parallel-in-time models, and match the performance of log neural controlled differential equations on six multivariate time-series classification datasets while cutting the average time per training step by a factor of twenty.

cs.LG

Transforming CCTV cameras into NO$_2$ sensors at city scale for adaptive policymaking

Air pollution in cities, especially NO\textsubscript{2}, is linked to numerous health problems, ranging from mortality to mental health challenges and attention deficits in children. While cities globally have initiated policies to curtail emissions, real-time monitoring remains challenging due to limited environmental sensors and their inconsistent distribution. This gap hinders the creation of adaptive urban policies that respond to the sequence of events and daily activities affecting pollution in cities. Here, we demonstrate how city CCTV cameras can act as a pseudo-NO\textsubscript{2} sensors. Using a predictive graph deep model, we utilised traffic flow from London's cameras in addition to environmental and spatial factors, generating NO\textsubscript{2} predictions from over 133 million frames. Our analysis of London's mobility patterns unveiled critical spatiotemporal connections, showing how specific traffic patterns affect NO\textsubscript{2} levels, sometimes with temporal lags of up to 6 hours. For instance, if trucks only drive at night, their effects on NO\textsubscript{2} levels are most likely to be seen in the morning when people commute. These findings cast doubt on the efficacy of some of the urban policies currently being implemented to reduce pollution. By leveraging existing camera infrastructure and our introduced methods, city planners and policymakers could cost-effectively monitor and mitigate the impact of NO\textsubscript{2} and other pollutants.

cs.LG

High-degree cubature on Wiener space through unshuffle expansions

Utilising classical results on the structure of Hopf algebras, we develop a novel approach for the construction of cubature formulae on Wiener space based on unshuffle expansions. We demonstrate the effectiveness of this approach by constructing the first explicit degree-7 cubature formula on $d$-dimensional Wiener space with drift, in the sense of Lyons and Victoir. The support of our degree-7 formula is significantly smaller than that of currently implemented or proposed constructions.

math.PR

Combining Hough Transform and Deep Learning Approaches to Reconstruct ECG Signals From Printouts

This work presents our team's (SignalSavants) winning contribution to the 2024 George B. Moody PhysioNet Challenge. The Challenge had two goals: reconstruct ECG signals from printouts and classify them for cardiac diseases. Our focus was the first task. Despite many ECGs being digitally recorded today, paper ECGs remain common throughout the world. Digitising them could help build more diverse datasets and enable automated analyses. However, the presence of varying recording standards and poor image quality requires a data-centric approach for developing robust models that can generalise effectively. Our approach combines the creation of a diverse training set, Hough transform to rotate images, a U-Net based segmentation model to identify individual signals, and mask vectorisation to reconstruct the signals. We assessed the performance of our models using the 10-fold stratified cross-validation (CV) split of 21,799 recordings proposed by the PTB-XL dataset. On the digitisation task, our model achieved an average CV signal-to-noise ratio of 17.02 and an official Challenge score of 12.15 on the hidden set, securing first place in the competition. Our study shows the challenges of building robust, generalisable, digitisation approaches. Such models require large amounts of resources (data, time, and computational power) but have great potential in diversifying the data available.

cs.LG

Deep Signature: Characterization of Large-Scale Molecular Dynamics

Understanding protein dynamics are essential for deciphering protein functional mechanisms and developing molecular therapies. However, the complex high-dimensional dynamics and interatomic interactions of biological processes pose significant challenge for existing computational techniques. In this paper, we approach this problem for the first time by introducing Deep Signature, a novel computationally tractable framework that characterizes complex dynamics and interatomic interactions based on their evolving trajectories. Specifically, our approach incorporates soft spectral clustering that locally aggregates cooperative dynamics to reduce the size of the system, as well as signature transform that collects iterated integrals to provide a global characterization of the non-smooth interactive dynamics. Theoretical analysis demonstrates that Deep Signature exhibits several desirable properties, including invariance to translation, near invariance to rotation, equivariance to permutation of atomic coordinates, and invariance under time reparameterization. Furthermore, experimental results on three benchmarks of biological processes verify that our approach can achieve superior performance compared to baseline methods.

q-bio.QM

Higher Order Lipschitz Greedy Recombination Interpolation Method (HOLGRIM)

In this paper we introduce the Higher Order Lipschitz Greedy Recombination Interpolation Method (HOLGRIM) for finding sparse approximations of Lip$(\gamma)$ functions, in the sense of Stein, given as a linear combination of a (large) number of simpler Lip$(\gamma)$ functions. HOLGRIM is developed as a refinement of the Greedy Recombination Interpolation Method (GRIM) in the setting of Lip$(\gamma)$ functions. HOLGRIM combines dynamic growth-based interpolation techniques with thinning-based reduction techniques in a data-driven fashion. The dynamic growth is driven by a greedy selection algorithm in which multiple new points may be selected at each step. The thinning reduction is carried out by recombination, the linear algebra technique utilised by GRIM. We establish that the number of non-zero weights for the approximation returned by HOLGRIM is controlled by a particular packing number of the data. The level of data concentration required to guarantee that HOLGRIM returns a good sparse approximation is decreasing with respect to the regularity parameter $\gamma > 0$. Further, we establish complexity cost estimates verifying that implementing HOLGRIM is feasible.

math.NA

Higher Order Lipschitz Sandwich Theorems

We investigate the consequence of two Lip$(\gamma)$ functions, in the sense of Stein, being close throughout a subset of their domain. A particular consequence of our results is the following. Given $K_0 > \varepsilon > 0$ and $\gamma > \eta > 0$ there is a constant $\delta = \delta(\gamma,\eta,\varepsilon,K_0) > 0$ for which the following is true. Let $\Sigma \subset \mathbb{R}^d$ be closed and $f , h : \Sigma \to \mathbb{R}$ be Lip$(\gamma)$ functions whose Lip$(\gamma)$ norms are both bounded above by $K_0$. Suppose $B \subset \Sigma$ is closed and that $f$ and $h$ coincide throughout $B$. Then over the set of points in $\Sigma$ whose distance to $B$ is at most $\delta$ we have that the Lip$(\eta)$ norm of the difference $f-h$ is bounded above by $\varepsilon$. More generally, we establish that this phenomenon remains valid in a less restrictive Banach space setting under the weaker hypothesis that the two Lip$(\gamma)$ functions $f$ and $h$ are only close in a pointwise sense throughout the closed subset $B$. We require only that the subset $\Sigma$ be closed; in particular, the case that $\Sigma$ is finite is covered by our results. The restriction that $\eta < \gamma$ is sharp in the sense that our result is false for $\eta := \gamma$.

math.CA

Log-PDE Methods for Rough Signature Kernels

Signature kernels, inner products of path signatures, underpin several machine learning algorithms for multivariate time series analysis. For bounded variation paths, signature kernels were recently shown to solve a Goursat PDE. However, existing PDE solvers only use increments as input data, leading to first order approximation errors. These approaches become computationally intractable for highly oscillatory input paths, as they have to be resolved at a fine enough scale to accurately recover their signature kernel, resulting in significant time and memory complexities. In this paper, we extend the analysis to rough paths, and show, leveraging the framework of smooth rough paths, that the resulting rough signature kernels can be approximated by a novel system of PDEs whose coefficients involve higher order iterated integrals of the input rough paths. We show that this system of PDEs admits a unique solution and establish quantitative error bounds yielding a higher order approximation to rough signature kernels.

cs.LG

Multimodal deep learning approach to predicting neurological recovery from coma after cardiac arrest

This work showcases our team's (The BEEGees) contributions to the 2023 George B. Moody PhysioNet Challenge. The aim was to predict neurological recovery from coma following cardiac arrest using clinical data and time-series such as multi-channel EEG and ECG signals. Our modelling approach is multimodal, based on two-dimensional spectrogram representations derived from numerous EEG channels, alongside the integration of clinical data and features extracted directly from EEG recordings. Our submitted model achieved a Challenge score of $0.53$ on the hidden test set for predictions made $72$ hours after return of spontaneous circulation. Our study shows the efficacy and limitations of employing transfer learning in medical classification. With regard to prospective implementation, our analysis reveals that the performance of the model is strongly linked to the selection of a decision threshold and exhibits strong variability across data splits.

cs.LG