SearcharxivSearch

arXiv subjects

Ben Hayes

Publications and source records attributed to Ben Hayes.

At least 19 recordsLinked to original sources

Questions on the structure of random embeddings of $L(\mathbb{F}_2)$

Motivated by recent developments at the interface of operator algebras and random matrix theory, we propose new conjectures concerning the asymptotic structure of random matrix models of the countable free groups. The first conjecture predicts a random matrix analogue of the Akemann-Ostrand property for free groups, and reveals a succinct approach to recover the Peterson-Thom property for $L(\mathbb{F}_2)$. The second stronger conjecture is motivated by continuous model theory. It predicts that the \emph{random} embedding of the free group factor into a matrix ultraproduct is \emph{existential}. We discuss the interesting relationship between these conjectures.

math.OA

LiveBand: Live Accompaniment Generation in the Audio Domain

We present LiveBand, a real-time system that generates high-fidelity music accompaniments to live audio input, respecting strict causal constraints. Our method trains a causal transformer generator in the continuous latent space of a pre-trained causal audio autoencoder, using adversarial sequence-level supervision from a discriminator. At each timestep, the generator receives only the causally available mix context and Gaussian noise, and predicts accompaniment latents without access to future mix frames or ground-truth target latents. Training is performed in a single parallel forward pass under causal masking, while streaming inference proceeds autoregressively with a rolling attention state. The model's training and inference computations are matched by design, eliminating teacher forcing and the associated exposure bias. On a multi-instrument music accompaniment benchmark, LiveBand improves over prior work on objective measures of audio quality, beat alignment, and mix adherence, while enabling real-time streaming generation without lookahead into the future on consumer hardware.

cs.SD

Coamenability and strong ergodicity

Following methods of Bannon-Marrakchi-Ozawa, we show that for coamenable inclusion $\mathcal{S}\leq \mathcal{R}$ of ergodic, probability measure-preserving relations, we have that $\mathcal{R}$ is strongly ergodic if and only if $\mathcal{S}$ is strongly ergodic. More general results are given when $\mathcal{S}\leq \mathcal{R}$ is coamenable, $\mathcal{R}$ is strongly ergodic, but we do not assume ergodicity of $\mathcal{S}$. As a consequence, if $\Lambda\leq \Gamma$ is a coamenable inclusion of groups, then any strongly ergodic $\Gamma$ action has countably many ergodic components for the $\Lambda$ action, each of which is strongly ergodic.

math.DS

Selfless Inclusions of C*-Algebras

We introduce and study a natural notion of selflessness for inclusions of C*-probability spaces, which in particular implies that all intermediate C*-algebras are selfless in the sense of Robert. We identify natural sources of selfless inclusions in the realms of Z-stable and free product C*-algebras. As an application of this, we prove selflessness for a new family of C*-probability spaces outside the regime of free products and group C*-algebras. These include the reduced free unitary compact quantum groups.

math.OA

PESTO: Real-Time Pitch Estimation with Self-supervised Transposition-equivariant Objective

In this paper, we introduce PESTO, a self-supervised learning approach for single-pitch estimation using a Siamese architecture. Our model processes individual frames of a Variable-$Q$ Transform (VQT) and predicts pitch distributions. The neural network is designed to be equivariant to translations, notably thanks to a Toeplitz fully-connected layer. In addition, we construct pitch-shifted pairs by translating and cropping the VQT frames and train our model with a novel class-based transposition-equivariant objective, eliminating the need for annotated data. Thanks to this architecture and training objective, our model achieves remarkable performances while being very lightweight ($130$k parameters). Evaluations on music and speech datasets (MIR-1K, MDB-stem-synth, and PTDB) demonstrate that PESTO not only outperforms self-supervised baselines but also competes with supervised methods, exhibiting superior cross-dataset generalization. Finally, we enhance PESTO's practical utility by developing a streamable VQT implementation using cached convolutions. Combined with our model's low latency (less than 10 ms) and minimal parameter count, this makes PESTO particularly suitable for real-time applications.

cs.SD

Audio synthesizer inversion in symmetric parameter spaces with approximately equivariant flow matching

Many audio synthesizers can produce the same signal given different parameter configurations, meaning the inversion from sound to parameters is an inherently ill-posed problem. We show that this is largely due to intrinsic symmetries of the synthesizer, and focus in particular on permutation invariance. First, we demonstrate on a synthetic task that regressing point estimates under permutation symmetry degrades performance, even when using a permutation-invariant loss function or symmetry-breaking heuristics. Then, viewing equivalent solutions as modes of a probability distribution, we show that a conditional generative model substantially improves performance. Further, acknowledging the invariance of the implicit parameter distribution, we find that performance is further improved by using a permutation equivariant continuous normalizing flow. To accommodate intricate symmetries in real synthesizers, we also propose a relaxed equivariance strategy that adaptively discovers relevant symmetries from data. Applying our method to Surge XT, a full-featured open source synthesizer used in real world audio production, we find our method outperforms regression and generative baselines across audio reconstruction metrics.

cs.SD

Selfless reduced free product $C^*$-algebras

We study selflessness in the general setting of reduced free products of $C^*$-algebras. Towards this end, we develop a suitable theory of rapid decay for filtrations in arbitrary $C^*$-probability spaces. We provide several natural examples and permanence properties of this phenomenon. By using this framework in combination with von Neumann algebraic techniques involving approximate forms of orthogonality, we are able to prove selflessness for general families of reduced free product $C^*$-algebras. As an instance of our results, we prove selflessness and thus strict comparison for the canonical $C^*$-algebras generated by Voiculescu's free semicircular systems. Our results also provide new examples of purely infinite reduced free products.

math.OA

DiffVox: A Differentiable Model for Capturing and Analysing Vocal Effects Distributions

This study introduces a novel and interpretable model, DiffVox, for matching vocal effects in music production. DiffVox, short for ``Differentiable Vocal Fx", integrates parametric equalisation, dynamic range control, delay, and reverb with efficient differentiable implementations to enable gradient-based optimisation for parameter estimation. Vocal presets are retrieved from two datasets, comprising 70 tracks from MedleyDB and 365 tracks from a private collection. Analysis of parameter correlations reveals strong relationships between effects and parameters, such as the high-pass and low-shelf filters often working together to shape the low end, and the delay time correlating with the intensity of the delayed signals. Principal component analysis reveals connections to McAdams' timbre dimensions, where the most crucial component modulates the perceived spaciousness while the secondary components influence spectral brightness. Statistical testing confirms the non-Gaussian nature of the parameter distribution, highlighting the complexity of the vocal effects space. These initial findings on the parameter distributions set the foundation for future research in vocal effects modelling and automatic mixing. Our source code and datasets are accessible at https://github.com/SonyResearch/diffvox.

cs.SD

Coamenability and cospectral radius for orbit equivalence relations

We consider inclusions $\mathcal{S}\leq \mathcal{R}$ of discrete, probability measure-preserving orbit equivalence relations. In previous work with Ab\'{e}rt-Fra\c{c}zyk, we established the pointwise almost sure existence of the cospectral radius of a random walk on the $\mathcal{R}$-classes. In this paper, we investigate the connections of this cospectral radius to the coamenability of the inclusion $\mathcal{S}\leq \mathcal{R}$. We also undertake a systematic study of coamenability for inclusions of relations, establishing several equivalence formulations of this notion.

math.DS

General solidity phenomena and anticoarse spaces for type $\mathrm{III}_1$ factors

By developing a theory of anticoarse spaces in the purely infinite setting and using 1-bounded entropy techniques along with recent strong convergence results in random matrix theory, we show that free Araki--Woods factors offer the first examples of type $\mathrm{III}$ factors satisfying vastly general degrees of indecomposability phenomena. Notably this includes strong solidity with respect to any weakening of the normalizer currently in the literature and the Peterson--Thom property.

math.OA

On the structure of graph product von Neumann algebras

We undertake a comprehensive study of structural properties of graph products of von Neumann algebras equipped with faithful, normal states, as well as properties of the graph products relative to subalgebras coming from induced subgraphs. Among the technical contributions in this paper include a complete bimodule calculation for subalgebras arising from subgraphs. As an application, we obtain a complete classification of when two subalgebras coming from induced subgraphs can be amenable relative to each other. We also give complete characterizations of when the graph product can be full, diffuse, or a factor. Our results are obtained in a broad generality, and we emphasize that they are new even in the tracial setting. They also allow us to deduce new results about when graph products of groups can be amenable relative to each other.

math.OA

Random permutation matrix models for graph products

Graph independence (also known as $\epsilon$-independence or $\lambda$-independence) is a mixture of classical independence and free independence corresponding to graph products or groups and operator algebras. Using conjugation by certain random permutation matrices, we construct random matrix models for graph independence with amalgamation over the diagonal matrices. This yields a new probabilist,ic proof that graph products of sofic groups are sofic.

math.OA

Spectral non-concentration near the top for unimodular random graphs

In recent work on equiangular lines, Jiang, Tidor, Yuan, Zhang, and Zhao showed that a connected bounded degree graph has sublinear second eigenvalue multiplicity. More generally they show that there cannot be too many eigenvalues near the top of the spectrum. We extend this result to infinite unimodular random graphs. As a corollary, the spectral distribution of the adjacency operator cannot have an atom at the top. For an infinite regular expander, we deduce that the singularity of the spectral measure at the top satisfies $\mu_G[(1-\theta)\rho,\rho] \lesssim \theta^c$ for some constant $c>0$, where $\rho$ is the spectral radius of the adjacency operator of the graph. This implies new general estimates on the return probabilities of random walks.

math.PR

Growth dichotomy for unimodular random rooted trees

We show that the growth of a unimodular random rooted tree $(T,o)$ of degree bounded by $d$ always exists, assuming its upper growth passes the critical threshold $\sqrt{d-1}$. This complements Timar's work who showed the possible nonexistence of growth below this threshold. The proof goes as follows. By Benjamini-Lyons-Schramm, we can realize $(T,o)$ as the cluster of the root for some invariant percolation on the $d$-regular tree. Then we show that for such a percolation, the limiting exponent with which the lazy random walk returns to the cluster of its starting point always exists. We develop a new method to get this, that we call the 2-3-method, as the usual pointwise ergodic theorems do not seem to work here. We then define and prove the Cohen-Grigorchuk co-growth formula to the invariant percolation setting. This establishes and expresses the growth of the cluster from the limiting exponent, assuming we are above the critical threshold.

math.PR

AI (r)evolution -- where are we heading? Thoughts about the future of music and sound technologies in the era of deep learning

Artificial Intelligence (AI) technologies such as deep learning are evolving very quickly bringing many changes to our everyday lives. To explore the future impact and potential of AI in the field of music and sound technologies a doctoral day was held between Queen Mary University of London (QMUL, UK) and Sciences et Technologies de la Musique et du Son (STMS, France). Prompt questions about current trends in AI and music were generated by academics from QMUL and STMS. Students from the two institutions then debated these questions. This report presents a summary of the student debates on the topics of: Data, Impact, and the Environment; Responsible Innovation and Creative Practice; Creativity and Bias; and From Tools to the Singularity. The students represent the future generation of AI and music researchers. The academics represent the incumbent establishment. The student debates reported here capture visions, dreams, concerns, uncertainties, and contentious issues for the future of AI and music as the establishment is rightfully challenged by the next generation.

cs.CY

A Review of Differentiable Digital Signal Processing for Music & Speech Synthesis

The term "differentiable digital signal processing" describes a family of techniques in which loss function gradients are backpropagated through digital signal processors, facilitating their integration into neural networks. This article surveys the literature on differentiable audio signal processing, focusing on its use in music & speech synthesis. We catalogue applications to tasks including music performance rendering, sound matching, and voice transformation, discussing the motivations for and implications of the use of this methodology. This is accompanied by an overview of digital signal processing operations that have been implemented differentiably. Finally, we highlight open challenges, including optimisation pathologies, robustness to real-world conditions, and design trade-offs, and discuss directions for future research.

cs.SD

Consequences of the random matrix solution to the Peterson-Thom conjecture

In this paper we show various new structural properties of free group factors using the recent resolution (due independently to Belinschi-Capitaine and Bordenave-Collins) of the Peterson-Thom conjecture. These results include the resolution to the coarseness conjecture independently due to the first-named author and Popa, a generalization of Ozawa-Popa's celebrated strong solidity result using vastly more general versions of the normalizer (and in an ultraproduct setting), a dichotomy result for intertwining of maximal amenable subalgebras of interpolated free group factors, as well as application to ultraproduct embeddings of nonamenable subalgebras of interpolated free group factors.

math.OA