SearcharxivSearch

arXiv subjects

Xavier Gonzalez

Publications and source records attributed to Xavier Gonzalez.

9 recordsLinked to original sources

Unifying Optimization and Dynamics to Parallelize Sequential Computation: A Guide to Parallel Newton Methods for Breaking Sequential Bottlenecks

Massively parallel hardware (GPUs) and long sequence data have made parallel algorithms essential for machine learning at scale. Yet dynamical systems, like recurrent neural networks and Markov chain Monte Carlo, were thought to suffer from sequential bottlenecks. Recent work showed that dynamical systems can in fact be parallelized across the sequence length by reframing their evaluation as a system of nonlinear equations, which can be solved with Newton's method using a parallel associative scan. However, these parallel Newton methods struggled with limitations, primarily inefficiency, instability, and lack of convergence guarantees. This thesis addresses these limitations with methodological and theoretical contributions, drawing particularly from optimization. Methodologically, we develop scalable and stable parallel Newton methods, based on quasi-Newton and trust-region approaches. The quasi-Newton methods are faster and more memory efficient, while the trust-region approaches are significantly more stable. Theoretically, we unify many fixed-point methods into our parallel Newton framework, including Picard and Jacobi iterations. We establish a linear convergence rate for these techniques that depends on the method's approximation accuracy and stability. Moreover, we give a precise condition, rooted in dynamical stability, that characterizes when parallelization provably accelerates a dynamical system and when it cannot. Specifically, the sign of the Largest Lyapunov Exponent of a dynamical system determines whether or not parallel Newton methods converge quickly. In sum, this thesis unlocks scalable and stable methods for parallelizing sequential computation, and provides a firm theoretical basis for when such techniques will and will not work. This thesis also serves as a guide to parallel Newton methods for researchers who want to write the next chapter in this ongoing story.

math.NA

A Unifying Framework for Parallelizing Sequential Models with Linear Dynamical Systems

Harnessing parallelism in seemingly sequential models is a central challenge for modern machine learning. Several approaches have been proposed for evaluating sequential processes in parallel using iterative fixed-point methods, like Newton, Picard, and Jacobi iterations. In this work, we show that these methods can be understood within a common framework based on linear dynamical systems (LDSs), where different iteration schemes arise naturally as approximate linearizations of a nonlinear recursion. Moreover, we theoretically analyze the rates of convergence of these methods, and we verify the predictions of this theory with several case studies. This unifying framework highlights shared principles behind these techniques and clarifies when particular fixed-point methods are most likely to be effective. By bridging diverse algorithms through the language of LDSs, the framework provides a clearer theoretical foundation for parallelizing sequential models and points toward new opportunities for efficient and scalable computation.

cs.LG

Parallelizing MCMC Across the Sequence Length

Markov chain Monte Carlo (MCMC) methods are foundational algorithms for Bayesian inference and probabilistic modeling. However, most MCMC algorithms are inherently sequential and their time complexity scales linearly with the sequence length. Previous work on adapting MCMC to modern hardware has therefore focused on running many independent chains in parallel. Here, we take an alternative approach: we propose algorithms to evaluate MCMC samplers in parallel across the chain length. To do this, we build on recent methods for parallel evaluation of nonlinear recursions that formulate the state sequence as a solution to a fixed-point problem and solve for the fixed-point using a parallel form of Newton's method. We show how this approach can be used to parallelize Gibbs, Metropolis-adjusted Langevin, and Hamiltonian Monte Carlo sampling across the sequence length. In several examples, we demonstrate the simulation of up to hundreds of thousands of MCMC samples with only tens of parallel Newton iterations. Additionally, we develop two new parallel quasi-Newton methods to evaluate nonlinear recursions with lower memory costs and reduced runtime. We find that the proposed parallel algorithms accelerate MCMC sampling across multiple examples, in some cases by more than an order of magnitude compared to sequential evaluation.

stat.CO

Predictability Enables Parallelization of Nonlinear State Space Models

The rise of parallel computing hardware has made it increasingly important to understand which nonlinear state space models can be efficiently parallelized. Recent advances like DEER (arXiv:2309.12252) and DeepPCR (arXiv:2309.16318) recast sequential evaluation as a parallelizable optimization problem, sometimes yielding dramatic speedups. However, the factors governing the difficulty of these optimization problems remained unclear, limiting broader adoption. In this work, we establish a precise relationship between a system's dynamics and the conditioning of its corresponding optimization problem, as measured by its Polyak-Lojasiewicz (PL) constant. We show that the predictability of a system, defined as the degree to which small perturbations in state influence future behavior and quantified by the largest Lyapunov exponent (LLE), impacts the number of optimization steps required for evaluation. For predictable systems, the state trajectory can be computed in at worst $O((\log T)^2)$ time, where $T$ is the sequence length: a major improvement over the conventional sequential approach. In contrast, chaotic or unpredictable systems exhibit poor conditioning, with the consequence that parallel evaluation converges too slowly to be useful. Importantly, our theoretical analysis shows that predictable systems always yield well-conditioned optimization problems, whereas unpredictable systems lead to severe conditioning degradation. We validate our claims through extensive experiments, providing practical guidance on when nonlinear dynamical systems can be efficiently parallelized. We highlight predictability as a key design principle for parallelizable models.

math.OC

Towards Scalable and Stable Parallelization of Nonlinear RNNs

Transformers and linear state space models can be evaluated in parallel on modern hardware, but evaluating nonlinear RNNs appears to be an inherently sequential problem. Recently, however, Lim et al. '24 developed an approach called DEER, which evaluates nonlinear RNNs in parallel by posing the states as the solution to a fixed-point problem. They derived a parallel form of Newton's method to solve the fixed-point problem and achieved significant speedups over sequential evaluation. However, the computational complexity of DEER is cubic in the state size, and the algorithm can suffer from numerical instability. We address these limitations with two novel contributions. To reduce the computational complexity, we apply quasi-Newton approximations and show they converge comparably to Newton, use less memory, and are faster. To stabilize DEER, we leverage a connection between the Levenberg-Marquardt algorithm and Kalman smoothing, which we call ELK. This connection allows us to stabilize Newton's method while using efficient parallelized Kalman smoothing algorithms to retain performance. Through several experiments, we show that these innovations allow for parallel evaluation of nonlinear RNNs at larger scales and with greater stability.

cs.LG

From partitions to Hodge numbers of Hilbert Schemes of Surfaces

We celebrate the 100th anniversary of Srinivasa Ramanujan's election as a Fellow of the Royal Society, which was largely based on his work with G. H. Hardy on the asymptotic properties of the partition function. After recalling this revolutionary work, marking the birth of the "circle method", we present a contemporary example of its legacy in topology. We deduce the equidistribution of Hodge numbers for Hilbert schemes of suitable smooth projective surfaces.

math.NT

Exact Formulas for Invariants of Hilbert Schemes

A theorem of Göttsche establishes a connection between cohomological invariants of a complex projective surface $S$ and corresponding invariants of the Hilbert scheme of $n$ points on $S.$ This relationship is encoded in certain infinite product $q$-series which are essentially modular forms. Here we make use of the circle method to arrive at exact formulas for certain specializations of these $q$-series, yielding convergent series for the signature and Euler characteristic of these Hilbert schemes. We also analyze the asymptotic and distributional properties of the $q$-series' coefficients.

math.NT

Moonshine for All Finite Groups

In recent literature, moonshine has been explored for some groups beyond the Monster, for example the sporadic O'Nan and Thompson groups. This collection of examples may suggest that moonshine is a rare phenomenon, but a fundamental and largely unexplored question is how general the correspondence is between modular forms and finite groups. For every finite group $G$, we give constructions of infinitely many graded infinite-dimensional $\mathbb{C}[G]$-modules where the McKay-Thompson series for a conjugacy class $[g]$ is a weakly holomorphic modular function properly on $Γ_0(\text{ord}(g))$. As there are only finitely many normalized Hauptmoduln, groups whose McKay-Thompson series are normalized Hauptmoduln are rare, but not as rare as one might naively expect. We give bounds on the powers of primes dividing the order of groups which have normalized Hauptmoduln of level $\text{ord}(g)$ as the graded trace functions for any conjugacy class $[g]$, and completely classify the finite abelian groups with this property. In particular, these include $(\mathbb{Z} / 5 \mathbb{Z})^5$ and $(\mathbb{Z} / 7 \mathbb{Z})^4$, which are not subgroups of the Monster.

math.NT