SearcharxivSearch

arXiv subjects

Michael Lindsey

Publications and source records attributed to Michael Lindsey.

At least 19 recordsLinked to original sources

GoForth: Language Models for RNA Design under Structure, Sequence, and Coding Constraints

RNA inverse sequence design has broad biological and engineering applications, but computational methods for practical design queries remain limited. Such queries may impose several constraints at once, including target folds or motifs, fixed bases, and coding restrictions, while leaving arbitrary sequence and structure in unspecified regions. Because these constraints may permit many acceptable sequences, we study RNA design as a conditional generative modeling problem. The basic object is a conditional law over RNA sequences given a user-specified condition, with full inverse folding as a special case. We introduce GoForth, a forward-trained RNA language model that conditions on structure, sequence, and coding targets. The formulation separates three ingredients that are often entangled in RNA design: a sequence prior, a forward folding sampler, and a reward or likelihood oracle. We train encoder-decoder models on witnessed folds rather than on outputs from an inverse-design teacher and validate our methodology on full inverse-folding benchmarks, as well as tasks involving constraints on structure, sequence, and coding. The resulting models achieve fast and high-quality candidate generation for mixed RNA design specifications. Moreover they furnish useful semantic embeddings of design tasks and a robust learned notion of designability.

q-bio.QM

Direct RNA sequence design under codon constraints using expressive tensor-based secondary structure models

Nucleic acid sequence design via codon optimization is a fundamental task with applications across synthetic biology, mRNA therapeutics, and vaccine design. Given a target protein, it is a major open challenge to navigate the combinatorially large design space of codon sequences mapping to its amino acid sequence. Computational approaches generally seek to optimize simple objectives based on the codon sequence, possibly together with more complicated contributions based on secondary structure analysis. In this work, we demonstrate a direct and efficient algorithm to sample sequences from a suitable Boltzmann distribution defined in terms of the codon sequence and a fully detailed secondary structure free energy model, as well as related algorithms for exact computation of statistical quantities such as free energies, base pairing probabilities, and base and codon marginals. These algorithms draw upon a recently developed tensor-based formulation of secondary structure thermodynamics and demonstrate, for the first time, that global sequence design can be accomplished with respect to a highly accurate free energy model. Moreover, the algorithms can leverage any available CPU and GPU resources in parallel for massive computational speedups.

q-bio.QM

Fast adaptive discontinuous basis sets for electronic structure

We develop a discontinuous Galerkin (DG) framework for automatically constructing adaptive basis sets for electronic structure calculations. By allowing basis functions to be discontinuous across element interfaces, our approach supports flexible combinations of atom-centered and polynomial basis sets, maintains favorable numerical conditioning, and induces structured sparsity of the one- and two-electron integrals, which we compute using specialised numerical integration strategies. In addition, we introduce a simple post-processing procedure to obtain continuous solutions if desired. We also introduce multigrid-preconditioned Poisson solvers that enable fast algorithms for both Hartree-Fock (HF) and density functional theory (DFT) calculations within our DG basis sets. Moreover, these basis sets naturally support adaptive multigrid preconditioning for the linear eigensolvers employed within the self-consistent field iteration for HF and DFT. Numerical experiments for HF and DFT demonstrate that our approach achieves chemical accuracy with modest basis sizes that compare favorably to the sizes of ordinary GTO basis sets achieving similar accuracy, while offering additional structured sparsity and improved computational scalability in the size-extensive limit. The framework thus provides a flexible route toward the construction of systematically improvable and structured adaptive basis sets for electronic structure theory.

physics.comp-ph

An approximation theory for Markov chain compression

We develop a framework for the compression of reversible Markov chains with rigorous error control. Given a subset of selected states, we construct reduced dynamics that can be lifted to an approximation of the full dynamics, and we prove simple spectral and nuclear norm bounds on the recovery error in terms of a suitably interpreted Nystr\"{o}m approximation error. We introduce two compression schemes: a projective compression based on committor functions and a structure-preserving compression defined in terms of an induced Markov chain over the selected states. The Nystr\"{o}m error appearing in our bounds can be controlled using recent results on column subset selection by nuclear maximization. Numerical experiments validate our theory and demonstrate the scalability of our approach.

math.NA

Fast entropy-regularized SDP relaxations for permutation synchronization

We introduce fast randomized algorithms for solving semidefinite programming (SDP) relaxations of the partial permutation synchronization (PPS) problem, a core task in multi-image matching with significant relevance to 3D reconstruction. Our methods build on recent advances in entropy-regularized semidefinite programming and are tailored to the unique structure of PPS, in which the unknowns are partial permutation matrices aligning sparse and noisy pairwise correspondences across images. We prove that entropy regularization resolves optimizer non-uniqueness in standard relaxations, and we develop a randomized solver with nearly optimal scaling in the number of observed correspondences. We also develop several rounding procedures for recovering combinatorial solutions from the implicitly represented primal solution variable, maintaining cycle consistency if desired without harming computational scaling. We demonstrate that our approach achieves state-of-the-art performance on synthetic and real-world datasets in terms of speed and accuracy. Our results highlight PPS as a paradigmatic setting in which entropy-regularized SDP admits both theoretical and practical advantages over traditional low-rank or spectral techniques.

math.OC

Improved energies and wave function accuracy with Weighted Variational Monte Carlo

Neural network parametrizations have increasingly been used to represent the ground and excited states in variational Monte Carlo (VMC) with promising results. However, traditional VMC methods only optimize the wave function in regions of peak probability. The wave function is uncontrolled in the tails of the probability distribution, which can limit the accuracy of the trained wavefunction approximation. To improve the approximation accuracy in the probability tails, this paper interprets VMC as a gradient flow in the space of wave functions, followed by a projection step. From this perspective, arbitrary probability distributions can be used in the projection step, allowing the user to prioritize accuracy in different regions of state space. Motivated by this theoretical perspective, the paper tests a new weighted VMC method on the antiferromagnetic Heisenberg model for a periodic spin chain. Compared to traditional VMC, weighted VMC reduces the error in the ground state energy by a factor of 2 and it reduces the errors in the local energies away from the mode by large factors of $10^2$--$10^4$.

physics.comp-ph

Non-Euclidean dual gradient ascent for entropically regularized linear and semidefinite programming

We present an optimization framework that exhibits dimension-independent convergence on a broad class of semidefinite programs (SDPs). Our approach first regularizes the primal problem with the von Neumann entropy, then solve the regularized problem using dual gradient ascent with respect to a problem-adapted norm. In particular, we show that the dual gradient norm converges to zero at a rate independent of the ambient dimension and, via rounding arguments, construct primal-feasible solutions in certain special cases. We also derive explicit convergence rates for the objective. In order to achieve optimal computational scaling, we must accommodate the use of stochastic gradients constructed via randomized trace estimators. Throughout we illustrate the generality of our framework via three important special cases -- the Goemans-Williamson SDP relaxation of the Max-Cut problem, the optimal transport linear program, and several SDP relaxations of the permutation synchronization problem. Numerical experiments confirm that our methods achieve dimension-independent convergence in practice.

math.OC

Toward optimal-scaling DFT: stochastic Hartree theory in the thermodynamic and complete basis set limits at arbitrary temperature

We present the first mathematical analysis of stochastic density functional theory (DFT) in the context of the Hartree approximation. We motivate our analysis via the notion of nearly-optimal or $\tilde{O}(n)$ scaling with respect to the number $n$ of computational degrees of freedom, independent of the number of electrons, in both the thermodynamic and complete basis set limits. Indeed, the promise of such scaling is the primary motivation for stochastic DFT relative to conventional orbital-based approaches, as well as deterministic orbital-free alternatives. We highlight three key targets for mathematical attention, which are synthesized in our algorithm and analysis. First, we identify a particular stochastic estimator for the Hartree potential whose sample complexity is essentially independent of the discretization size. Second, we reformulate the self-consistent field iteration as a stochastic mirror descent method where the Fermi-Dirac entropy plays the role of the Bregman potential, and we prove a nearly discretization-independent bound on the number of iterations needed to reach fixed accuracy. Third, motivated by the estimator, we introduce a novel pole expansion scheme for the square-root Fermi-Dirac operator, preserving $\tilde{O}(n)$ cost per mirror descent iteration even in the complete basis set limit. Combining these ingredients, we establish nearly-optimal scaling in both limits of interest under reasonable assumptions on the basis sets chosen for discretization. Extensive numerical experiments on problems with as many as $10^{6}$ degrees of freedom validate our algorithm and support the theory of nearly-optimal scaling.

math.NA

Geometric adaptive smoothed aggregation multigrid for discontinuous Galerkin discretisations

We present a geometric multigrid solver based on adaptive smoothed aggregation suitable for Discontinuous Galerkin (DG) discretisations. Mesh hierarchies are formed via domain decomposition techniques, and the method is applicable to fully unstructured meshes using arbitrary element shapes. Furthermore, the method can be employed for a wide range of commonly used DG numerical fluxes for first- and second-order PDEs including the Interior Penalty and the Local Discontinuous Galerkin methods. We demonstrate excellent and mesh-independent convergence for a range of problems including the Poisson equation, and convection-diffusion for a range of P\'eclet numbers.

math.NA

Fast and Accurate Interpolative Decompositions for General, Sparse, and Structured Tensors

In this work, we develop deterministic and random sketching-based algorithms for two types of tensor interpolative decompositions (ID): the core interpolative decomposition (CoreID, also known as the structure-preserving HOSVD) and the satellite interpolative decomposition (SatID, also known as the HOID or CURT). We adopt a new adaptive approach that leads to ID error bounds independent of the size of the tensor. In addition to the adaptive approach, we use tools from random sketching to enable an efficient and provably accurate calculation of these decompositions. We also design algorithms specialized to tensors that are sparse or given as a sum of rank-one tensors, i.e., in the CP format. Besides theoretical analyses, numerical experiments on both synthetic and real-world data demonstrate the power of the proposed algorithms.

math.NA

Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks

We establish the asymptotic implicit bias of gradient descent (GD) for generic non-homogeneous deep networks under exponential loss. Specifically, we characterize three key properties of GD iterates starting from a sufficiently small empirical risk, where the threshold is determined by a measure of the network's non-homogeneity. First, we show that a normalized margin induced by the GD iterates increases nearly monotonically. Second, we prove that while the norm of the GD iterates diverges to infinity, the iterates themselves converge in direction. Finally, we establish that this directional limit satisfies the Karush-Kuhn-Tucker (KKT) conditions of a margin maximization problem. Prior works on implicit bias have focused exclusively on homogeneous networks; in contrast, our results apply to a broad class of non-homogeneous networks satisfying a mild near-homogeneity condition. In particular, our results apply to networks with residual connections and non-homogeneous activation functions, thereby resolving an open problem posed by Ji and Telgarsky (2020).

cs.LG

MNE: overparametrized neural evolution with applications to diffusion processes and sampling

We propose a framework for solving evolution equations within parametric function classes, especially ones that are specified by neural networks. We call this framework the minimal neural evolution (MNE) because it is motivated by the goal of seeking the smallest instantaneous change in the neural network parameters that is compatible with exact solution of the evolution equation at a set of evolving collocation points. Formally, the MNE is quite similar to the recently introduced Neural Galerkin framework, but a difference in perspective motivates an alternative sketching procedure that effectively reduces the linear systems solved within the integrator to a size that is interpretable as an effective rank of the evolving neural tangent kernel, while maintaining a smooth evolution equation for the neural network parameters. We focus specifically on the application of this framework to diffusion processes, where the score function allows us to define intuitive dynamics for the collocation points. These can in turn be propagated jointly with the neural network parameters using a high-order adaptive integrator. In particular, we demonstrate how the Ornstein-Uhlenbeck diffusion process can be used for the task of sampling from a probability distribution given a formula for the density but no training data. This framework extends naturally to allow for conditional sampling and marginalization, and we show how to systematically remove the sampling bias due to parametric approximation error. We validate the efficiency, systematic improvability, and scalability of our approach on illustrative examples in low and high spatial dimensions.

math.NA

A gradient-based and determinant-free framework for fully Bayesian Gaussian process regression

Gaussian Process Regression (GPR) is widely used for inferring functions from noisy data. GPR crucially relies on the choice of a kernel, which might be specified in terms of a collection of hyperparameters that must be chosen or learned. Fully Bayesian GPR seeks to infer these kernel hyperparameters in a Bayesian sense, and the key computational challenge in sampling from their posterior distribution is the need for frequent determinant evaluations of large kernel matrices. This paper introduces a gradient-based, determinant-free approach for fully Bayesian GPR that combines a Gaussian integration trick for avoiding the determinant with Hamiltonian Monte Carlo (HMC) sampling. Our framework permits a matrix-free formulation and reduces the difficulty of dealing with hyperparameter gradients to a simple automatic differentiation. Our implementation is highly flexible and leverages GPU acceleration with linear-scaling memory footprint. Numerical experiments demonstrate the method's ability to scale gracefully to both high-dimensional hyperparameter spaces and large kernel matrices.

math.NA

Fast and spectrally accurate construction of adaptive diagonal basis sets for electronic structure

In this article, we combine the periodic sinc basis set with a curvilinear coordinate system for electronic structure calculations. This extension allows for variable resolution across the computational domain, with higher resolution close to the nuclei and lower resolution in the inter-atomic regions. We address two key challenges that arise while using basis sets obtained by such a coordinate transformation. First, we use pseudospectral methods to evaluate the integrals needed to construct the Hamiltonian in this basis. Second, we demonstrate how to construct an appropriate coordinate transformation by solving the Monge-Amp\`ere equation using a new approach that we call the cyclic Knothe-Rosenblatt flow. The solution of both of these challenges enables mean-field calculations at a cost that is log-linear in the number of basis functions. We demonstrate that our method approaches the complete basis set limit faster than basis sets with uniform resolution. We also emphasize how these basis sets satisfy the diagonal approximation, which is shown to be a consequence of the pseudospectral method. The diagonal approximation is highly desirable for the solution of the electronic structure problem in many frameworks, including mean field theories, tensor network methods, quantum computing, and quantum Monte Carlo.

physics.comp-ph

Gaussian process regression with log-linear scaling for common non-stationary kernels

We introduce a fast algorithm for Gaussian process regression in low dimensions, applicable to a widely-used family of non-stationary kernels. The non-stationarity of these kernels is induced by arbitrary spatially-varying vertical and horizontal scales. In particular, any stationary kernel can be accommodated as a special case, and we focus especially on the generalization of the standard Mat\'ern kernel. Our subroutine for kernel matrix-vector multiplications scales almost optimally as $O(N\log N)$, where $N$ is the number of regression points. Like the recently developed equispaced Fourier Gaussian process (EFGP) methodology, which is applicable only to stationary kernels, our approach exploits non-uniform fast Fourier transforms (NUFFTs). We offer a complete analysis controlling the approximation error of our method, and we validate the method's practical performance with numerical experiments. In particular we demonstrate improved scalability compared to to state-of-the-art rank-structured approaches in spatial dimension $d>1$.

math.NA

Column and row subset selection using nuclear scores: algorithms and theory for Nystr\"{o}m approximation, CUR decomposition, and graph Laplacian reduction

Column selection is an essential tool for structure-preserving low-rank approximation, with wide-ranging applications across many fields, such as data science, machine learning, and theoretical chemistry. In this work, we develop unified methodologies for fast, efficient, and theoretically guaranteed column selection. First we derive and implement a sparsity-exploiting deterministic algorithm applicable to tasks including kernel approximation and CUR decomposition. Next, we develop a matrix-free formalism relying on a randomization scheme satisfying guaranteed concentration bounds, applying this construction both to CUR decomposition and to the approximation of matrix functions of graph Laplacians. Importantly, the randomization is only relevant for the computation of the scores that we use for column selection, not the selection itself given these scores. For both deterministic and matrix-free algorithms, we bound the performance favorably relative to the expected performance of determinantal point process (DPP) sampling and, in select scenarios, that of exactly optimal subset selection. The general case requires new analysis of the DPP expectation. Finally, we demonstrate strong real-world performance of our algorithms on a diverse set of example approximation tasks.

math.NA

Large Stepsize Gradient Descent for Non-Homogeneous Two-Layer Networks: Margin Improvement and Fast Optimization

The typical training of neural networks using large stepsize gradient descent (GD) under the logistic loss often involves two distinct phases, where the empirical risk oscillates in the first phase but decreases monotonically in the second phase. We investigate this phenomenon in two-layer networks that satisfy a near-homogeneity condition. We show that the second phase begins once the empirical risk falls below a certain threshold, dependent on the stepsize. Additionally, we show that the normalized margin grows nearly monotonically in the second phase, demonstrating an implicit bias of GD in training non-homogeneous predictors. If the dataset is linearly separable and the derivative of the activation function is bounded away from zero, we show that the average empirical risk decreases, implying that the first phase must stop in finite steps. Finally, we demonstrate that by choosing a suitably large stepsize, GD that undergoes this phase transition is more efficient than GD that monotonically decreases the risk. Our analysis applies to networks of any width, beyond the well-known neural tangent kernel and mean-field regimes.

stat.ML

Multiscale interpolative construction of quantized tensor trains

Quantized tensor trains (QTTs) have recently emerged as a framework for the numerical discretization of continuous functions, with the potential for widespread applications in numerical analysis. However, the theory of QTT approximation is not fully understood. In this work, we advance this theory from the point of view of multiscale polynomial interpolation. This perspective clarifies why QTT ranks decay with increasing depth, quantitatively controls QTT rank in terms of smoothness of the target function, and explains why certain functions with sharp features and poor quantitative smoothness can still be well approximated by QTTs. The perspective also motivates new practical and efficient algorithms for the construction of QTTs from function evaluations on multiresolution grids.

math.NA