SearcharxivSearch

arXiv subjects

Jonas Köhler

Publications and source records attributed to Jonas Köhler.

At least 19 recordsLinked to original sources

An ab initio foundation model of wavefunctions that accurately describes chemical bond breaking

Reliable description of bond breaking remains a major challenge for quantum chemistry due to the multireference character of the electronic structure in dissociating species. Multireference methods in particular suffer from large computational cost, which under the normal paradigm has to be paid anew for each system at a full price, ignoring commonalities in electronic structure across molecules. Quantum Monte Carlo with deep neural networks uniquely offers to exploit such commonalities by pretraining transferable wavefunction models, but all such attempts were so far limited in scope. Here, we bring this paradigm to fruition with Orbformer, a transferable wavefunction model pretrained on 22,000 equilibrium and dissociating structures that can be fine-tuned on unseen molecules reaching an accuracy-cost ratio rivalling classical multireference methods. On established benchmarks as well as more challenging bond dissociations and Diels-Alder reactions, Orbformer is the only method that consistently converges to chemical accuracy (1 kcal/mol). This work turns the idea of amortizing the cost of solving the Schrödinger equation over many molecules into a practical approach in quantum chemistry.

physics.chem-ph

Efficient higher-order local time integration for Friedrichs' systems

In this paper, we construct an efficient higher-order local time integration scheme for spatially discretized linear Friedrichs' systems. In particular, our interest is in problems where only a few of the mesh elements are small while the majority of the elements is much larger. The special combination of two methods like the leapfrog method on the coarse part of the mesh and the Crank-Nicolson method on the fine part as was done in Hochbruck, Sturm 2016 and Hochbruck, Köhler 2022 is not suitable for higher-order time integration. Therefore, we suggest to approximate the solution of the linear systems arising in each time step by a preconditioned Krylov subspace method, e.g., the quasi-minimal residual method by Freund and Nachtigal 1991. The techniques developed here for linear problems also carry over to nonlinear problems, where linear systems of the same type arise within a Newton-type iteration. Motivated by the analysis of locally implicit methods by Hochbruck and Sturm 2016, we show how to construct a preconditioner in such a way that the number of iterations required by the Krylov subspace method to achieve a certain accuracy is bounded independently of the diameter of the small mesh elements. We prove this behavior by using Faber polynomials and complex approximation theory. The cost to apply the preconditioner consists of the solution of a small linear system, whose dimension corresponds to the degrees of freedom within the fine part of the mesh (and its next coarse neighbors). If this dimension is small compared to the size of the full mesh, the preconditioner is very efficient. We conclude by verifying our theoretical results with numerical experiments for the fourth-order Gauss-Legendre Runge--Kutta method.

math.NA

Probabilistic and Alarm-Based Evaluation of a b-Value-Driven Deep Learning Earthquake Forecast

We evaluate the forecasting performance of a deep learning model, originally introduced as a pattern-extraction framework, that operates on the spatiotemporal evolution of seismic b-values in a short-term forecasting context. Model output is rescaled to account for training on balanced datasets and evaluated relative to a spatial base-rate model using the Brier Skill Score (BSS). Absolute skill values are small, but mean BSS values are consistently positive, including at locations where Mw geq 5 earthquakes occurred during the test period, indicating information content beyond historical seismicity alone. Alarm-based evaluation using Molchan diagrams shows elevated event capture rates at low alarm fractions (5.88 percent of events captured at 1 percent area under alarm), indicating discrimination exceeding random and purely spatial reference models under constrained alarm conditions. Comparison with ETAS-derived triggered probabilities further reveals a weak positive correlation, suggesting partial sensitivity of the model output to seismic regimes characterized by enhanced clustering and recent activity, while remaining distinct from classical aftershock-based descriptions. Together, these results indicate that spatiotemporal variations in b-values contain a persistent, though limited, signal relevant to probabilistic earthquake forecasting, yielding marginal but consistent improvements over baseline models across complementary evaluation frameworks.

physics.geo-ph

Detecting Spatiotemporal b-Value Anomalies with a Progressive Deep Learning Architecture

Identifying systematic patterns in seismicity that precede large earthquakes remains a central challenge in statistical seismology. In this work, we present a methodological framework for detecting spatiotemporal anomalies in seismicity using the evolution of gridded b-values. Focusing on the Japanese subduction zone, we construct daily b-value fields on a fine spatial grid by aggregating local seismicity over moving time windows, yielding a continuous 2+1D representation of seismic-state evolution. We formulate the problem as a binary classification task in which spatiotemporal blocks extracted from these $b$-value fields are labeled according to the occurrence of a target earthquake with \Mw $\geq 5$ in the central region within the next day. To model this data, we introduce a hybrid deep-learning architecture that combines a spatial convolutional encoder with a temporal convolutional network, enabling joint learning of spatial structure and temporal dynamics. A progressive meta-epoch training scheme is employed, in which the model is iteratively updated using a time-forward strategy that mirrors operational deployment and mitigates issues related to nonstationarity. This paper is strictly methodological in scope. It describes the construction of b-value fields, the spatiotemporal sampling strategy, the network architecture, and the progressive training and internal validation framework used for model development and parameter selection.

physics.geo-ph

Altermagnetic magnon transport in the \textit{d}-wave altermagnet \ch{LuFeO3}

Altermagnets exhibit a spin-split band structure despite having zero net magnetization, leading to special magnonic properties such as anisotropic magnon lifetimes and field-free spin transport. Here, we present a direct experimental demonstration of non-local magnon transport in the \textit{d}-wave altermagnet \ch{LuFeO3}, using both spin Seebeck and spin Hall effect-based injection and detection. We observe a non-local spin signal at zero magnetic field when the transport is along an altermagnetic direction, but not for transport along other directions. The observed sign reversal between two distinct altermagnetic directions in the spin Seebeck response demonstrates the altermagnetic nature of the magnon transport. In contrast, when transport is aligned along or perpendicular to the easy axis, both the first-harmonic signal and the sign-reversal effect vanish, consistent with symmetry-imposed suppression. These findings are supported by atomistic spin dynamics simulations, as well as linear spin wave theory calculations, which explain how our altermagnetic system hosts anisotropic spin Seebeck transport. Our results provide direct evidence of direction-dependent magnon splitting in altermagnets and highlight their potential for field-free magnonic spin transport, offering a promising pathway for low-power spintronic applications.

cond-mat.mtrl-sci

A Deep Learning Pipeline for Large Earthquake Analysis using High-Rate Global Navigation Satellite System Data

Deep learning techniques for processing large and complex datasets have unlocked new opportunities for fast and reliable earthquake analysis using Global Navigation Satellite System (GNSS) data. This work presents a deep learning model, MagEs, to estimate earthquake magnitudes using data from high-rate GNSS stations. Furthermore, MagEs is integrated with the DetEQ model for earthquake detection within the SAIPy package, creating a comprehensive pipeline for earthquake detection and magnitude estimation using HR-GNSS data. The MagEs model provides magnitude estimates within seconds of detection when using stations within 3 degrees of the epicenter, which are the most relevant for real-time applications. However, since it has been trained on data from stations up to 7.5 degrees away, it can also analyze data from larger distances. The model can process data from a single station at a time or combine data from up to three stations. The model was trained using synthetic data reflecting rupture scenarios in the Chile subduction zone, and the results confirm strong performance for Chilean earthquakes. Although tests from other tectonic regions also yielded good results, incorporating regional data through transfer learning could further improve its performance in diverse seismic settings. The model has not yet been deployed in an operational real-time monitoring system, but simulation tests that update data in a second-by-second manner demonstrate its potential for future real-time adaptation.

physics.geo-ph

Highly Accurate Real-space Electron Densities with Neural Networks

Variational ab-initio methods in quantum chemistry stand out among other methods in providing direct access to the wave function. This allows in principle straightforward extraction of any other observable of interest, besides the energy, but in practice this extraction is often technically difficult and computationally impractical. Here, we consider the electron density as a central observable in quantum chemistry and introduce a novel method to obtain accurate densities from real-space many-electron wave functions by representing the density with a neural network that captures known asymptotic properties and is trained from the wave function by score matching and noise-contrastive estimation. We use variational quantum Monte Carlo with deep-learning ansätze (deep QMC) to obtain highly accurate wave functions free of basis set errors, and from them, using our novel method, correspondingly accurate electron densities, which we demonstrate by calculating dipole moments, nuclear forces, contact densities, and other density-based properties.

physics.chem-ph

Rigid Body Flows for Sampling Molecular Crystal Structures

Normalizing flows (NF) are a class of powerful generative models that have gained popularity in recent years due to their ability to model complex distributions with high flexibility and expressiveness. In this work, we introduce a new type of normalizing flow that is tailored for modeling positions and orientations of multiple objects in three-dimensional space, such as molecules in a crystal. Our approach is based on two key ideas: first, we define smooth and expressive flows on the group of unit quaternions, which allows us to capture the continuous rotational motion of rigid bodies; second, we use the double cover property of unit quaternions to define a proper density on the rotation group. This ensures that our model can be trained using standard likelihood-based methods or variational inference with respect to a thermodynamic target density. We evaluate the method by training Boltzmann generators for two molecular examples, namely the multi-modal density of a tetrahedral system in an external field and the ice XI phase in the TIP4P water model. Our flows can be combined with flows operating on the internal degrees of freedom of molecules and constitute an important step towards the modeling of distributions of many interacting molecules.

cs.LG

Flow-matching -- efficient coarse-graining of molecular dynamics without forces

Coarse-grained (CG) molecular simulations have become a standard tool to study molecular processes on time- and length-scales inaccessible to all-atom simulations. Parameterizing CG force fields to match all-atom simulations has mainly relied on force-matching or relative entropy minimization, which require many samples from costly simulations with all-atom or CG resolutions, respectively. Here we present flow-matching, a new training method for CG force fields that combines the advantages of both methods by leveraging normalizing flows, a generative deep learning method. Flow-matching first trains a normalizing flow to represent the CG probability density, which is equivalent to minimizing the relative entropy without requiring iterative CG simulations. Subsequently, the flow generates samples and forces according to the learned distribution in order to train the desired CG free energy model via force matching. Even without requiring forces from the all-atom simulations, flow-matching outperforms classical force-matching by an order of magnitude in terms of data efficiency, and produces CG models that can capture the folding and unfolding transitions of small proteins.

physics.comp-ph

Smooth Normalizing Flows

Normalizing flows are a promising tool for modeling probability distributions in physical systems. While state-of-the-art flows accurately approximate distributions and energies, applications in physics additionally require smooth energies to compute forces and higher-order derivatives. Furthermore, such densities are often defined on non-trivial topologies. A recent example are Boltzmann Generators for generating 3D-structures of peptides and small proteins. These generative models leverage the space of internal coordinates (dihedrals, angles, and bonds), which is a product of hypertori and compact intervals. In this work, we introduce a class of smooth mixture transformations working on both compact intervals and hypertori. Mixture transformations employ root-finding methods to invert them in practice, which has so far prevented bi-directional flow training. To this end, we show that parameter gradients and forces of such inverses can be computed from forward evaluations via the inverse function theorem. We demonstrate two advantages of such smooth flows: they allow training by force matching to simulation data and can be used as potentials in molecular dynamics simulations.

stat.ML

Generating stable molecules using imitation and reinforcement learning

Chemical space is routinely explored by machine learning methods to discover interesting molecules, before time-consuming experimental synthesizing is attempted. However, these methods often rely on a graph representation, ignoring 3D information necessary for determining the stability of the molecules. We propose a reinforcement learning approach for generating molecules in cartesian coordinates allowing for quantum chemical prediction of the stability. To improve sample-efficiency we learn basic chemical rules from imitation learning on the GDB-11 database to create an initial model applicable for all stoichiometries. We then deploy multiple copies of the model conditioned on a specific stoichiometry in a reinforcement learning setting. The models correctly identify low energy molecules in the database and produce novel isomers not found in the training set. Finally, we apply the model to larger molecules to show how reinforcement learning further refines the imitation learning model in domains far from the training data.

physics.chem-ph

Training Invertible Linear Layers through Rank-One Perturbations

Many types of neural network layers rely on matrix properties such as invertibility or orthogonality. Retaining such properties during optimization with gradient-based stochastic optimizers is a challenging task, which is usually addressed by either reparameterization of the affected parameters or by directly optimizing on the manifold. This work presents a novel approach for training invertible linear layers. In lieu of directly optimizing the network parameters, we train rank-one perturbations and add them to the actual weight matrices infrequently. This P$^{4}$Inv update allows keeping track of inverses and determinants without ever explicitly computing them. We show how such invertible blocks improve the mixing and thus the mode separation of the resulting normalizing flows. Furthermore, we outline how the P$^4$ concept can be utilized to retain properties other than invertibility.

stat.ML

Stochastic Normalizing Flows

The sampling of probability distributions specified up to a normalization constant is an important problem in both machine learning and statistical mechanics. While classical stochastic sampling methods such as Markov Chain Monte Carlo (MCMC) or Langevin Dynamics (LD) can suffer from slow mixing times there is a growing interest in using normalizing flows in order to learn the transformation of a simple prior distribution to the given target distribution. Here we propose a generalized and combined approach to sample target densities: Stochastic Normalizing Flows (SNF) -- an arbitrary sequence of deterministic invertible functions and stochastic sampling blocks. We show that stochasticity overcomes expressivity limitations of normalizing flows resulting from the invertibility constraint, whereas trainable transformations between sampling steps improve efficiency of pure MCMC/LD along the flow. By invoking ideas from non-equilibrium statistical mechanics we derive an efficient training procedure by which both the sampler's and the flow's parameters can be optimized end-to-end, and by which we can compute exact importance weights without having to marginalize out the randomness of the stochastic blocks. We illustrate the representational power, sampling efficiency and asymptotic correctness of SNFs on several benchmarks including applications to sampling molecular systems in equilibrium.

stat.ML

Equivariant Flows: Exact Likelihood Generative Learning for Symmetric Densities

Normalizing flows are exact-likelihood generative neural networks which approximately transform samples from a simple prior distribution to samples of the probability distribution of interest. Recent work showed that such generative models can be utilized in statistical mechanics to sample equilibrium states of many-body systems in physics and chemistry. To scale and generalize these results, it is essential that the natural symmetries in the probability density -- in physics defined by the invariances of the target potential -- are built into the flow. We provide a theoretical sufficient criterion showing that the distribution generated by \textit{equivariant} normalizing flows is invariant with respect to these symmetries by design. Furthermore, we propose building blocks for flows which preserve symmetries which are usually found in physical/chemical many-body particle systems. Using benchmark systems motivated from molecular physics, we demonstrate that those symmetry preserving flows can provide better generalization capabilities and sampling efficiency.

stat.ML

DP-MAC: The Differentially Private Method of Auxiliary Coordinates for Deep Learning

Developing a differentially private deep learning algorithm is challenging, due to the difficulty in analyzing the sensitivity of objective functions that are typically used to train deep neural networks. Many existing methods resort to the stochastic gradient descent algorithm and apply a pre-defined sensitivity to the gradients for privatizing weights. However, their slow convergence typically yields a high cumulative privacy loss. Here, we take a different route by employing the method of auxiliary coordinates, which allows us to independently update the weights per layer by optimizing a per-layer objective function. This objective function can be well approximated by a low-order Taylor's expansion, in which sensitivity analysis becomes tractable. We perturb the coefficients of the expansion for privacy, which we optimize using more advanced optimization routines than SGD for faster convergence. We empirically show that our algorithm provides a decent trained model quality under a modest privacy budget.

cs.LG

Equivariant Flows: sampling configurations for multi-body systems with symmetric energies

Flows are exact-likelihood generative neural networks that transform samples from a simple prior distribution to the samples of the probability distribution of interest. Boltzmann Generators (BG) combine flows and statistical mechanics to sample equilibrium states of strongly interacting many-body systems such as proteins with 1000 atoms. In order to scale and generalize these results, it is essential that the natural symmetries of the probability density - in physics defined by the invariances of the energy function - are built into the flow. Here we develop theoretical tools for constructing such equivariant flows and demonstrate that a BG that is equivariant with respect to rotations and particle permutations can generalize to sampling nontrivially new configurations where a nonequivariant BG cannot.

stat.ML

Boltzmann Generators -- Sampling Equilibrium States of Many-Body Systems with Deep Learning

Computing equilibrium states in condensed-matter many-body systems, such as solvated proteins, is a long-standing challenge. Lacking methods for generating statistically independent equilibrium samples in "one shot", vast computational effort is invested for simulating these system in small steps, e.g., using Molecular Dynamics. Combining deep learning and statistical mechanics, we here develop Boltzmann Generators, that are shown to generate unbiased one-shot equilibrium samples of representative condensed matter systems and proteins. Boltzmann Generators use neural networks to learn a coordinate transformation of the complex configurational equilibrium distribution to a distribution that can be easily sampled. Accurate computation of free energy differences and discovery of new configurations are demonstrated, providing a statistical mechanics tool that can avoid rare events during sampling without prior knowledge of reaction coordinates.

stat.ML

Double-slit photoelectron interference in strong-field ionization of the neon dimer

Wave-particle duality is an inherent peculiarity of the quantum world. The double-slit experiment has been frequently used for understanding different aspects of this fundamental concept. The occurrence of interference rests on the lack of which-way information and on the absence of decoherence mechanisms, which could scramble the wave fronts. In this letter, we report on the observation of two-center interference in the molecular frame photoelectron momentum distribution upon ionization of the neon dimer by a strong laser field. Postselection of ions, which were measured in coincidence with electrons, allowed choosing the symmetry of the continuum electronic wave function, leading to observation of both, gerade and ungerade, types of interference.

physics.atom-ph