SearcharxivSearch

arXiv subjects

Mark K. Transtrum

Publications and source records attributed to Mark K. Transtrum.

At least 19 recordsLinked to original sources

Comparative study of ensemble-based uncertainty quantification methods for neural network interatomic potentials

Machine learning interatomic potentials (MLIPs) enable atomistic simulations with near first-principles accuracy at substantially reduced computational cost, making them powerful tools for large-scale materials modeling. The accuracy of MLIPs is typically validated on a held-out dataset of \emph{ab initio} energies and atomic forces. However, accuracy on these small-scale properties does not guarantee reliability for emergent, system-level behavior -- precisely the regime where atomistic simulations are most needed, but for which direct validation is often computationally prohibitive. As a practical heuristic, predictive precision -- quantified as inverse uncertainty -- is commonly used as a proxy for accuracy, but its reliability remains poorly understood, particularly for system-level predictions. In this work, we systematically assess the relationship between predictive precision and accuracy in both in-distribution (ID) and out-of-distribution (OOD) regimes, focusing on ensemble-based uncertainty quantification methods for neural network potentials, including bootstrap, dropout, random initialization, and snapshot ensembles. We use held-out cross-validation for ID assessment and calculate cold curve energies and phonon dispersion relations for OOD testing. These evaluations are performed across various carbon allotropes as representative test systems. We find that uncertainty estimates can behave counterintuitively in OOD settings, often plateauing or even decreasing as predictive errors grow. These results highlight fundamental limitations of current uncertainty quantification approaches and underscore the need for caution when using predictive precision as a stand-in for accuracy in large-scale, extrapolative applications.

cond-mat.mtrl-sci

Inverse design of bespoke interatomic potentials via active learning by information-matching

Interatomic potentials (IPs) enable large-scale atomistic simulations beyond the reach of first-principles methods, but their predictive reliability depends critically on the selection of training data, quantified uncertainty, and model expressiveness. Active learning (AL) provides a principled framework for constructing efficient and accurate IPs, yet most strategies reduce parameter uncertainty without explicitly accounting for the specific material properties being predicted. The information-matching (IM) approach addresses this limitation by requiring that the selected training data provide at least as much parameter space information as needed to achieve prescribed uncertainty targets for selected quantities of interest (QoIs). Here, we apply IM to develop bespoke IPs specifically tailored for predicting plastic strength in metals. Due to the high computational cost of simulating plastic strength, we employ an indirect IM strategy that targets inexpensive intermediate QoIs that correlate with strength. The IM method enables precise parameter constraints with minimal training data, yielding precise predictions for both the intermediate QoIs and plastic strength. Yet, model error remains a key limitation, and a post hoc uncertainty inflation correction provides a viable means to mitigate this limitation. These findings illustrate both the promise and limits of uncertainty-aware AL for predicting complex material properties.

cond-mat.mtrl-sci

Composable and adaptive design of machine learning interatomic potentials guided by Fisher-information analysis

An adaptive physics-inspired model design strategy for machine-learning interatomic potentials (MLIPs) is proposed. This strategy relies on iterative reconfigurations of composite models from single-term models, followed by a unified training procedure. A model evaluation method based on the Fisher information matrix (FIM) and multiple-property error metrics is also proposed to guide the model reconfiguration and hyperparameter optimization. By combining the reconfiguration and the evaluation subroutines, we provide an adaptive MLIP design strategy that balances flexibility and extensibility. In a case study of designing models against a structurally diverse niobium dataset, we managed to obtain an optimal model configuration with 75 parameters generated by our framework that achieved a force RMSE of 0.172 eV/Å and an energy RMSE of 0.013 eV/atom.

cond-mat.mtrl-sci

An information-matching approach to optimal experimental design and active learning

The efficacy of mathematical models heavily depends on the quality of the training data, yet collecting sufficient data is often expensive and challenging. Many modeling applications require inferring parameters only as a means to predict other quantities of interest (QoI). Because models often contain many unidentifiable (sloppy) parameters, QoIs often depend on a relatively small number of parameter combinations. Therefore, we introduce an information-matching criterion based on the Fisher Information Matrix to select the most informative training data from a candidate pool. This method ensures that the selected data contain sufficient information to learn only those parameters that are needed to constrain downstream QoIs. It is formulated as a convex optimization problem, making it scalable to large models and datasets. We demonstrate the effectiveness of this approach across various modeling problems in diverse scientific fields, including power systems and underwater acoustics. Finally, we use information-matching as a query function within an Active Learning loop for material science applications. In all these applications, we find that a relatively small set of optimal training data can provide the necessary information for achieving precise predictions. These results are encouraging for diverse future applications, particularly active learning in large machine learning models.

cs.LG

An Analytical Characterization of Sloppiness in Neural Networks: Insights from Linear Models

Recent experiments have shown that training trajectories of multiple deep neural networks with different architectures, optimization algorithms, hyper-parameter settings, and regularization methods evolve on a remarkably low-dimensional "hyper-ribbon-like" manifold in the space of probability distributions. Inspired by the similarities in the training trajectories of deep networks and linear networks, we analytically characterize this phenomenon for the latter. We show, using tools in dynamical systems theory, that the geometry of this low-dimensional manifold is controlled by (i) the decay rate of the eigenvalues of the input correlation matrix of the training data, (ii) the relative scale of the ground-truth output to the weights at the beginning of training, and (iii) the number of steps of gradient descent. By analytically computing and bounding the contributions of these quantities, we characterize phase boundaries of the region where hyper-ribbons are to be expected. We also extend our analysis to kernel machines and linear models that are trained with stochastic gradient descent.

cs.LG

A Time-Dependent Ginzburg-Landau Framework for Sample-Specific Simulation of Superconductors for SRF Applications

Modern superconducting radio frequency (SRF) applications demand precise control over material properties across multiple length scales - from microscopic composition, to mesoscopic defect structures, to macroscopic cavity geometry. We present a time-dependent Ginzburg-Landau (TDGL) framework that incorporates spatially varying parameters derived from experimental measurements and ab initio calculations, enabling realistic, sample-specific simulations. As a demonstration, we model Sn-deficient islands in Nb$_3$Sn and calculate the field at which vortex nucleation first occurs for various defect configurations. These thresholds serve as a predictive tool for identifying defects likely to degrade SRF cavity performance. We then simulate the resulting dissipation and show how aggregate contributions from multiple small defects can reproduce trends consistent with high-field $Q$-slope behavior observed experimentally. Our results offer a pathway for connecting microscopic defect properties to macroscopic SRF performance using a computationally efficient mesoscopic model.

cond-mat.supr-con

A CRISP approach to QSP: XAI enabling fit-for-purpose models

Quantitative Systems Pharmacology (QSP) promises to accelerate drug development, enable personalized medicine, and improve the predictability of clinical outcomes. Realizing this potential requires effectively managing the complexity of mathematical models representing biological systems. Here, we present and validate a novel QSP workflow--CRISP (Contextualized Reduction for Identifiability and Scientific Precision)--that addresses a central challenge in QSP: the problem of complexity and over-parameterization, in which models contain irrelevant parameters that obscure interpretation and hinder predictive reliability. The CRISP workflow begins with a literature-derived model, constructed to be comprehensive and unbiased by integrating prior mechanistic insights. At the core of the workflow is the Manifold Boundary Approximation Method (MBAM), a reduction technique that simplifies models while preserving mechanistic structure and predictive fidelity. By applying MBAM in a context-specific manner, CRISP links parsimonious models directly to predictions of interest, clarifying causal structure and enhancing interpretability. The resulting models are computationally efficient and well-suited to key QSP tasks, including virtual population generation, experimental design, toxicology, and target discovery. We demonstrate the utility of CRISP on case studies involving the coagulation cascade and SHIV infection, and identify promising directions for improving the efficacy of bNAb therapies for HIV. Together, these results establish CRISP as a general-purpose QSP workflow for turning complex mechanistic models into tools for precise scientific reasoning to guide pharmacological and regulatory decision-making.

q-bio.QM

eGAD! double descent is explained by Generalized Aliasing Decomposition

A central problem in data science is to use potentially noisy samples of an unknown function to predict values for unseen inputs. In classical statistics, predictive error is understood as a trade-off between the bias and the variance that balances model simplicity with its ability to fit complex functions. However, over-parameterized models exhibit counterintuitive behaviors, such as "double descent" in which models of increasing complexity exhibit decreasing generalization error. Others may exhibit more complicated patterns of predictive error with multiple peaks and valleys. Neither double descent nor multiple descent phenomena are well explained by the bias-variance decomposition. We introduce a novel decomposition that we call the generalized aliasing decomposition (GAD) to explain the relationship between predictive performance and model complexity. The GAD decomposes the predictive error into three parts: 1) model insufficiency, which dominates when the number of parameters is much smaller than the number of data points, 2) data insufficiency, which dominates when the number of parameters is much greater than the number of data points, and 3) generalized aliasing, which dominates between these two extremes. We demonstrate the applicability of the GAD to diverse applications, including random feature models from machine learning, Fourier transforms from signal processing, solution methods for differential equations, and predictive formation enthalpy in materials discovery. Because key components of the GAD can be explicitly calculated from the relationship between model class and samples without seeing any data labels, it can answer questions related to experimental design and model selection before collecting data or performing experiments. We further demonstrate this approach on several examples and discuss implications for predictive modeling and data science.

math.ST

Comparing analytic and data-driven approaches to parameter identifiability: A power systems case study

Parameter identifiability refers to the capability of accurately inferring the parameter values of a model from its observations (data). Traditional analysis methods exploit analytical properties of the closed form model, in particular sensitivity analysis, to quantify the response of the model predictions to variations in parameters. Techniques developed to analyze data, specifically manifold learning methods, have the potential to complement, and even extend the scope of the traditional analytical approaches. We report on a study comparing and contrasting analytical and data-driven approaches to quantify parameter identifiability and, importantly, perform parameter reduction tasks. We use the infinite bus synchronous generator model, a well-understood model from the power systems domain, as our benchmark problem. Our traditional analysis methods use the Fisher Information Matrix to quantify parameter identifiability analysis, and the Manifold Boundary Approximation Method to perform parameter reduction. We compare these results to those arrived at through data-driven manifold learning schemes: Output - Diffusion Maps and Geometric Harmonics. For our test case, we find that the two suites of tools (analytical when a model is explicitly available, as well as data-driven when the model is lacking and only measurement data are available) give (correct) comparable results; these results are also in agreement with traditional analysis based on singular perturbation theory. We then discuss the prospects of using data-driven methods for such model analysis.

cs.LG

The Training Process of Many Deep Networks Explores the Same Low-Dimensional Manifold

We develop information-geometric techniques to analyze the trajectories of the predictions of deep networks during training. By examining the underlying high-dimensional probabilistic models, we reveal that the training process explores an effectively low-dimensional manifold. Networks with a wide range of architectures, sizes, trained using different optimization methods, regularization techniques, data augmentation techniques, and weight initializations lie on the same manifold in the prediction space. We study the details of this manifold to find that networks with different architectures follow distinguishable trajectories but other factors have a minimal influence; larger networks train along a similar manifold as that of smaller networks, just faster; and networks initialized at very different parts of the prediction space converge to the solution along a similar manifold.

cs.LG

Sloppy model analysis identifies bifurcation parameters without Normal Form analysis

Bifurcation phenomena are common in multi-dimensional multi-parameter dynamical systems. Normal form theory suggests that the bifurcations themselves are driven by relatively few parameters; however, these are often nonlinear combinations of the bare parameters in which the equations are expressed. Discovering reparameterizations to transform such complex original equations into normal-form is often very difficult, and the reparameterization may not even exist in a closed-form. Recent advancements have tied both information geometry and bifurcations to the Renormalization Group. Here, we show that sloppy model analysis (a method of information geometry) can be used directly on bifurcations of increasing time scales to rapidly characterize the system's topological inhomogeneities, whether the system is in normal form or not. We anticipate that this novel analytical method, which we call time-widening information geometry (TWIG), will be useful in applied network analysis.

math.DS

Emergent super-antiferromagnetic correlations in monolayers of Fe3O4 nanoparticles throughout the superparamagnetic blocking transition

We report nanoscale inter-particle magnetic orderings in self-assemblies of Fe3O4 nanoparticles (NPs), and the emergence of inter-particle antiferromagnetic (AF) (super-antiferromagnetic) correlations near the coercive field at low temperature. The magnetic ordering is probed via x-ray resonant magnetic scattering (XRMS), with the x-ray energy tuned to the Fe-L3 edge and using circular polarized light. By exploiting dichroic effects, a magnetic scattering signal is isolated from the charge scattering signal. The magnetic signal informs about nanoscale spatial orderings at various stages throughout the magnetization process and at various temperatures throughout the superparamagnetic blocking transition, for two different sizes of NPs, 5 and 11 nm, with blocking temperatures TB of 28 K and 170 K, respectively. At 300 K, while the magnetometry data essentially shows superparamagnetism and absence of hysteresis for both particle sizes, the XRMS data reveals the presence of non-zero (up to 9/100) inter-particle AF couplings when the applied field is released to zero for the 11 nm NPs. These AF couplings are drastically amplified when the NPs are cooled down below TB and reach up to 12/100 for the 5 nm NPs and 48/100 for the 11 nm NPs, near the coercive point. The data suggests that the particle size affects the prevalence of the AF couplings: compared to ferromagnetic (F) couplings, the relative prevalence of AF couplings at the coercive point increases from a factor ~ 1.6 to 3.8 when the NP size increases from 5 to 11 nm.

cond-mat.mes-hall

ZrNb(CO) RF superconducting thin film with high critical temperature in the theoretical limit

Superconducting radio-frequency (SRF) resonators are critical components for particle accelerator applications, such as free-electron lasers, and for emerging technologies in quantum computing. Developing advanced materials and their deposition processes to produce RF superconductors that yield nanoohms surface resistances is a key metric for the wider adoption of SRF technology. Here we report ZrNb(CO) RF superconducting films with high critical temperatures (Tc) achieved for the first time under ambient pressure. The attainment of a Tc near the theoretical limit for this material without applied pressure is promising for its use in practical applications. A range of Tc, likely arising from Zr doping variation, may allow a tunable superconducting coherence length that lowers the sensitivity to material defects when an ultra-low surface resistance is required. Our ZrNb(CO) films are synthesized using a low-temperature (100 - 200 C) electrochemical recipe combined with thermal annealing. The phase transformation as a function of annealing temperature and time is optimized by the evaporated Zr-Nb diffusion couples. Through phase control, we avoid hexagonal Zr phases that are equilibrium-stable but degrade Tc. X-ray and electron diffraction combined with photoelectron spectroscopy reveal a system containing cubic ZrNb mixed with rocksalt NbC and low-dielectric-loss ZrO2. We demonstrate proof-of-concept RF performance of ZrNb(CO) on an SRF sample test system. BCS resistance trends lower than reference Nb, while quench fields occur at approximately 35 mT. Our results demonstrate the potential of ZrNb(CO) thin films for particle accelerator and other SRF applications.

cond-mat.mtrl-sci

Information geometry for multiparameter models: New perspectives on the origin of simplicity

Complex models in physics, biology, economics, and engineering are often sloppy, meaning that the model parameters are not well determined by the model predictions for collective behavior. Many parameter combinations can vary over decades without significant changes in the predictions. This review uses information geometry to explore sloppiness and its deep relation to emergent theories. We introduce the model manifold of predictions, whose coordinates are the model parameters. Its hyperribbon structure explains why only a few parameter combinations matter for the behavior. We review recent rigorous results that connect the hierarchy of hyperribbon widths to approximation theory, and to the smoothness of model predictions under changes of the control variables. We discuss recent geodesic methods to find simpler models on nearby boundaries of the model manifold -- emergent theories with fewer parameters that explain the behavior equally well. We discuss a Bayesian prior which optimizes the mutual information between model parameters and experimental data, naturally favoring points on the emergent boundary theories and thus simpler models. We introduce a `projected maximum likelihood' prior that efficiently approximates this optimal prior, and contrast both to the poor behavior of the traditional Jeffreys prior. We discuss the way the renormalization group coarse-graining in statistical mechanics introduces a flow of the model manifold, and connect stiff and sloppy directions along the model manifold with relevant and irrelevant eigendirections of the renormalization group. Finally, we discuss recently developed `intensive' embedding methods, allowing one to visualize the predictions of arbitrary probabilistic models as low-dimensional projections of an isometric embedding, and illustrate our method by generating the model manifold of the Ising model.

cond-mat.stat-mech

Theory of Nb-Zr Alloy Superconductivity and First Experimental Demonstration for Superconducting Radio-Frequency Cavity Applications

Niobium-zirconium (Nb-Zr) alloy is an old superconductor that is a promising new candidate for superconducting radio-frequency (SRF) cavity applications. Using density-functional and Eliashberg theories, we show that addition of Zr to a Nb surface in small concentrations increases the critical temperature $T_c$ and improves other superconducting properties. Furthermore, we calculate $T_c$ for Nb-Zr alloys across a broad range of Zr concentrations, showing good agreement with the literature for disordered alloys as well as the potential for significantly higher $T_c$ in ordered alloys near 75%Nb/25%Zr composition. We provide experimental verification on Nb-Zr alloy samples and SRF sample test cavities prepared with either physical vapor or our novel electrochemical deposition recipes. These samples have the highest measured $T_c$ of any Nb-Zr superconductor to date and indicate a reduction in BCS resistance compared to the conventional Nb reference sample; they represent the first steps along a new pathway to greatly enhanced SRF performance. Finally, we use Ginzburg-Landau theory to show that the addition of Zr to a Nb surface increases the superheating field $B_{sh}$, a key figure of merit for SRF which determines the maximum accelerating gradient at which cavities can operate.

cond-mat.supr-con

Extending OpenKIM with an Uncertainty Quantification Toolkit for Molecular Modeling

Atomistic simulations are an important tool in materials modeling. Interatomic potentials (IPs) are at the heart of such molecular models, and the accuracy of a model's predictions depends strongly on the choice of IP. Uncertainty quantification (UQ) is an emerging tool for assessing the reliability of atomistic simulations. The Open Knowledgebase of Interatomic Models (OpenKIM) is a cyberinfrastructure project whose goal is to collect and standardize the study of IPs to enable transparent, reproducible research. Part of the OpenKIM framework is the Python package, KIM-based Learning-Integrated Fitting Framework (KLIFF), that provides tools for fitting parameters in an IP to data. This paper introduces a UQ toolbox extension to KLIFF. We focus on two sources of uncertainty: variations in parameters and inadequacy of the functional form of the IP. Our implementation uses parallel-tempered Markov chain Monte Carlo (PTMCMC), adjusting the sampling temperature to estimate the uncertainty due to the functional form of the IP. We demonstrate on a Stillinger--Weber potential that makes predictions for the atomic energies and forces for silicon in a diamond configuration. Finally, we highlight some potential subtleties in applying and using these tools with recommendations for practitioners and IP developers.

physics.comp-ph

Bayesian, frequentist, and information geometric approaches to parametric uncertainty quantification of classical empirical interatomic potentials

In this paper, we consider the problem of quantifying parametric uncertainty in classical empirical interatomic potentials (IPs) using both Bayesian (Markov Chain Monte Carlo) and frequentist (profile likelihood) methods. We interface these tools with the Open Knowledgebase of Interatomic Models and study three models based on the Lennard-Jones, Morse, and Stillinger--Weber potentials. We confirm that IPs are typically sloppy, i.e., insensitive to coordinated changes in some parameter combinations. Because the inverse problem in such models is ill-conditioned, parameters are unidentifiable. This presents challenges for traditional statistical methods, as we demonstrate and interpret within both Bayesian and frequentist frameworks. We use information geometry to illuminate the underlying cause of this phenomenon and show that IPs have global properties similar to those of sloppy models from fields such as systems biology, power systems, and critical phenomena. IPs correspond to bounded manifolds with a hierarchy of widths, leading to low effective dimensionality in the model. We show how information geometry can motivate new, natural parameterizations that improve the stability and interpretation of uncertainty quantification analysis and further suggest simplified, less-sloppy models.

cond-mat.mtrl-sci

The supremum principle selects simple, transferable models

We consider how mathematical models enable predictions for conditions that are qualitatively different from the training data. We propose techniques based on information topology to find models that can apply their learning in regimes for which there is no data. The first step is to use the Manifold Boundary Approximation Method to construct simple, reduced models of target phenomena in a data-driven way. We consider the set of all such reduced models and use the topological relationships among them to reason about model selection for new, unobserved phenomena. Given minimal models for several target behaviors, we introduce the supremum principle as a criterion for selecting a new, transferable model. The supremal model, i.e., the least upper bound, is the simplest model that reduces to each of the target behaviors. We illustrate how to discover supremal models with several examples; in each case, the supremal model unifies causal mechanisms to transfer successfully to new target domains. We use these examples to motivate a general algorithm that has formal connections to theories of analogical reasoning in cognitive psychology.

q-bio.QM