Searcharxiv⌕ Search

arXiv subjects

Gus L. W. Hart

Publications and source records attributed to Gus L. W. Hart.

At least 19 recordsLinked to original sources

POPSICLE: Benchmark Datasets for Segmentation and Localization in CryoET

Cryo-electron tomography (cryoET) has emerged as a powerful tool in structural and cellular biology by enabling direct visualization of macromolecular structures within intact cells, thereby linking molecular architecture to cellular organization in a native context. Realizing the full potential of cryoET, however, increasingly depends on advances in computational analysis, particularly machine learning (ML), to interpret its complex and information-rich data. Despite rapid progress, ML development for cryoET remains bottlenecked by the lack of standardized, well-annotated benchmarks. Existing evaluations are typically small, task-specific, and are assembled in isolation, limiting robust comparisons across methods. Here, we present POPSICLE, a benchmark suite for cryoET segmentation and macromolecular localization built from the CryoET Data Portal - an open, ML-ready repository of tomographic data, metadata, and annotations. POPSICLE spans eukaryotic and prokaryotic systems, both purified and fully in situ samples, and dense voxel-wise segmentation as well as sparse localization tasks. Built on a living data resource, it can expand as new datasets and annotations become available. Baseline experiments reveal substantial variation in model rankings across tasks, underscoring the need for benchmarks tailored to the unique characteristics of cryoET rather than evaluation practices adapted from adjacent biomedical imaging domains. POPSICLE thus provides an open and extensible foundation for reproducible ML evaluation in cryoET.

eess.IV↗

eGAD! double descent is explained by Generalized Aliasing Decomposition

A central problem in data science is to use potentially noisy samples of an unknown function to predict values for unseen inputs. In classical statistics, predictive error is understood as a trade-off between the bias and the variance that balances model simplicity with its ability to fit complex functions. However, over-parameterized models exhibit counterintuitive behaviors, such as "double descent" in which models of increasing complexity exhibit decreasing generalization error. Others may exhibit more complicated patterns of predictive error with multiple peaks and valleys. Neither double descent nor multiple descent phenomena are well explained by the bias-variance decomposition. We introduce a novel decomposition that we call the generalized aliasing decomposition (GAD) to explain the relationship between predictive performance and model complexity. The GAD decomposes the predictive error into three parts: 1) model insufficiency, which dominates when the number of parameters is much smaller than the number of data points, 2) data insufficiency, which dominates when the number of parameters is much greater than the number of data points, and 3) generalized aliasing, which dominates between these two extremes. We demonstrate the applicability of the GAD to diverse applications, including random feature models from machine learning, Fourier transforms from signal processing, solution methods for differential equations, and predictive formation enthalpy in materials discovery. Because key components of the GAD can be explicitly calculated from the relationship between model class and samples without seeing any data labels, it can answer questions related to experimental design and model selection before collecting data or performing experiments. We further demonstrate this approach on several examples and discuss implications for predictive modeling and data science.

math.ST↗

Describe, Transform, Machine Learning: Feature Engineering for Grain Boundaries and Other Variable-Sized Atom Clusters

Obtaining microscopic structure-property relationships for grain boundaries are challenging because of the complex atomic structures that underlie their behavior. This has led to recent efforts to obtain these relationships with machine learning, but representing a grain boundary structure in a manner suitable for machine learning is not a trivial task. There are three key steps common to property prediction in grain boundaries and other variable-sized atom clustered structures. These are: (1) describe the atomic structure as a feature matrix, (2) transform the variable-sized feature matrices of different structures to a fixed length common to all structures, and (3) apply machine learning to predict properties from the transformed feature matrices. We examine these feature engineering steps to understand how they impact the accuracy of grain boundary energy predictions. A database of over 7000 grain boundaries serves to evaluate the different feature engineering combinations. We also examine how these combination of engineered features provide interpretability, or the ability to extract insightful physics from the obtained structure-property relationships.

cond-mat.mtrl-sci↗

Grain boundary solute segregation across the 5D space of crystallographic character

Solute segregation in materials with grain boundaries (GBs) has emerged as a popular method to thermodynamically stabilize nanocrystalline structures. However, the impact of varied GB crystallographic character on solute segregation has never been thoroughly examined. This work examines Co solute segregation in a dataset of 7272 Al bicrystal GBs that span the 5D space of GB crystallographic character. Considerable attention is paid to verification of the calculations in the diverse and large set of GBs. In addition, the results of this work are favorably validated against similar bicrystal and polycrystal simulations. As with other work, we show that Co atoms exhibit strong segregation to sites in Al GBs and that segregation correlates strongly with GB energy and GB excess volume. Segregation varies smoothly in the 5D crystallographic space but has a complex landscape without an obvious functional form.

cond-mat.mtrl-sci↗

Optimal Routes to Ultrafast Polarization Reversal in Ferroelectric LiNbO3

We use the frozen phonon method to calculate the anharmonic potential energy surface and to model the ultrafast ferroelectric polarization reversal in LiNbO3 driven by intense pulses of THz light. Before stable switching of the polarization occurs, there exists a region of excitation field-strengths where transient switching can occur, as observed experimentally [Physical Review Letters 118, 197601 (2017)]. By varying the excitation frequency from 4 to 20 THz, our modeling suggests that more efficient, permanent polarization switching can occur by directly exciting the soft mode at 7 THz, compared to nonlinear phononic-induced switching driven by exciting a high frequency mode at 18 THz. We also show that neglecting anharmonic coupling pathways in the modeled experiment can lead to significant differences in the modeled switching field strengths.

cond-mat.mtrl-sci↗

Facet and energy predictions in grain boundaries: lattice matching and molecular dynamics

Many material properties can be traced back to properties of their grain boundaries. Grain boundary energy (GBE), as a result, is a key quantity of interest in the analysis and modeling of microstructure. A standard method for calculating grain boundary energy is molecular dynamics (MD); however, on-the-fly MD calculations are not tenable due to the extensive computational time required. Lattice matching (LM) is a reduced-order method for estimating GBE quickly; however, it has only been tested against a relatively limited set of data, and does not have a suitable means for assessing error. In this work, we use the recently published dataset of Homer et al. [1] to assess the performance of LM over the full range of GB space, and to equip LM with a metric for error estimation. LM is used to generate energy estimates, along with predictions of facet morphology, for each of the 7,304 boundaries in the Homer dataset. In keeping with prior work, it is observed that LM predictions of low energy boundaries matches well with MD results. Moreover, there is a good general agreement between LM and MD, and it is apparent that the error scales approximately linearly with the predicted energy value; this makes it possible to establish an empirical estimate on error for future LM calculations. An essential part of the LM method is the faceting relaxation, which corrects the expected energy by convexification across the compact space (S2) of boundary plane orientations. The original Homer dataset did not allow for faceting, but upon extended annealing, it was shown that facet patterns similar to those predicted by LM were emerging.

cond-mat.mtrl-sci↗

Accelerating Training of MLIPs Through Small-Cell Training

While machine-learned interatomic potentials have become a mainstay for modeling materials, designing training sets that lead to robust potentials is challenging. Automated methods, such as active learning and on-the-fly learning, construct reliable training sets, but these processes can be resource-intensive. Current training approaches often use density functional theory (DFT) calculations that have the same cell size as the simulations that the potential is explicitly trained to model. Here, we demonstrate an easy-to-implement small-cell training protocol and use it to model the Zr-H system. This training leads to a potential that accurately predicts known stable Zr-H phases and reproduces the $α$-$β$ pure zirconium phase transition in molecular dynamics simulations. Compared to traditional active learning, small-cell training decreased the training time of the $α$-$β$ zirconium phase transition by approximately 20 times. The potential describes the phase transition with a degree of accuracy similar to that of the large-cell training method.

cond-mat.mtrl-sci↗

Machine Learning Predictions of High-Curie-Temperature Materials

Technologies that function at room temperature often require magnets with a high Curie temperature, $T_\mathrm{C}$, and can be improved with better materials. Discovering magnetic materials with a substantial $T_\mathrm{C}$ is challenging because of the large number of candidates and the cost of fabricating and testing them. Using the two largest known data sets of experimental Curie temperatures, we develop machine-learning models to make rapid $T_\mathrm{C}$ predictions solely based on the chemical composition of a material. We train a random forest model and a $k$-NN one and predict on an initial dataset of over 2,500 materials and then validate the model on a new dataset containing over 3,000 entries. The accuracy is compared for multiple compounds' representations ("descriptors") and regression approaches. A random forest model provides the most accurate predictions and is not improved by dimensionality reduction or by using more complex descriptors based on atomic properties. A random forest model trained on a combination of both datasets shows that cobalt-rich and iron-rich materials have the highest Curie temperatures for all binary and ternary compounds. An analysis of the model reveals systematic error that causes the model to over-predict low-$T_\mathrm{C}$ materials and under-predict high-$T_\mathrm{C}$ materials. For exhaustive searches to find new high-$T_\mathrm{C}$ materials, analysis of the learning rate suggests either that much more data is needed or that more efficient descriptors are necessary.

cond-mat.mtrl-sci↗

A set of moment tensor potentials for zirconium with increasing complexity

Machine learning force fields (MLFFs) are an increasingly popular choice for atomistic simulations due to their high fidelity and improvable nature. Here, we propose a hybrid small-cell approach that combines attributes of both offline and active learning to systematically expand a quantum mechanical (QM) database while constructing MLFFs with increasing model complexity. Our MLFFs employ the moment tensor potential formalism. During this process, we quantitatively assessed structural properties, elastic properties, dimer potential energies, melting temperatures, phase stability, point defect formation energies, point defect migration energies, free surface energies, and generalized stacking fault (GSF) energies of Zr as predicted by our MLFFs. Unsurprisingly, model complexity has a positive correlation with prediction accuracy. We also find that the MLFFs wee able to predict the properties of out-of-sample configurations without directly including these specific configurations in the training dataset. Additionally, we generated 100 MLFFs of high complexity (1513 parameters each) that reached different local optima during training. Their predictions cluster around the benchmark DFT values, but subtle physical features such as the location of local minima on the GSFE surface are washed out by statistical noise.

physics.comp-ph↗

Tensor-reduced atomic density representations

Density based representations of atomic environments that are invariant under Euclidean symmetries have become a widely used tool in the machine learning of interatomic potentials, broader data-driven atomistic modelling and the visualisation and analysis of materials datasets.The standard mechanism used to incorporate chemical element information is to create separate densities for each element and form tensor products between them. This leads to a steep scaling in the size of the representation as the number of elements increases. Graph neural networks, which do not explicitly use density representations, escape this scaling by mapping the chemical element information into a fixed dimensional space in a learnable way. We recast this approach as tensor factorisation by exploiting the tensor structure of standard neighbour density based descriptors. In doing so, we form compact tensor-reduced representations whose size does not depend on the number of chemical elements, but remain systematically convergeable and are therefore applicable to a wide range of data analysis and regression tasks.

physics.chem-ph↗

Relative Grain Boundary Energies from Triple Junction Geometry: Limitations to Assuming the Herring Condition in Nanocrystalline Thin Films

Grain boundary character distributions (GBCD) are routinely measured from bulk microcrystalline samples by electron backscatter diffraction (EBSD) and serial sectioning can be used to reconstruct relative grain boundary energy distributions (GBED) based on the 3D geometry of triple lines, assuming that the Herring condition of force balance is satisfied. These GBEDs correlate to those predicted from molecular dynamics (MD); furthermore, the GBCD and GBED are found to be inversely correlated. For nanocrystalline thin films, orientation mapping via precession enhanced electron diffraction (PED) has proven effective in measuring the GBCD, but the GBED has not been extracted. Here, the established relative energy extraction technique is adapted to PED data from four sputter deposited samples: a 40 nm-thick tungsten film and a 100 nm aluminum film as-deposited, after 30 and after 150 minutes annealing at 400°C. These films have columnar grain structures, so serial sectioning is not required to determine boundary inclination. Excepting the most energetically anisotropic and highest population boundaries, i.e. aluminum Σ3 boundaries, the relative GBED extracted from these data do not correlate with energies calculated using MD nor do they inversely correlate with the experimentally determined GBCD for either the tungsten or aluminum films. Failure to reproduce predicted energetic trends implies that the conventional Herring equation cannot be applied to determine relative GBEDs and thus geometries at triple junctions in these films are not well described by this condition; additional geometric factors must contribute to determining triple junction geometry and boundary network structure in spatially constrained, polycrystalline materials.

cond-mat.mtrl-sci↗

Examination of computed aluminum grain boundary structures and interface energies that span the 5D space of crystallographic character

The space of possible grain boundary structures is vast, with 5 macroscopic, crystallographic degrees of freedom that define the character of a grain boundary. While numerous datasets of grain boundaries have examined this space in part or in full, we present a computed dataset of over 7304 unique aluminum grain boundaries in the 5D crystallographic space. Our sampling also includes a range of possible microscopic, atomic configurations for each unique 5D crystallographic structure, which total over 43 million structures. We present an overview of the methods used to generate this dataset, an initial examination of the energy trends that follow the Read-Shockley relationship, hints at trends throughout the 5D space, variations in GB energy when non-minimum energy structures are examined, and insights gained in machine learning of grain boundary energy structure-property relationships. This dataset, which is available for download, has great potential for insight into GB structure-property relationships.

cond-mat.mtrl-sci↗

A general algorithm for calculating irreducible Brillouin zones

Calculations of properties of materials require performing numerical integrals over the Brillouin zone (BZ). Integration points in density functional theory codes are uniformly spread over the BZ (despite integration error being concentrated in small regions of the BZ) and preserve symmetry to improve computational efficiency. Integration points over an irreducible Brillouin zone (IBZ), a rotationally distinct region of the BZ, do not have to preserve crystal symmetry for greater efficiency. This freedom allows the use of adaptive meshes with higher concentrations of points at locations of large error, resulting in improved algorithmic efficiency. We have created an algorithm for constructing an IBZ of any crystal structure in 2D and 3D. The algorithm uses convex hull and half-space representations for the BZ and IBZ to make many aspects of construction and symmetry reduction of the BZ trivial. The algorithm is simple, general, and available as open-source software.

cond-mat.mtrl-sci↗

Effectiveness of smearing and tetrahedron methods: best practices in DFT codes

Density functional theory (DFT) codes are commonly treated as a "black box" in high-throughput screening of materials, with users opting for the default values of the input parameters. Often, non-experts may not sufficiently consider the effect of these parameters on prediction quality. In this work, we attempt to identify a robust set of parameters related to smearing and tetrahedron methods that return numerically accurate and efficient results for a wide variety of metallic systems. The effects of smearing and tetrahedron methods on the total energy, number of self-consistent field cycles, and forces on atoms are studied in two popular DFT codes: the Vienna Ab initio Simulation Package (VASP) and Quantum Espresso (QE). From nearly 40,000 computations, it is apparent that the optimal smearing depends on the system, smearing method, smearing parameter, and $k$-point density. The benefit of smearing is a minor reduction in the number of self-consistent field cycles, which is independent of the smearing method or parameter. A large smearing parameter -- what is considered large is system dependent -- leads to inaccurate total energies and forces. Blöchl's tetrahedron method leads to small improvements in total energies. When treating diverse systems with the same input parameters, we suggest using as little smearing as possible due to the system dependence of smearing and the risk of selecting a parameter that gives inaccurate energies and forces.

cond-mat.mtrl-sci↗

The AFLOW Library of Crystallographic Prototypes: Part 3

The AFLOW Library of Crystallographic Prototypes has been extended to include a total of 1,100 common crystal structural prototypes (510 new ones with Part 3), comprising all of the inorganic crystal structures defined in the seven-volume Strukturbericht series published in Germany from 1937 through 1943. We cover a history of the Strukturbericht designation system, the evolution of the system over time, and the first comprehensive index of inorganic Strukturbericht designations ever published.

cond-mat.mtrl-sci↗

Machine-learned Interatomic Potentials for Alloys and Alloy Phase Diagrams

We introduce machine-learned potentials for Ag-Pd to describe the energy of alloy configurations over a wide range of compositions. We compare two different approaches. Moment tensor potentials (MTP) are polynomial-like functions of interatomic distances and angles. The Gaussian Approximation Potential (GAP) framework uses kernel regression, and we use the Smooth Overlap of Atomic Positions (SOAP) representation of atomic neighbourhoods that consists of a complete set of rotational and permutational invariants provided by the power spectrum of the spherical Fourier transform of the neighbour density. Both types of potentials give excellent accuracy for a wide range of compositions and rival the accuracy of cluster expansion, a benchmark for this system. While both models are able to describe small deformations away from the lattice positions, SOAP-GAP excels at transferability as shown by sensible transformation paths between configurations, and MTP allows, due to its lower computational cost, the calculation of compositional phase diagrams. Given the fact that both methods perform as well as cluster expansion would but yield off-lattice models, we expect them to open new avenues in computational materials modeling for alloys.

cond-mat.mtrl-sci↗

A robust algorithm for $k$-point grid generation and symmetry reduction

We develop an algorithm for i) computing generalized regular $k$-point grids, ii) reducing the grids to their symmetrically distinct points, and iii) mapping the reduced grid points into the Brillouin zone. The algorithm exploits the connection between integer matrices and finite groups to achieve a computational complexity that is linear with the number of $k$-points. The favorable scaling means that, at a given $k$-point density, all possible commensurate grids can be generated (as suggested by Moreno and Soler) and quickly reduced to identify the grid with the fewest symmetrically unique $k$-points. These optimal grids provide significant speed-up compared to Monkhorst-Pack $k$-point grids; they have better symmetry reduction resulting in fewer irreducible $k$-points at a given grid density. The integer nature of this new reduction algorithm also simplifies issues with finite precision in current implementations. The algorithm is available as open source software.

physics.comp-ph↗

Machine-learned multi-system surrogate models for materials prediction

Surrogate machine-learning models are transforming computational materials science by predicting properties of materials with the accuracy of ab initio methods at a fraction of the computational cost. We demonstrate surrogate models that simultaneously interpolate energies of different materials on a dataset of 10 binary alloys (AgCu, AlFe, AlMg, AlNi, AlTi, CoNi, CuFe, CuNi, FeV, NbNi) with 10 different species and all possible fcc, bcc and hcp structures up to 8 atoms in the unit cell, 15\,950 structures in total. We find that the deviation of prediction errors when increasing the number of simultaneously modeled alloys is less than 1\,meV/atom. Several state-of-the-art materials representations and learning algorithms were found to qualitatively agree on the prediction errors of formation enthalpy with relative errors of $<$2.5\% for all systems.

cond-mat.mtrl-sci↗