SearcharxivSearch

arXiv subjects

Kenneth Kroenlein

Publications and source records attributed to Kenneth Kroenlein.

5 recordsLinked to original sources

Citrine Informatics: Chemical & Materials Development Platform

Today the Citrine Platform regularly powers data-driven materials discovery across industries, having moved beyond one-off demonstrations into routine industrial practice. Getting there required solving a core set of recurring obstacles: experimental data are scarce, costly, and published in formats that resist reuse; conventional accuracy metrics overstate model performance under the extrapolative conditions that define discovery; and realistic design spaces are bounded by physics, manufacturability, supply, and cost. Developed over more than a decade as an integrated response to these obstacles, the Citrine Platform is organized as four cooperating stages within a closed sequential learning loop. Stage 1 ingests and featurizes data through the Graphical Expression of Materials Data (GEMD) model, which treats process history, measurement uncertainty, and provenance as first-class features. Stage 2 builds machine learning models with well-calibrated uncertainty, including multivariate prediction intervals for correlated objectives, and validates them with extrapolative cross-validation and dynamic discovery metrics rather than random held-out splits. Stage 3 encodes compositional, physical, processing, and economic constraints directly into the design space, and Stage 4 applies the FUELS sequential learning framework with uncertainty-aware acquisition functions to navigate large constrained spaces under tight evaluation budgets. Published case studies spanning organic semiconductors, autonomous nanoparticle synthesis, and benchmark optimization tasks demonstrate two- to nine-fold reductions in experimental effort relative to random search, illustrating a stack in which data, modeling, and design-space layers continuously co-evolve.

cond-mat.mtrl-sci

HUGO-CS: A Hybrid-Labeled, Uncertainty-Aware, General-Purpose, Observational Dataset for Cold Spray

Cold spraying is an increasingly common approach for repairing and manufacturing components due to its solid-state manufacturing capabilities. However, process optimization remains difficult due to many interdependent parameters and the lack of large-scale, machine-readable data to support modeling. While the scientific literature contains many relevant experiments, results are inconsistently reported (often in tables and figures) and use non-uniform units, limiting utilization at scale. To address these limitations, this work presents HUGO-CS, a literature-derived dataset of 4,383 cold-spray experiments with 144 features from 1,124 sources, exceeding the previous largest dataset (137 samples) by 30x. With completely manual extraction requiring an average of 91 minutes per document, this work designs and leverages a Hybrid-labeled, Uncertainty-aware, General-purpose, Observational extraction framework, called HUGO, to support this extraction. HUGO combines automated LLM-based labeling with targeted manual label refinement to handle this experimental result extraction process from scientific literature. To balance labeling efficiency with extraction accuracy, HUGO introduces a Hierarchical Risk Mitigation (HRM) to route LLM outputs with a high risk of potential errors for manual review, while retaining low-risk records as auto-labeled. Lastly, HUGO post-processing consolidates categorical descriptors, maps reported feedstock chemistries into structured continuous compositions, and normalizes units across sources. Of the 4,383 reported experiments, 1,765 are hand-labeled, providing a high-quality labeled subset for benchmarking, error analysis, and higher-fidelity data points. All code to replicate this work, along with the complete HUGO-CS dataset, are released under a CC-BY license at https://github.com/sprice134/HUGO.

cs.LG

Predicting Magnetic Janus Particle Assembly with Differential Evolution Algorithm

Magnetic Janus particles allow access to complex, nonlinear assembled structures that may enable interesting new magnetorheological (MR) fluids with uniquely engineered field responses. However, the overwhelming size of the parameter space for Janus and patchy particles makes exploration of such systems by experimental trial and error or through detailed simulation impractical. Here, a differential evolution (DE)-based simulation method is explored to predict the assembly of magnetic Janus particles as an alternative method for assembly prediction. Structure predictions from the DE simulation for laterally- and radially-shifted magnetic Janus particles are compared to four published experimental and simulation case studies. The DE simulation captures the orientation and structure of magnetic Janus particles for a range of shifts and a variety of external field conditions using the point dipole approximation. Structural predictions that rely on the reorganization of large clusters of particles were less well represented by the DE predictions. Despite this limitation, the DE simulation method can be used to predict key structural factors for magnetic Janus particle assemblies, as demonstrated by favorable comparison with three of the four model studies.

cond-mat.soft

Nucleobase-functionalized graphene nanoribbons for accurate high-speed DNA sequencing

We propose a water-immersed nucleobase-functionalized suspended graphene nanoribbon as an intrinsically selective device for nucleotide detection. The proposed sensing method combines Watson-Crick selective base pairing with graphene's capacity for converting anisotropic lattice strain to changes in an electrical current at the nanoscale. Using detailed atomistic molecular dynamics simulations, we study sensor operation at ambient conditions. We combine simulated data with theoretical arguments to estimate the levels of measurable electrical signal variation in response to strains and determine that the proposed sensing mechanism shows significant promise for realistic DNA sensing devices without the need for advanced data processing, or highly restrictive operational conditions.

cond-mat.mes-hall

Towards Automated Benchmarking of Atomistic Forcefields: Neat Liquid Densities and Static Dielectric Constants from the ThermoML Data Archive

Atomistic molecular simulations are a powerful way to make quantitative predictions, but the accuracy of these predictions depends entirely on the quality of the forcefield employed. While experimental measurements of fundamental physical properties offer a straightforward approach for evaluating forcefield quality, the bulk of this information has been tied up in formats that are not machine-readable. Compiling benchmark datasets of physical properties from non-machine-readable sources require substantial human effort and is prone to accumulation of human errors, hindering the development of reproducible benchmarks of forcefield accuracy. Here, we examine the feasibility of benchmarking atomistic forcefields against the NIST ThermoML data archive of physicochemical measurements, which aggregates thousands of experimental measurements in a portable, machine-readable, self-annotating format. As a proof of concept, we present a detailed benchmark of the generalized Amber small molecule forcefield (GAFF) using the AM1-BCC charge model against measurements (specifically bulk liquid densities and static dielectric constants at ambient pressure) automatically extracted from the archive, and discuss the extent of available data. The results of this benchmark highlight a general problem with fixed-charge forcefields in the representation low dielectric environments such as those seen in binding cavities or biological membranes.

physics.chem-ph