SearcharxivSearch

arXiv subjects

Markus Stricker

Publications and source records attributed to Markus Stricker.

17 recordsLinked to original sources

Integrating Semantics into Research Data Management: Modelling and Validating Materials Science Experiment Workflows

The incorporation of Semantic Web technologies within scientific environments is becoming an increasingly popular Research Data Management (RDM) practice. While ontologies offer flexible, reusable and machine-readable vocabularies to describe domain-specific research data, Knowledge Graphs (KGs) facilitate the integration of heterogeneous data sources into an interoperable collection. Furthermore, KGs offer additional advantages, notably the use of expressive SPARQL queries, or the ability to define complex data validation rules with SHACL. This work describes the modelling of a relational RDM system with an ontology, and the subsequent construction of a KG based on it. Enabled by the highly interconnected nature of research data and experiment workflows present in the system, we not only show how we can easily and reliably build an efficient KG from such domain-specific RDM systems, but also how doing so enables more advanced use cases. This is demonstrated by the modelling of ideal counterparts for the experiment workflows logged in the KG, which are then used to programmatically generate SHACL shapes that fully validate the conformance of the latter. By integrating this functionality within a UI, we allow researchers to plan, reuse, share, and track the progress of their daily experiments.

cs.DB

Journal Research Data Policies in Materials Science

Open and reproducible research in materials science relies on the availability of data, code, and common metadata standards. Journal research data policies (RDPs) remain a primary mechanism by which publication norms are defined and enforced. We survey RDPs for 171 materials science journals spanning 17 publishers, using an expanded coding framework that captures both data-and-code sharing behavior as well as refereeing standards. We find clear signs of progress in comparison to earlier research on RDPs: nearly all journals provide an RDP, and most mention data availability statements. However, enforceable requirements remain uncommon, public deposition of underlying data is rarely mandatory, and FAIR publication is typically encouraged rather than required. Expectations for research software are substantially less developed than those for data, with limited attention to versioning and persistent identifiers, dependency disclosure, reproducible execution environments, or software quality practices. Aggregating the findings on policy features into an open research data score reveals pronounced heterogeneity across journals. Neither impact factor nor access model reliably predicts policy strength. Double-coding further shows that more complex policies and stricter policies can be more challenging to interpret consistently, and we highlight challenges in consistent RDP encoding across studies. Lastly, we conclude with recommended best practice directions for the future.

cs.DL

From Word2Vec to Transformers: Text-Derived Composition Embeddings for Filtering Combinatorial Electrocatalysts

Compositionally complex solid solution electrocatalysts span vast composition spaces, and even one materials system can contain more candidate compositions than can be measured exhaustively. Here we evaluate a label-free screening strategy that represents each composition using embeddings derived from scientific texts and prioritizes candidates based on similarity to two property concepts. We compare a corpus-trained Word2Vec baseline with transformer-based embeddings, where compositions are encoded either by linear element-wise mixing or by short composition prompts. Similarities to `concept directions', the terms conductivity and dielectric, define a 2-dimensional descriptor space, and a symmetric Pareto-front selection is used to filter candidate subsets without using electrochemical labels. Performance is assessed on 15 materials libraries including noble metal alloys and multicomponent oxides. In this setting, the lightweight Word2Vec baseline, which uses a simple linear combination of element embeddings, often achieves the highest number of reductions of possible candidate compositions while staying close to the best measured performance.

cond-mat.mtrl-sci

Ontology-aligned structuring and reuse of multimodal materials data and workflows towards automatic reproduction

Reproducibility of computational results remains a challenge in materials science, as simulation workflows and parameters are often reported only in unstructured text and tables. While literature data are valuable for validation and reuse, the lack of machine-readable workflow descriptions prevents large-scale curation and systematic comparison. Existing text-mining approaches are insufficient to extract complete computational workflows with their associated parameters. An ontology-driven, large language model (LLM)-assisted framework is introduced for the automated extraction and structuring of computational workflows from the literature. The approach focuses on density functional theory-based stacking fault energy (SFE) calculations in hexagonal close-packed magnesium and its binary alloys, and uses a multi-stage filtering strategy together with prompt-engineered LLM extraction applied to method sections and tables. Extracted information is unified into a canonical schema and aligned with established materials ontologies (CMSO, ASMO, and PLDO), enabling the construction of a knowledge graph using atomRDF. The resulting knowledge graph enables systematic comparison of reported SFE values and supports the structured reuse of computational protocols. While full computational reproducibility is still constrained by missing or implicit metadata, the framework provides a foundation for organizing and contextualizing published results in a semantically interoperable form, thereby improving transparency and reusability of computational materials data.

cond-mat.mtrl-sci

Multi-modal data-driven microstructure characterization

Electron backscatter diffraction is one of the most prevalent techniques used for microstructural characterization. In recent years, there has been an increase in the use of data-driven methods to analyze raw Kikuchi patterns. However, most of these require user input and the interpretation of the data-derived features is often challenging and subject to \textit{informed interpretation}. By using a combination of principal component analysis, constrained non-negative matrix factorization, and a variational autoencoder along with information-theoretical considerations on a multimodal dataset, it is shown that a) automated decision on method-specific hyperparameters, here the number of components in principal component analysis, the number of components for constrained non-negative matrix factorization, and the selection of reference constraints; and b) latent space features can be mapped to physically-meaningful quantities. In addition, the recommended region-of-interest (ROI) size for optimal model performance is approximated automatically to be twice the characteristic grain size based on information content of the dataset. Implemented in a workflow, this allows for a transferable, dataset-specific autonomous data-driven phase and grain segmentation including grain boundary detection and the analysis of very-small-angle intra-grain variations to complement conventional electron backscatter analysis.

cond-mat.mtrl-sci

Field report from Collaborative Research Center 1625: Heterogeneous research data management using ontology representations

The goal of the Collaborative Research Center 1625 is the establishment of a scientific basis for the atomic-scale understanding and design of multifunctional compositionally complex solid solution surfaces. Next to materials synthesis in form of thin-film materials libraries, various materials characterization and simulations techniques are used to explore the materials data space of the problem. Machine learning and artificial intelligence techniques guide its exploration and navigation. The effective use of the combined heterogeneous data requires more than just a simple research data management plan. Consequently, our research data management system maps different data modalities in different formats and resolutions from different labs to the correct spatial locations on physical samples. Besides a graphical user interface, the system can also be accessed through an application programming interface for reproducible data-driven workflows. It is implemented by a combination of a custom research data management system designed around a relational database, an ontology which builds upon materials science-specific ontologies, and the construction of a Knowledge Graph. Along with the technical solutions of research data management system and lessons learned, first use cases are shown which were not possible (or at least much harder to achieve) without it.

cond-mat.mtrl-sci

Iterative Corpus Refinement for Materials Property Prediction Based on Scientific Texts

The discovery and optimization of materials for specific applications is hampered by the practically infinite number of possible elemental combinations and associated properties, also known as the `combinatorial explosion'. By nature of the problem, data are scarce and all possible data sources should be used. In addition to simulations and experimental results, the latent knowledge in scientific texts is not yet used to its full potential. We present an iterative framework that refines a given scientific corpus by strategic selection of the most diverse documents, training Word2Vec models, and monitoring the convergence of composition-property correlations in embedding space. Our approach is applied to predict high-performing materials for oxygen reduction (ORR), hydrogen evolution (HER), and oxygen evolution (OER) reactions for a large number of possible candidate compositions. Our method successfully predicts the highest performing compositions among a large pool of candidates, validated by experimental measurements of the electrocatalytic performance in the lab. This work demonstrates and validates the potential of iterative corpus refinement to accelerate materials discovery and optimization, offering a scalable and efficient tool for screening large compositional spaces where reliable data are scarce or non-existent.

cs.CL

Electrocatalyst discovery through text mining and multi-objective optimization

The discovery and optimization of high-performance materials is the basis for advancing energy conversion technologies. To understand composition-property relationships, all available data sources should be leveraged: experimental results, predictions from simulations, and latent knowledge from scientific texts. Among these three, text-based data sources are still not used to their full potential. We present an approach combining text mining, Word2Vec representations of materials and properties, and Pareto front analysis for the prediction of high-performance candidate materials for electrocatalysis in regions where other data sources are scarce or non-existent. Candidate compositions are evaluated on the basis of their similarity to the terms `conductivity' and `dielectric', which enables reaction-specific candidate composition predictions for oxygen reduction (ORR), hydrogen evolution (HER), and oxygen evolution (OER) reactions. This, combined with Pareto optimization, allows us to significantly reduce the pool of candidate compositions to high-performing compositions. Our predictions, which are purely based on text data, match the measured electrochemical activity very well.

cond-mat.mtrl-sci

Control of ferroelectric domain wall dynamics by point defects: Insights from ab initio based simulations

The control of ferroelectric domain walls and their dynamics on the nanoscale becomes increasingly important for advanced nanoelectronics and novel computing schemes. One common approach to tackle this challenge is the pinning of walls by point defects. The fundamental understanding on how different defects influence the wall dynamics is, however, incomplete. In particular, the important class of defect dipoles in acceptor-doped ferroelectrics is currently underrepresented in theoretical work. In this study, we combine molecular dynamics simulations based on an \textit{ab\ initio}-derived effective Hamiltonian and methods from materials informatics, and analyze the impact of these defects on the motion of 180$^{\circ}$ domain walls in tetragonal BaTiO$_3$. We show how these defects can act as local pinning centers and restoring forces on the domain structure. Furthermore, we reveal how walls can flow around sparse defects by nucleation and growth of dipole clusters, and how pinning, roughening and bending of walls depend on the defect distribution. Surprisingly, the interaction between acceptor dopants and walls is short-ranged. We show that the limiting factor for the nucleation processes underlying wall motion is the defect-free area in front of the wall.

cond-mat.mtrl-sci

Composition-property extrapolation for compositionally complex solid solutions based on word embeddings

Mastering the challenge of predicting properties of unknown materials with multiple principal elements (high entropy alloys/compositionally complex solid solutions) is crucial for the speedup in materials discovery. We show and discuss three models, using property data from two ternary systems (Ag-Pd-Ru; Ag-Pd-Pt), to predict material performance in the shared quaternary system (Ag-Pd-Pt-Ru). First, we apply Gaussian Process Regression (GPR) based on composition, which includes both Ag and Pd, achieving an initial correlation coefficient for the prediction ($r$) of 0.63 and a determination coefficient ($r^2$) of 0.08. Second, we present a version of the GPR model using word embedding-derived materials vectors as representations. Using materials-specific embedding vectors significantly improves the predictive capability, evident from an improved $r^2$ of 0.65. The third model is based on a `standard vector method' which synthesizes weighted vector representations of material properties, then creating a reference vector that results in a very good correlation with the quaternary system's material performance (resulting $r$ of 0.89). Our approach demonstrates that existing experimental data combined with latent knowledge of word embedding-based representations of materials can be used effectively for materials discovery where data is typically sparse.

cond-mat.mtrl-sci

Dislocation cartography: Representations and unsupervised classification of dislocation networks with unique fingerprints

Detecting structure in data is the first step to arrive at meaningful representations for systems. This is particularly challenging for dislocation networks evolving as a consequence of plastic deformation of crystalline systems. Our study employs Isomap, a manifold learning technique, to unveil the intrinsic structure of high-dimensional density field data of dislocation structures from different compression axis. The resulting maps provide a systematic framework for quantitatively comparing dislocation structures, offering unique fingerprints based on density fields. Our novel, unbiased approach contributes to the quantitative classification of dislocation structures which can be systematically extended.

cond-mat.mtrl-sci

Employing constrained non-negative matrix factorization for microstructure segmentation

Materials characterization using electron backscatter diffraction (EBSD) requires indexing the orientation of the measured region from Kikuchi patterns. The quality of Kikuchi patterns can degrade due to pattern overlaps arising from two or more orientations, in the presence of defects or grain boundaries. In this work we employ constrained non-negative matrix factorization to segment a microstructure with small grain misorientations,< 1 degree, and predict the amount of pattern overlap. First we implement the method on mixed simulated patterns - that replicates a pattern overlap scenario, and demonstrate the resolution limit of pattern mixing or factorization resolution using a weight metric. Subsequently, we segment a single-crystal dendritic microstructure and compare the results with high resolution EBSD. By utilizing weight metrics across a low angle grain boundary we demonstrate how very small misorientations/low-angle grain boundaries can be resolved at a pixel level. Our approach constitutes a versatile and robust tool, complementing other fast indexing methods for microstructure characterization.

cond-mat.mtrl-sci

MatNexus: A Comprehensive Text Mining and Analysis Suite for Materials Discover

MatNexus is a specialized software for the automated collection, processing, and analysis of text from scientific articles. Through an integrated suite of modules, the MatNexus facilitates the retrieval of scientific articles, processes textual data for insights, generates vector representations suitable for machine learning, and offers visualization capabilities for word embeddings. With the vast volume of scientific publications, MatNexus stands out as an end-to-end tool for researchers aiming to gain insights from scientific literature in material science, making the exploration of materials, such as the electrocatalyst examples we show here, efficient and insightful.

cond-mat.mtrl-sci

Microscopic insights on field induced switching and domain wall motion in orthorhombic ferroelectrics

Surprisingly little is known about the microscopic processes that govern ferroelectric switching in orthorhombic ferroelectrics. To study microscopic switching processes we combine ab initio-based molecular dynamics simulations and data science on the prototypical material BaTiO$_3$. We reveal two different field regimes: For moderate field strengths, the switching is dominated by domain wall motion while a fast bulk-like switching can be induced for large fields. Switching in both field regimes follows a multi-step process via polarization directions perpendicular to the applied field. In the former case, the moving wall is of Bloch character and hosts dipole vortices due to nucleation, growth, and crossing of two dimensional 90$^{\circ}$ domains. In the second case, the local polarization shows a continuous correlated rotation via a an intermediate tetragonal multidomain state.

cond-mat.mtrl-sci

Statistical analysis of Discrete Dislocation Dynamics simulations: initial structures, cross-slip and microstructure evolution

Over the past decades, discrete dislocation dynamics simulations have been shown to reliably predict the evolution of dislocation microstructures for micrometer-sized metallic samples. Such simulations provide insight into the governing deformation mechanisms and the interplay between different physical phenomena such as dislocation reactions or cross-slip. This work is focused on a detailed analysis of the influence of the cross-slip on the evolution of dislocation systems. A tailored data mining strategy using the ``\ab*{discrete-to-continuous} framework'' allows to quantify differences and to quantitatively compare dislocation structures. We analyze the quantitative effects of the cross-slip on the microstructure in the course of a tensile test and a subsequent relaxation to present the role of cross-slip in the microstructure evolution. The precision of the extracted quantitative information using D2C strongly depends on the resolution of the domain averaging. We also analyze how the resolution of the averaging influences the distribution of total dislocation density and curvature fields of the specimen. Our analyzes are important approaches for interpreting the resulting structures calculated by dislocation dynamics simulations.

cond-mat.mtrl-sci

Computationally accelerated experimental materials characterization -- drawing inspiration from high-throughput simulation workflows

Computational materials science increasingly benefits from data management, automation, and algorithm-based decision-making for the simulation of material properties and behavior. Experimental materials science also changes rapidly by incorporation of `machine learning' in materials discovery campaigns. The obvious benefits which include automation, reproducibility, data provenance, and reusability of managed data, however, is not widely available in the experimental domain. We present an implementation of a Active Learning loop with a direct interface to an experimental measurement device in pyiron, a framework designed for high-throughput simulations, as demonstrator how to combine experimental and simulated data in one framework. Apart from the acceleration provided by the active learning approach, additional acceleration of the experimental characterization is achieved by using prior knowledge from density functional theory simulations as well as composition-property predictions from literature mining using correlations in word embeddings. With data from all domains in the same framework, a heretofore untapped and much-needed potential for the acceleration of materials characterization and materials discovery campaigns becomes available.

cond-mat.mtrl-sci

Efficient implementation of atom-density representations

Physically-motivated and mathematically robust atom-centred representations of molecular structures are key to the success of modern atomistic machine learning (ML) methods. They lie at the foundation of a wide range of methods to predict the properties of both materials and molecules as well as to explore and visualize the chemical compound and configuration space. Recently, it has become clear that many of the most effective representations share a fundamental formal connection: that they can all be expressed as a discretization of N-body correlation functions of the local atom density, suggesting the opportunity of standardizing and, more importantly, optimizing the calculation of such representations. We present an implementation, named librascal, whose modular design lends itself both to developing refinements to the density-based formalism and to rapid prototyping for new developments of rotationally equivariant atomistic representations. As an example, we discuss SOAP features, perhaps the most widely used member of this family of representations, to show how the expansion of the local density can be optimized for any choice of radial basis set. We discuss the representation in the context of a kernel ridge regression model, commonly used with SOAP features, and analyze how the computational effort scales for each of the individual steps of the calculation. By applying data reduction techniques in feature space, we show how to further reduce the total computational cost by at up to a factor of 4 or 5 without affecting the model's symmetry properties and without significantly impacting its accuracy.

physics.chem-ph