SearcharxivSearch

arXiv subjects

Angel Yanguas-Gil

Publications and source records attributed to Angel Yanguas-Gil.

At least 19 recordsLinked to original sources

Evaluating LLM-based AI agents integrated with materials synthesis tools: the case of atomic layer deposition

This work provides an overview of the different strategies that can be used to evaluate the performance of AI models and agents based on large language models (LLMs) for materials synthesis. After providing a brief overview of the key technologies behind the current generation of AI agents based on LLMs, we summarize the different approaches to evaluating these models in the context of materials science and in particular on materials synthesis, with a specific emphasis on scenarios in which the models are directly integrated with experimental tools. We discuss evaluation strategies spanning knowledge and reasoning benchmarks, tool-use benchmarks, and closed loop benchmarks involving the interaction with experimental systems or realistic virtual tools. We use atomic layer deposition (ALD) as a case study, emphasizing how existing approaches in the literature both build from general approaches used beyond materials science and can be generalized to other materials synthesis techniques. Finally, we provide a practical evaluation framework to evaluate LLMs in the context of materials synthesis

cond-mat.mtrl-sci

Performance of AI agents based on reasoning language models on ALD process optimization tasks

In this work we explore the performance and behavior of reasoning large language models to autonomously optimize atomic layer deposition (ALD) processes. In the ALD process optimization task, an agent built on top of a reasoning LLM has to find optimal dose times for an ALD precursor and a coreactant without any prior knowledge on the process, including whether it is actually self-limited. The agent is meant to interact iteratively with an ALD reactor in a fully unsupervised way. We evaluate this agent using a simple model of an ALD tool that incorporates ALD processes with different self-limited surface reaction pathways as well as a non self-limited component. Our results show that agents based on reasoning models like OpenAI's o3 and GPT5 consistently succeeded at completing this optimization task. However, we observed significant run-to-run variability due to the non deterministic nature of the model's response. In order to understand the logic followed by the reasoning model, the agent uses a two step process in which the model first generates an open response detailing the reasoning process. This response is then transformed into a structured output. An analysis of these reasoning traces showed that the logic of the model was sound and that its reasoning was based on the notions of self-limited process and saturation expected in the case of ALD. However, the agent can sometimes be misled by its own prior choices when exploring the optimization space.

cond-mat.mtrl-sci

Microstructural and preliminary optical and microwave characterization of erbium doped CaMoO$_4$ thin films

This work explores erbium-doped calcium molybdate (CaMoO$_4$) thin films grown on silicon and yttria stabilized zirconia (YSZ) substrates, as a potential solid state system for C-band (utilizing the $\sim$1.5 $μ$m Er$^{3+}$ 4f-4f transition) quantum emitters for quantum network applications. Through molecular beam epitaxial growth experiments and electron microscopy, X-ray diffraction and reflection electron diffraction studies, we identify an incorporation limited deposition regime that enables a 1:1 Ca:Mo ratio in the growing film leading to single phase CaMoO$_4$ formation that can be in-situ doped with Er (typically 2-100 ppm). We further show that growth on silicon substrates is single phase but polycrystalline in morphology; while growth on YSZ substrates leads to high-quality epitaxial single crystalline CaMoO$_4$ films. We perform preliminary optical and microwave characterization on the suspected $Y_1 - Z_1$ transition of 2 ppm, 200 nm epitaxial CaMoO$_4$ annealed thin films and extract an optical inhomogeneous linewidth of 9.1(1) GHz, an optical excited state lifetime of 6.7(2) ms, a spectral diffusion-limited homogeneous linewidth of 6.7(4) MHz, and an EPR linewidth of 1.10(2) GHz.

cond-mat.mtrl-sci

Surrogate models to optimize plasma assisted atomic layer deposition in high aspect ratio features

In this work we explore surrogate models to optimize plasma enhanced atomic layer deposition (PEALD) in high aspect ratio features. In plasma-based processes such as PEALD and atomic layer etching, surface recombination can dominate the reactivity of plasma species with the surface, which can lead to unfeasibly long exposure times to achieve full conformality inside nanostructures like high aspect ratio vias. Using a synthetic dataset based on simulations of PEALD, we train artificial neural networks to predict saturation times based on cross section thickness data obtained for partially coated conditions. The results obtained show that just two experiments in undersaturated conditions contain enough information to predict saturation times within 10% of the ground truth. A surrogate model trained to determine whether surface recombination dominates the plasma-surface interactions in a PEALD process achieves 99% accuracy. This demonstrates that machine learning can provide a new pathway to accelerate the optimization of PEALD processes in areas such as microelectronics. Our approach can be easily extended to atomic layer etching and more complex structures.

cond-mat.mtrl-sci

EAIRA: Establishing a Methodology for Evaluating AI Models as Scientific Research Assistants

Recent advancements have positioned AI, and particularly Large Language Models (LLMs), as transformative tools for scientific research, capable of addressing complex tasks that require reasoning, problem-solving, and decision-making. Their exceptional capabilities suggest their potential as scientific research assistants but also highlight the need for holistic, rigorous, and domain-specific evaluation to assess effectiveness in real-world scientific applications. This paper describes a multifaceted methodology for Evaluating AI models as scientific Research Assistants (EAIRA) developed at Argonne National Laboratory. This methodology incorporates four primary classes of evaluations. 1) Multiple Choice Questions to assess factual recall; 2) Open Response to evaluate advanced reasoning and problem-solving skills; 3) Lab-Style Experiments involving detailed analysis of capabilities as research assistants in controlled environments; and 4) Field-Style Experiments to capture researcher-LLM interactions at scale in a wide range of scientific domains and applications. These complementary methods enable a comprehensive analysis of LLM strengths and weaknesses with respect to their scientific knowledge, reasoning abilities, and adaptability. Recognizing the rapid pace of LLM advancements, we designed the methodology to evolve and adapt so as to ensure its continued relevance and applicability. This paper describes the methodology state at the end of February 2025. Although developed within a subset of scientific domains, the methodology is designed to be generalizable to a wide range of scientific domains.

cs.AI

Benchmarking large language models for materials synthesis: the case of atomic layer deposition

In this work we introduce an open-ended question benchmark, ALDbench, to evaluate the performance of large language models (LLMs) in materials synthesis, and in particular in the field of atomic layer deposition, a thin film growth technique used in energy applications and microelectronics. Our benchmark comprises questions with a level of difficulty ranging from graduate level to domain expert current with the state of the art in the field. Human experts reviewed the questions along the criteria of difficulty and specificity, and the model responses along four different criteria: overall quality, specificity, relevance, and accuracy. We ran this benchmark on an instance of OpenAI's GPT-4o. The responses from the model received a composite quality score of 3.7 on a 1 to 5 scale, consistent with a passing grade. However, 36% of the questions received at least one below average score. An in-depth analysis of the responses identified at least five instances of suspected hallucination. Finally, we observed statistically significant correlations between the difficulty of the question and the quality of the response, the difficulty of the question and the relevance of the response, and the specificity of the question and the accuracy of the response as graded by the human experts. This emphasizes the need to evaluate LLMs across multiple criteria beyond difficulty or accuracy.

cs.LG

Modeling scale-up of particle coating by atomic layer deposition

Atomic layer deposition (ALD) is a promising technique to functionalize particle surfaces for energy applications including energy storage, catalysis, and decarbonization. In this work, we present a set of models of ALD particle coating to explore the transition from lab scale to manufacturing. Our models encompass the main particle coating manufacturing approaches including rotary bed, fluidized bed, and continuously vibrating reactors. These models provide key metrics, such as throughput and precursor utilization, required to evaluate the scalability of ALD manufacturing approaches and their feasibility in the context of energy applications. Our results show that designs that force the precursor to flow through fluidized particles transition faster to a transport-limited regime where throughput is maximized. They also exhibit higher precursor utilization. In the context of continuous processes, our models indicate that it is possible to achieve self-extinguishing processes with almost 100% precursor utilization. A comparison with past experimental results of ALD in fluidized bed reactors shows excellent qualitative and quantitative agreement.

cond-mat.mtrl-sci

Design Principles for Lifelong Learning AI Accelerators

Lifelong learning - an agent's ability to learn throughout its lifetime - is a hallmark of biological learning systems and a central challenge for artificial intelligence (AI). The development of lifelong learning algorithms could lead to a range of novel AI applications, but this will also require the development of appropriate hardware accelerators, particularly if the models are to be deployed on edge platforms, which have strict size, weight, and power constraints. Here, we explore the design of lifelong learning AI accelerators that are intended for deployment in untethered environments. We identify key desirable capabilities for lifelong learning accelerators and highlight metrics to evaluate such accelerators. We then discuss current edge AI accelerators and explore the future design of lifelong learning accelerators, considering the role that different emerging technologies could play.

cs.LG

Improving Performance in Continual Learning Tasks using Bio-Inspired Architectures

The ability to learn continuously from an incoming data stream without catastrophic forgetting is critical to designing intelligent systems. Many approaches to continual learning rely on stochastic gradient descent and its variants that employ global error updates, and hence need to adopt strategies such as memory buffers or replay to circumvent its stability, greed, and short-term memory limitations. To address this limitation, we have developed a biologically inspired lightweight neural network architecture that incorporates synaptic plasticity mechanisms and neuromodulation and hence learns through local error signals to enable online continual learning without stochastic gradient descent. Our approach leads to superior online continual learning performance on Split-MNIST, Split-CIFAR-10, and Split-CIFAR-100 datasets compared to other memory-constrained learning approaches and matches that of the state-of-the-art memory-intensive replay-based approaches. We further demonstrate the effectiveness of our approach by integrating key design concepts into other backpropagation-based continual learning algorithms, significantly improving their accuracy. Our results provide compelling evidence for the importance of incorporating biological principles into machine learning models and offer insights into how we can leverage them to design more efficient and robust systems for online continual learning.

cs.LG

AutoML for neuromorphic computing and application-driven co-design: asynchronous, massively parallel optimization of spiking architectures

In this work we have extended AutoML inspired approaches to the exploration and optimization of neuromorphic architectures. Through the integration of a parallel asynchronous model-based search approach with a simulation framework to simulate spiking architectures, we are able to efficiently explore the configuration space of neuromorphic architectures and identify the subset of conditions leading to the highest performance in a targeted application. We have demonstrated this approach on an exemplar case of real time, on-chip learning application. Our results indicate that we can effectively use optimization approaches to optimize complex architectures, therefore providing a viable pathway towards application-driven codesign.

cs.NE

A Domain-Agnostic Approach for Characterization of Lifelong Learning Systems

Despite the advancement of machine learning techniques in recent years, state-of-the-art systems lack robustness to "real world" events, where the input distributions and tasks encountered by the deployed systems will not be limited to the original training context, and systems will instead need to adapt to novel distributions and tasks while deployed. This critical gap may be addressed through the development of "Lifelong Learning" systems that are capable of 1) Continuous Learning, 2) Transfer and Adaptation, and 3) Scalability. Unfortunately, efforts to improve these capabilities are typically treated as distinct areas of research that are assessed independently, without regard to the impact of each separate capability on other aspects of the system. We instead propose a holistic approach, using a suite of metrics and an evaluation framework to assess Lifelong Learning in a principled way that is agnostic to specific domains or system techniques. Through five case studies, we show that this suite of metrics can inform the development of varied and complex Lifelong Learning systems. We highlight how the proposed suite of metrics quantifies performance trade-offs present during Lifelong Learning system development - both the widely discussed Stability-Plasticity dilemma and the newly proposed relationship between Sample Efficient and Robust Learning. Further, we make recommendations for the formulation and use of metrics to guide the continuing development of Lifelong Learning systems and assess their progress in the future.

cs.LG

General policy mapping: online continual reinforcement learning inspired on the insect brain

We have developed a model for online continual or lifelong reinforcement learning (RL) inspired on the insect brain. Our model leverages the offline training of a feature extraction and a common general policy layer to enable the convergence of RL algorithms in online settings. Sharing a common policy layer across tasks leads to positive backward transfer, where the agent continuously improved in older tasks sharing the same underlying general policy. Biologically inspired restrictions to the agent's network are key for the convergence of RL algorithms. This provides a pathway towards efficient online RL in resource-constrained scenarios.

cs.LG

Machine learning and atomic layer deposition: predicting saturation times from reactor growth profiles using artificial neural networks

In this work we explore the application of deep neural networks to the optimization of atomic layer deposition processes based on thickness values obtained at different points of an ALD reactor. We introduce a dataset designed to train neural networks to predict saturation times based on the dose time and thickness values measured at different points of the reactor for a single experimental condition. We then explore different artificial neural network configurations, including depth (number of hidden layers) and size (number of neurons in each layers) to better understand the size and complexity that neural networks should have to achieve high predictive accuracy. The results obtained show that trained neural networks can accurately predict saturation times without requiring any prior information on the surface kinetics. This provides a viable approach to minimize the number of experiments required to optimize new ALD processes in a known reactor. However, the datasets and training procedure depend on the reactor geometry.

cs.LG

Reactor scale simulations of ALD and ALE: ideal and non-ideal self-limited processes in a cylindrical and a 300 mm wafer cross-flow reactor

We have developed a simulation tool to model self-limited processes such as atomic layer deposition and atomic layer etching inside reactors of arbitrary geometry. In this work, we have applied this model to two standard types of cross-flow reactors: a cylindrical reactor and a model 300 mm wafer reactor, and explored both ideal and non-ideal self-limited kinetics. For the cylindrical tube reactor the full simulation results agree well with analytic expressions obtained using a simple plug flow model, though the presence of axial diffusion tends to soften growth profiles with respect to the plug flow case. Our simulations also allowed us to model the output of in-situ techniques such as quartz crystal microbalance and mass spectrometry, providing a way of discriminating between ideal and non-ideal surface kinetics using in-situ measurements. We extended the simulations to consider two non-ideal self-limited processes: soft-saturating processes characterized by a slow reaction pathway, and processes where surface byproducts can compete with the precursor for the same pool of adsorption sites, allowing us to quantify their impact in the thickness variability across 300 mm wafer substrates.

physics.app-ph

Fast, Smart Neuromorphic Sensors Based on Heterogeneous Networks and Mixed Encodings

Neuromorphic architectures are ideally suited for the implementation of smart sensors able to react, learn, and respond to a changing environment. Our work uses the insect brain as a model to understand how heterogeneous architectures, incorporating different types of neurons and encodings, can be leveraged to create systems integrating input processing, evaluation, and response. Here we show how the combination of time and rate encodings can lead to fast sensors that are able to generate a hypothesis on the input in only a few cycles and then use that hypothesis as secondary input for more detailed analysis.

cs.NE

Neuromodulated Neural Architectures with Local Error Signals for Memory-Constrained Online Continual Learning

The ability to learn continuously from an incoming data stream without catastrophic forgetting is critical for designing intelligent systems. Many existing approaches to continual learning rely on stochastic gradient descent and its variants. However, these algorithms have to implement various strategies, such as memory buffers or replay, to overcome well-known shortcomings of stochastic gradient descent methods in terms of stability, greed, and short-term memory. To that end, we develop a biologically-inspired light weight neural network architecture that incorporates local learning and neuromodulation to enable input processing over data streams and online learning. Next, we address the challenge of hyperparameter selection for tasks that are not known in advance by implementing transfer metalearning: using a Bayesian optimization to explore a design space spanning multiple local learning rules and their hyperparameters, we identify high performing configurations in classical single task online learning and we transfer them to continual learning tasks with task-similarity considerations. We demonstrate the efficacy of our approach on both single task and continual learning setting. For the single task learning setting, we demonstrate superior performance over other local learning approaches on the MNIST, Fashion MNIST, and CIFAR-10 datasets. Using high performing configurations metalearned in the single task learning setting, we achieve superior continual learning performance on Split-MNIST, and Split-CIFAR-10 data as compared with other memory-constrained learning approaches, and match that of the state-of-the-art memory-intensive replay-based approaches.

cs.LG

Coarse scale representation of spiking neural networks: backpropagation through spikes and application to neuromorphic hardware

In this work we explore recurrent representations of leaky integrate and fire neurons operating at a timescale equal to their absolute refractory period. Our coarse time scale approximation is obtained using a probability distribution function for spike arrivals that is homogeneously distributed over this time interval. This leads to a discrete representation that exhibits the same dynamics as the continuous model, enabling efficient large scale simulations and backpropagation through the recurrent implementation. We use this approach to explore the training of deep spiking neural networks including convolutional, all-to-all connectivity, and maxpool layers directly in Pytorch. We found that the recurrent model leads to high classification accuracy using just 4-long spike trains during training. We also observed a good transfer back to continuous implementations of leaky integrate and fire neurons. Finally, we applied this approach to some of the standard control problems as a first step to explore reinforcement learning using neuromorphic chips.

cs.NE

Optical and structural properties of Si doped $β$-Ga$_2$O$_3$ (010) thin films homoepitaxially grown by halide vapor phase epitaxy

We report the optical, electrical, and structural properties of Si doped $β$-Ga$_2$O$_3$ films grown on (010)-oriented $β$-Ga$_2$O$_3$ substrate via HVPE. Our results show that, despite growth rates that are more than one order of magnitude faster than MOCVD, films with mobility values of up to 95 cm$^2$V$^{-1}$s$^{-1}$ at a carrier concentration of 1.3$\times$10$^{17}$ cm$^{-3}$ can be achieved using this technique, with all Si-doped samples showing n-type behavior with carrier concentrations in the range of 10$^{17}$ to 10$^{19}$ cm$^{-3}$. All samples showed similar room temperature photoluminescence, with only the samples with the lowest carrier concentration showing the presence of a blue luminescence, and the Raman spectra exhibiting only phonon modes that belong to $β$-Ga$_2$O$_3$, indicating that the Ga$_2$O$_3$ films are phase pure and of high crystal quality. We further evaluated the epitaxial quality of the films by carrying out grazing incidence X-ray scattering measurements, which allowed us to discriminate the bulk and film contributions. Finally, MOS capacitors were fabricated using ALD HfO$_2$ to perform C-V measurements. The carrier concentration and dielectric values extracted from the C-V characteristics are in good agreement with Hall probe measurements. These results indicate that HVPE has a strong potential to yield device-quality $β$-Ga$_2$O$_3$ films that can be utilized to develop vertical devices for high-power electronics applications.

physics.app-ph