SearcharxivSearch

arXiv subjects

Angela Violi

Publications and source records attributed to Angela Violi.

10 recordsLinked to original sources

Machine learning models for Si nanoparticle growth in nonthermal plasma

Nanoparticles (NPs) formed in nonthermal plasmas (NTPs) can have unique properties and applications. However, modeling their growth in these environments presents significant challenges due to the non-equilibrium nature of NTPs, making them computationally expensive to describe. In this work, we address the challenges associated with accelerating the estimation of parameters needed for these models. Specifically, we explore how different machine learning models can be tailored to improve prediction outcomes. We apply these methods to reactive classical molecular dynamics data, which capture the processes associated with colliding silane fragments in NTPs. These reactions exemplify processes where qualitative trends are clear, but their quantification is challenging, hard to generalize, and requires time-consuming simulations. Our results demonstrate that good prediction performance can be achieved when appropriate loss functions are implemented and correct invariances are imposed. While the diversity of molecules used in the training set is critical for accurate prediction, our findings indicate that only a fraction (15-25\%) of the energy and temperature sampling is required to achieve high levels of accuracy. This suggests a substantial reduction in computational effort is possible for similar systems.

physics.comp-ph

Joint Optimization of Piecewise Linear Ensembles

Tree ensembles achieve state-of-the-art performance on numerous prediction tasks. We propose $\textbf{J}$oint $\textbf{O}$ptimization of $\textbf{P}$iecewise $\textbf{L}$inear $\textbf{En}$sembles (JOPLEn), which jointly fits piecewise linear models at all leaf nodes of an existing tree ensemble. In addition to enhancing the ensemble expressiveness, JOPLEn allows several common penalties, including sparsity-promoting and subspace-norms, to be applied to nonlinear prediction. For example, JOPLEn with a nuclear norm penalty learns subspace-aligned functions. Additionally, JOPLEn (combined with a Dirty LASSO penalty) is an effective feature selection method for nonlinear prediction in multitask learning. Finally, we demonstrate the performance of JOPLEn on 153 regression and classification datasets and with a variety of penalties. JOPLEn leads to improved prediction performance relative to not only standard random forest and boosted tree ensembles, but also other methods for enhancing tree ensembles.

cs.LG

Universal Feature Selection for Simultaneous Interpretability of Multitask Datasets

Extracting meaningful features from complex, high-dimensional datasets across scientific domains remains challenging. Current methods often struggle with scalability, limiting their applicability to large datasets, or make restrictive assumptions about feature-property relationships, hindering their ability to capture complex interactions. BoUTS's general and scalable feature selection algorithm surpasses these limitations to identify both universal features relevant to all datasets and task-specific features predictive for specific subsets. Evaluated on seven diverse chemical regression datasets, BoUTS achieves state-of-the-art feature sparsity while maintaining prediction accuracy comparable to specialized methods. Notably, BoUTS's universal features enable domain-specific knowledge transfer between datasets, and suggest deep connections in seemingly-disparate chemical datasets. We expect these results to have important repercussions in manually-guided inverse problems. Beyond its current application, BoUTS holds immense potential for elucidating data-poor systems by leveraging information from similar data-rich systems. BoUTS represents a significant leap in cross-domain feature selection, potentially leading to advancements in various scientific fields.

cs.LG

Antiviral Drug-Membrane Permeability: the Viral Envelope and Cellular Organelles

To shorten the time required to find effective new drugs, like antivirals, a key parameter to consider is membrane permeability, as a compound intended for an intracellular target with poor permeability will have low efficacy. Here, we present a computational model that considers both drug characteristics and membrane properties for the rapid assessment of drugs permeability through the coronavirus envelope and various cellular membranes. We analyze 79 drugs that are considered as potential candidates for the treatment of SARS-CoV-2 and determine their time of permeation in different organelle membranes grouped by viral baits and mammalian processes. The computational results are correlated with experimental data, present in the literature, on bioavailability of the drugs, showing a negative correlation between fast permeation and most promising drugs. This model represents an important tool capable of evaluating how permeability affects the ability of compounds to reach both intended and unintended intracellular targets in an accurate and rapid way. The method is general and flexible and can be employed for a variety of molecules, from small drugs to nanoparticles, as well to a variety of biological membranes.

q-bio.BM

On sparse identification of complex dynamical systems: A study on discovering influential reactions in chemical reaction networks

A wide variety of real life complex networks are prohibitively large for modeling, analysis and control. Understanding the structure and dynamics of such networks entails creating a smaller representative network that preserves its relevant topological and dynamical properties. While modern machine learning methods have enabled identification of governing laws for complex dynamical systems, their inability to produce white-box models with sufficient physical interpretation renders such methods undesirable to domain experts. In this paper, we introduce a hybrid black-box, white-box approach for the sparse identification of the governing laws for complex, highly coupled dynamical systems with particular emphasis on finding the influential reactions in chemical reaction networks for combustion applications, using a data-driven sparse-learning technique. The proposed approach identifies a set of influential reactions using species concentrations and reaction rates,with minimal computational cost without requiring additional data or simulations. The new approach is applied to analyze the combustion chemistry of H2 and C3H8 in a constant-volume homogeneous reactor. The influential reactions determined by the sparse-learning method are consistent with the current kinetics knowledge of chemical mechanisms. Additionally, we show that a reduced version of the parent mechanism can be generated as a combination of the significantly reduced influential reactions identified at different times and conditions and that for both H2 and C3H8 fuel, the reduced mechanisms perform closely to the parent mechanisms as a function of the ignition delay time over a wide range of conditions. Our results demonstrate the potential of the sparse-learning approach as an effective and efficient tool for dynamical system analysis and reduction. The uniqueness of this approach as applied to combustion systems lies in the ability to identify influential reactions in specified conditions and times during the evolution of the combustion process. This ability is of great interest to understand chemical reaction systems.

physics.chem-ph

Spatial Dependence of Polycyclic Aromatic Compounds Growth in Counterflow Flames

The formation mechanisms of aromatic compounds in flame are strongly influenced by the chemical and thermal history that leads to their formation. Indeed, the complex environments that characterize combustion systems do not only affect the composition of gas-phase species, but they also determine the structure and the characteristics of the soot precursors generated. To illustrate the importance of these effects, in this work we investigate the growth mechanisms of soot precursors in an atmospheric-pressure ethylene/oxygen/argon counterflow diffusion flame, using a combination of computational and experimental techniques. In diffusion flames, flow characteristics play an important role in the formation, growth, and oxidation of particles, and soot precursors are strongly affected by the flame location. Fluid dynamics simulations and stochastic discrete modeling were employed together to identify key reaction pathways along various flow streamlines. The models were validated with experimental mass spectra obtained using aerosol mass spectrometry coupled with vacuum-ultraviolet photoionization. Results show that both the hydrogen-abstraction-acetylene-addition mechanism and oxygen-insertion reactions are responsible for the molecular growth, and their relative importance is determined by the flame conditions along the streamlines. Oxygenated species were detected in regions of high temperature, high atomic oxygen concentration, and relatively low acetylene abundance. This study also emphasizes the need to model the counterflow flame in three dimensions to capture the spatial dependence on growth mechanisms of soot precursors.

physics.chem-ph

The Role of Molecular Properties on the Dimerization of Aromatic Compounds

Recent results have shown the presence and importance of oxygen chemistry during the growth of aromatic compounds, leading to the formation of oxygenated structures that have been identified in various environments. Since the formation of polycyclic aromatic compounds (PAC) bridge the formation of gas-phase species with particle inception, in this work we report a detailed analysis of the effects of molecular characteristics on physical growth of PAC via dimerization. We have included oxygen content, mass, type of bonds (rigid versus rotatable), and shape as main properties of the molecules and studied their effect on the propensity of these structures to form homo-molecular and hetero-molecular dimers. Using enhanced sampling molecular dynamics techniques, we have computed the free energy of dimerization in the temperature range $500-1680$~K. Initial structures used in this study were obtained from experimental data. The results show that although the effects of shape, presence of oxygen, mass, and internal bonds are tightly intertwined, and their relative importance changes with temperature. In general, mass and the presence of rotatable bonds are the best indicators. The results provide knowledge on the inception step and the role that particle characteristics play during inception. In addition, our study highlights the fact that current models that use stabilomers as monomers for physical aggregation are overestimating the importance of this process during particle nucleation.

physics.chem-ph

A New Data-Driven Sparse-Learning Approach to Study Chemical Reaction Networks

Chemical kinetic mechanisms can be represented by sets of elementary reactions that are easily translated into mathematical terms using physicochemical relationships. The schematic representation of reactions captures the interactions between reacting species and products. Determining the minimal chemical interactions underlying the dynamic behavior of systems is a major task. In this paper, we introduce a novel approach for the identification of the influential reactions in chemical reaction networks for combustion applications, using a data-driven sparse-learning technique. The proposed approach identifies a set of influential reactions using species concentrations and reaction rates, with minimal computational cost without requiring additional data or simulations. The new approach is applied to analyze the combustion chemistry of H2 and C3H8 in a constant-volume homogeneous reactor. The influential reactions identified by the sparse-learning method are consistent with the current kinetics knowledge of chemical mechanisms. Additionally, we show that a reduced version of the parent mechanism can be generated as a combination of the influential reactions identified at different times and conditions and that for both H2 and C3H8 this reduced mechanism performs closely to the parent mechanism as a function of ignition delay over a wide range of conditions. Our results demonstrate the potential of the sparse-learning approach as an effective and efficient tool for mechanism analysis and mechanism reduction.

math.OC

A Data-Driven Sparse-Learning Approach to Model Reduction in Chemical Reaction Networks

In this paper, we propose an optimization-based sparse learning approach to identify the set of most influential reactions in a chemical reaction network. This reduced set of reactions is then employed to construct a reduced chemical reaction mechanism, which is relevant to chemical interaction network modeling. The problem of identifying influential reactions is first formulated as a mixed-integer quadratic program, and then a relaxation method is leveraged to reduce the computational complexity of our approach. Qualitative and quantitative validation of the sparse encoding approach demonstrates that the model captures important network structural properties with moderate computational load.

math.OC

Fast exploration of chemical reaction networks

A variety of natural phenomena comprises a huge number of competing reactions and short-lived intermediates. Any study of such processes requires the discovery and accurate modeling of their underlying reaction network. However, this task is challenging due to the complexity in exploring all the possible pathways and the high computational cost in accurately modeling a large number of reactions. Fortunately, very often these processes are dominated by only a limited subset of the network's reaction pathways. In this work we propose a novel computationally inexpensive method to identify and select the key pathways of complex reaction networks, so that high-level ab-initio calculations can be more efficiently targeted at these critical reactions. The method estimates the relative importance of the reaction pathways for given reactants by analyzing the accelerated evolution of hundreds of replicas of the system and detecting products formation. This acceleration-detection method is able to tremendously speed up the reactivity of uni- and bimolecular reactions, without requiring any previous knowledge of products or transition states. Importantly, the method is efficiently iterative, as it can be straightforwardly applied for the most frequently observed products, therefore providing an efficient algorithm to identify the key reactions of extended chemical networks. We verified the validity of our approach on three different systems, including the reactivity of t-decalin with a methyl radical, and in all cases the expected behavior was recovered within statistical error.

physics.chem-ph