SearcharxivSearch

arXiv subjects

Manabu Ihara

Publications and source records attributed to Manabu Ihara.

16 recordsLinked to original sources

Band Structure Modulation of ZrO2 Nanoparticles for Control of CO Adsorption Properties: A Combined Density Functional Theory - Density Functional Tight Binding Study

We present a combined density functional theory (DFT) and density functional tight binding (DFTB) study of zirconia (ZrO2) nanoparticles of experimentally relevant sizes of several nanometers and their interactions with the CO molecule. A hybrid DFTB - Force Field (DFTB-FF) framework is developed, whereby band structure calculations rely on an existing Slater-Koster framework, while the accuracy of structural optimization and adsorption properties is controlled by the introduction of classical long-range interatomic potentials into DFTB instead of the traditional repulsive potentials. Additionally, coordination-dependent Zr-C potentials are introduced to account for the distinct local chemical environments of bulk-like facet sites and under-coordinated tip and edge sites, thereby improving the description of CO adsorption. This hybrid DFTB-FF approach substantially improves the robustness of geometry optimization and provides a practical strategy for extending the applicability of DFTB to complex oxide nanostructures. The calculations reveal that termination stoichiometry can be used to engineer intrinsic, p-type, or n-type electronic structures and thereby tune the adsorption activity of zirconia nanoparticles. While stoichiometric nanoparticles do not activate the C-O bond, low-coordinated sites in Zr-rich (n-type) nanoparticles exhibit chemisorption accompanied by charge donation into a CO antibonding LUMO-derived orbital, resulting in C-O bond activation. These results demonstrate that stoichiometry-controlled electronic structure and under-coordinated surface sites introduced by nanostructuring play a key role in governing the adsorption strength and reactivity of zirconia nanoparticles.

cond-mat.mtrl-sci

Neural networks for neurocomputing circuits: a computational study of tolerance to noise and activation function non-uniformity when machine learning materials properties

Dedicated analog neurocomputing circuits are promising for high-throughput, low power consumption applications of machine learning (ML) and for applications where implementing a digital computer is unwieldy (remote locations; small, mobile, and autonomous devices, extreme conditions, etc.). Neural networks (NN) implemented in such circuits, however, must contend with circuit noise and the non-uniform shapes of the neuron activation function (NAF) due to the dispersion of performance characteristics of circuit elements (such as transistors or diodes implementing the neurons). We present a computational study of the impact of circuit noise and NAF inhomogeneity in function of NN architecture and training regimes. We focus on one application that requires high-throughput ML: materials informatics, using as representative problem ML of formation energies vs. lowest-energy isomer of peri-condensed hydrocarbons, formation energies and band gaps of double perovskites, and zero point vibrational energies of molecules from QM9 dataset. We show that NNs generally possess low noise tolerance with the model accuracy rapidly degrading with noise level. Single-hidden layer NNs, and NNs with larger-than-optimal sizes are somewhat more noise-tolerant. Models that show less overfitting (not necessarily the lowest test set error) are more noise-tolerant. Importantly, we demonstrate that the effect of activation function inhomogeneity can be palliated by retraining the NN using practically realized shapes of NAFs.

cs.NE

Gaussian Process Regression -- Neural Network Hybrid with Optimized Redundant Coordinates

Recently, a Gaussian Process Regression - neural network (GPRNN) hybrid machine learning method was proposed, which is based on additive-kernel GPR in redundant coordinates constructed by rules [J. Phys. Chem. A 127 (2023) 7823]. The method combined the expressive power of an NN with the robustness of linear regression, in particular, with respect to overfitting when the number of neurons is increased beyond optimal. We introduce opt-GPRNN, in which the redundant coordinates of GPRNN are optimized with a Monte Carlo algorithm and show that when combined with optimization of redundant coordinates, GPRNN attains the lowest test set error with much fewer terms / neurons and retains the advantage of avoiding overfitting when the number of neurons is increased beyond optimal value. The method, opt-GPRNN possesses an expressive power closer to that of a multilayer NN and could obviate the need for deep NNs in some applications. With optimized redundant coordinates, a dimensionality reduction regime is also possible. Examples of application to machine learning an interatomic potential and materials informatics are given.

stat.ML

Machine learning of kinetic energy densities with target and feature averaging: better results with fewer training data

Machine learning of kinetic energy functionals (KEF), in particular kinetic energy density (KED) functionals, has recently attracted attention as a promising way to construct KEFs for orbital-free density functional theory (OF-DFT). Neural networks (NN) and kernel methods including Gaussian process regression (GPR) have been used to learn Kohn-Sham (KS) KED from density-based descriptors derived from KS DFT calculations. The descriptors are typically expressed as functions of different powers and derivatives of the electron density. This can generate large and extremely unevenly distributed datasets, which complicates effective application of machine learning techniques. Very uneven data distributions require many training data points, can cause overfitting, and ultimately lower the quality of a ML KED model. We show that one can produce more accurate ML models from fewer data by working with partially averaged density-dependent variables and KED. Averaging palliates the issue of very uneven data distributions and associated difficulties of sampling, while retaining enough spatial structure necessary for working within the paradigm of KEDF. We use GPR as a function of partially spatially averaged terms of the 4th order gradient expansion and the Kohn-Sham effective potential and obtain accurate and stable (with respect to different random choices of training points) kinetic energy models for Al, Mg, and Si simultaneously from as few as 2000 samples (about 0.3% of the total KS DFT data). In particular, accuracies on the order of 1% in a measure of the quality of energy-volume dependence B' = \frac{E(V_0-ΔV)-2E(V_0)+E(V_0 + ΔV)}{\big(\frac{ΔV}{V_0}\big)^2} are obtained simultaneously for all three materials.

cond-mat.mtrl-sci

Machine learning-guided construction of an analytic kinetic energy functional for orbital free density functional theory

Machine learning (ML) of kinetic energy functionals (KEF) for orbital-free density functional theory (OF-DFT) holds the promise of addressing an important bottleneck in large-scale ab initio materials modeling where sufficiently accurate analytic KEFs are lacking. However, ML models are not as easily handled as analytic expressions; they need to be provided in the form of algorithms and associated data. Here, we bridge the two approaches and construct an analytic expression for a KEF guided by interpretative machine learning of crystal cell-averaged kinetic energy densities (τ) of several hundred materials. A previously published dataset including multiple phases of 433 unary, binary, and ternary compounds containing Li, Al, Mg, Si, As, Ga, Sb, Na, Sn, P, and In was used for training, including data at the equilibrium geometry as well as strained structures. A hybrid Gaussian process regression - neural network (GPR-NN) method was used to understand the type of functional dependence of τ on the features which contained cell-averaged terms of the 4th order gradient expansion and the product of the electron density and Kohn-Sham effective potential. Based on this analysis, an analytic model is constructed that can reproduce Kohn-Sham DFT energy-volume curves with sufficient accuracy (pronounced minima that are sufficiently close to the minima of the Kohn-Sham DFT-based curves and with sufficiently close curvatures) to enable structure optimizations and elastic response calculations.

cond-mat.mtrl-sci

A machine-learned kinetic energy model for light weight metals and compounds of group III-V elements

We present a machine-learned (ML) model of kinetic energy for orbital-free density functional theory (OF-DFT) suitable for bulk light weight metals and compounds made of group III-V elements. The functional is machine-learned with Gaussian process regression (GPR) from data computed with Kohn-Sham DFT with plane wave bases and local pseudopotentials. The dataset includes multiple phases of unary, binary, and ternary compounds containing Li, Al, Mg, Si, As, Ga, Sb, Na, Sn, P, and In. A total of 433 materials were used for training, and 18 strained structures were used for each material. Averaged (over the unit cell) kinetic energy density is fitted as a function of averaged terms of the 4th order gradient expansion and the product of the density and effective potential. The kinetic energy predicted by the model allows reproducing energy-volume curves around equilibrium geometry with good accuracy. We show that the GPR model beats linear and polynomial regressions. We also find that unary compounds sample a wider region of the descriptor space than binary and ternary compounds, and it is therefore important to include them in the training set; a GPR model trained on a small number of unary compounds is able to extrapolate relatively well to binary and ternary compounds but not vice versa.

cond-mat.mtrl-sci

Machine learning the screening factor in the soft bond valence approach for rapid crystal structure estimation

Development of new functional ceramics is important for several applications, including electrochemical batteries and fuel cells. Computational prescreening and selection of such materials can help discover novel materials but is challenging due to the high cost of electronic structure calculations which would be needed to compute the structures and properties of interest such as the material's stability and ion diffusion properties. The soft bond valence (SoftBV) approach is attractive for rapid prescreening among multiple compositions and structures, but the simplicity of the approximation can make the results inaccurate. We explore the possibility of enhancing the accuracy of the SoftBV approach when estimating crystal structures by adapting the parameters of the approximation to the chemical composition. Specifically, on the examples of perovskite- and spinel-type oxides that have been proposed as promising solid-state ionic conductors, the screening factor, an independent parameter of the SoftBV approximation, is modeled using linear and non-linear methods as a function of descriptors of chemical composition. We find that making the screening factor a function of composition can noticeably improve the ability of SoftBV to correctly model structures, in particular new, putative crystal structures whose structural parameters are yet unknown. We also analyze the relative importance of nonlinearity and coupling in improving the model and find that while the quality of the model is improved by including nonlinearity, coupling is relatively unimportant. While using a neural network showed no improvement over linear regression, the recently proposed GPR-NN method that is a hybrid between a single hidden layer neural network and kernel regression showed substantial improvement, enabling the prediction of structural parameters of new ceramics with accuracy on the order of 1%.

cond-mat.mtrl-sci

Degeneration of kernel regression with Matern kernels into low-order polynomial regression in high dimension

Kernel methods such as kernel ridge regression and Gaussian process regressions with Matern type kernels have been increasingly used, in particular, to fit potential energy surfaces (PES) and density functionals, and for materials informatics. When the dimensionality of the feature space is high, these methods are used with necessarily sparse data. In this regime, the optimal length parameter of a Matern-type kernel tends to become so large that the method effectively degenerates into a low-order polynomial regression and therefore loses any advantage over such regression. This is demonstrated theoretically as well as numerically on the examples of six- and fifteen-dimensional molecular PES using squared exponential and simple exponential kernels. The results shed additional light on the success of polynomial approximations such as PIP for medium size molecules and on the importance of orders-of-coupling based models for preserving the advantages of kernel methods with Matern type kernels or on the use of physically-motivated (reproducing) kernels.

physics.comp-ph

Orders-of-coupling representation with a single neural network with optimal neuron activation functions and without nonlinear parameter optimization

Representations of multivariate functions with low-dimensional functions that depend on subsets of original coordinates (corresponding of different orders of coupling) are useful in quantum dynamics and other applications, especially where integration is needed. Such representations can be conveniently built with machine learning methods, and previously, methods building the lower-dimensional terms of such representations with neural networks [e.g. Comput. Phys. Comm. 180 (2009) 2002] and Gaussian process regressions [e.g. Mach. Learn. Sci. Technol. 3 (2022) 01LT02] were proposed. Here, we show that neural network models of orders-of-coupling representations can be easily built by using a recently proposed neural network with optimal neuron activation functions computed with a first-order additive Gaussian process regression [arXiv:2301.05567] and avoiding non-linear parameter optimization. Examples are given of representations of molecular potential energy surfaces.

cs.LG

Neural network with optimal neuron activation functions based on additive Gaussian process regression

Feed-forward neural networks (NN) are a staple machine learning method widely used in many areas of science and technology. While even a single-hidden layer NN is a universal approximator, its expressive power is limited by the use of simple neuron activation functions (such as sigmoid functions) that are typically the same for all neurons. More flexible neuron activation functions would allow using fewer neurons and layers and thereby save computational cost and improve expressive power. We show that additive Gaussian process regression (GPR) can be used to construct optimal neuron activation functions that are individual to each neuron. An approach is also introduced that avoids non-linear fitting of neural network parameters. The resulting method combines the advantage of robustness of a linear regression with the higher expressive power of a NN. We demonstrate the approach by fitting the potential energy surfaces of the water molecule and formaldehyde. Without requiring any non-linear optimization, the additive GPR based approach outperforms a conventional NN in the high accuracy regime, where a conventional NN suffers more from overfitting.

stat.ML

The loss of the property of locality of the kernel in high-dimensional Gaussian process regression on the example of the fitting of molecular potential energy surfaces

Kernel based methods including Gaussian process regression (GPR) and generally kernel ridge regression (KRR) have been finding increasing use in computational chemistry, including the fitting of potential energy surfaces and density functionals in high-dimensional feature spaces. Kernels of the Matern family such as Gaussian-like kernels (basis functions) are often used, which allows imparting them the meaning of covariance functions and formulating GPR as an estimator of the mean of a Gaussian distribution. The notion of locality of the kernel is critical for this interpretation. It is also critical to the formulation of multi-zeta type basis functions widely used in computational chemistry We show, on the example of fitting of molecular potential energy surfaces of increasing dimensionality, the practical disappearance of the property of locality of a Gaussian-like kernel in high dimensionality. We also formulate a multi-zeta approach to the kernel and show that it significantly improves the quality of regression in low dimensionality but loses any advantage in high dimensionality, which is attributed to the loss of the property of locality.

stat.ML

Rectangularization of Gaussian process regression for optimization of hyperparameters

Gaussian process regression (GPR) is a powerful machine learning method which has recently enjoyed wider use, in particular in physical sciences. In its original formulation, GPR uses a square matrix of covariances among training data and can be viewed as linear regression problem with equal numbers of training data and basis functions. When data are sparse, avoidance of overfitting and optimization of hyperparameters of GPR are difficult, in particular in high-dimensional spaces where the data sparsity issue cannot practically be resolved by adding more data. Optimal choice of hyperparameters, however, determines success or failure of the application of the GPR method. We show that parameter optimization is facilitated by rectangularization of the defining equation of GPR. On the example of a 15-dimensional molecular potential energy surface we demonstrate that this approach allows effective hyperparameter tuning even with very sparse data.

math.NA

Non-invasive improvement of machining by reversible electrochemical doping: a proof of principle with computational modeling

We propose that the machinability of hard ceramics can be improved by reversible electrochemical doping. On the example of TiO2, we show in a combined density functional theory-molecular dynamics computational study that a small amount of intercalated lithium, which preserves the host structure and can be introduced reversibly, leads to a lowering of the strength of work materials and the cutting force. This is in spite of the fact that there are no significant modifications of the elastic constants at room temperature, i.e. the effect is mostly on plastic properties. This approach is expected to be applicable to a class of ceramics exhibiting similar mechanisms of host-dopant interactions and presents a reversible and non-destructive way of modifying mechanical properties.

cond-mat.mtrl-sci

On the optimization of hyperparameters in Gaussian process regression with the help of low-order high-dimensional model representation

When the data are sparse, optimization of hyperparameters of the kernel in Gaussian process regression by the commonly used maximum likelihood estimation (MLE) criterion often leads to overfitting. We show that choosing hyperparameters (in this case, kernel length parameter and regularization parameter) based on a criterion of the completeness of the basis in the corresponding linear regression problem is superior to MLE. We show that this is facilitated by the use of high-dimensional model representation (HDMR) whereby a low-order HDMR representation can provide reliable reference functions and large synthetic test data sets needed for basis parameter optimization even when the original data are few.

stat.ME

Easy representation of multivariate functions with low-dimensional terms via Gaussian process regression kernel design: applications to machine learning of potential energy surfaces and kinetic energy densities from sparse data

We show that Gaussian process regression (GPR) allows representing multivariate functions with low-dimensional terms via kernel design. When using a kernel built with HDMR (High-dimensional model representation), one obtains a similar type of representation as the previously proposed HDMR-GPR scheme while being faster and simpler to use. We tested the approach on cases where highly accurate machine learning is required from sparse data by fitting potential energy surfaces and kinetic energy densities.

math.NA

Random Sampling High Dimensional Model Representation Gaussian Process Regression (RS-HDMR-GPR) for representing multidimensional functions with machine-learned lower-dimensional terms allowing insight with a general method

We present a Python implementation for RS-HDMR-GPR (Random Sampling High Dimensional Model Representation Gaussian Process Regression). The method builds representations of multivariate functions with lower-dimensional terms, either as an expansion over orders of coupling or using terms of only a given dimensionality. This facilitates, in particular, recovering functional dependence from sparse data. The code also allows for imputation of missing values of the variables and for a significant pruning of the useful number of HDMR terms. The code can also be used for estimating relative importance of different combinations of input variables, thereby adding an element of insight to a general machine learning method. The capabilities of this regression tool are demonstrated on test cases involving synthetic analytic functions, the potential energy surface of the water molecule, kinetic energy densities of materials (crystalline magnesium, aluminum, and silicon), and financial market data.

stat.CO