SearcharxivSearch

arXiv subjects

Bernard R. Brooks

Publications and source records attributed to Bernard R. Brooks.

7 recordsLinked to original sources

Boltzmann Classifier: A Thermodynamic-Inspired Approach to Supervised Learning

We present the Boltzmann classifier, a novel distance based probabilistic classification algorithm inspired by the Boltzmann distribution. Unlike traditional classifiers that produce hard decisions or uncalibrated probabilities, the Boltzmann classifier assigns class probabilities based on the average distance to the nearest neighbors within each class, providing interpretable, physically meaningful outputs. We evaluate the performance of the method across three application domains: molecular activity prediction, oxidation state classification of transition metal complexes, and breast cancer diagnosis. In the molecular activity task, the classifier achieved the highest accuracy in predicting active compounds against two protein targets, with strong correlations observed between the predicted probabilities and experimental pIC50 values. For metal complexes, the classifier accurately distinguished between oxidation states II and III for Fe, Mn, and Co, using only metal-ligand bond lengths extracted from crystallographic data, and demonstrated high consistency with known chemical trends. In the breast cancer dataset, the classifier achieved 97% accuracy, with low confidence predictions concentrated in inherently ambiguous cases. Across all tasks, the Boltzmann classifier performed competitively or better than standard models such as logistic regression, support vector machines, random forests, and k-nearest neighbors. Its probabilistic outputs were found to correlate with continuous physical or biological properties, highlighting its potential utility in both classification and regression contexts. The results suggest that the Boltzmann classifier is a robust and interpretable alternative to conventional machine learning approaches, particularly in scientific domains where underlying structure property relationships are important.

cs.LG

GraphVAMPNet, using graph neural networks and variational approach to markov processes for dynamical modeling of biomolecules

Finding low dimensional representation of data from long-timescale trajectories of biomolecular processes such as protein-folding or ligand-receptor binding is of fundamental importance and kinetic models such as Markov modeling have proven useful in describing the kinetics of these systems. Recently, an unsupervised machine learning technique called VAMPNet was introduced to learn the low dimensional representation and linear dynamical model in an end-to-end manner. VAMPNet is based on variational approach to Markov processes (VAMP) and relies on neural networks to learn the coarse-grained dynamics. In this contribution, we combine VAMPNet and graph neural networks to generate an end-to-end framework to efficiently learn high-level dynamics and metastable states from the long-timescale molecular dynamics trajectories. This method bears the advantages of graph representation learning and uses graph message passing operations to generate an embedding for each datapoint which is used in the VAMPNet to generate a coarse-grained representation. This type of molecular representation results in a higher resolution and more interpretable Markov model than the standard VAMPNet enabling a more detailed kinetic study of the biomolecular processes. Our GraphVAMPNet approach is also enhanced with an attention mechanism to find the important residues for classification into different metastable states.

physics.comp-ph

Variational embedding of protein folding simulations using gaussian mixture variational autoencoders

Conformational sampling of biomolecules using molecular dynamics simulations often produces large amount of high dimensional data that makes it difficult to interpret using conventional analysis techniques. Dimensionality reduction methods are thus required to extract useful and relevant information. Here we devise a machine learning method, Gaussian mixture variational autoencoder (GMVAE) that can simultaneously perform dimensionality reduction and clustering of biomolecular conformations in an unsupervised way. We show that GMVAE can learn a reduced representation of the free energy landscape of protein folding with highly separated clusters that correspond to the metastable states during folding. Since GMVAE uses a mixture of Gaussians as the prior, it can directly acknowledge the multi-basin nature of protein folding free-energy landscape. To make the model end-to-end differentialble, we use a Gumbel-softmax distribution. We test the model on three long-timescale protein folding trajectories and show that GMVAE embedding resembles the folding funnel with folded states down the funnel and unfolded states outer in the funnel path. Additionally, we show that the latent space of GMVAE can be used for kinetic analysis and Markov state models built on this embedding produce folding and unfolding timescales that are in close agreement with other rigorous dynamical embeddings such as time independent component analysis (TICA).

q-bio.BM

Finding dominant transition pathways via global optimization of action

We present a new computational approach, Action-CSA, to sample multiple reaction pathways with fixed initial and final states through global optimization of the Onsager-Machlup action using the conformational space annealing method. This approach successfully samples not only the most dominant pathway but also many other possible paths without initial guesses on reaction pathways. Pathway space is efficiently sampled by crossover operations of a set of paths and preserving the diversity of sampled pathways. The sampling ability of the approach is assessed by finding pathways for the conformational changes of alanine dipeptide and hexane. The benchmarks demonstrate that the rank order and the transition time distribution of multiple pathways identified by the new approach are in good agreement with those of long molecular dynamics simulations. We also show that the lowest action folding pathway of the mini-protein FSD-1 identified by the new approach is consistent with previous molecular dynamics simulations and experiments.

physics.chem-ph

Link community detection through global optimization and the inverse resolution limit of partition density

We investigate the possibility of global optimization-based overlapping community detection, using link community framework. We first show that partition density, the original quality function used in link community detection method, is not suitable as a quality function for global optimization because it prefers breaking communities into triangles except in highly limited conditions. We analytically derive those conditions and confirm it with computational results on direct optimization of various synthetic and real-world networks. To overcome this limitation, we propose alternative approaches combining the weighted line graph transformation and existing quality functions for node-based communities. We suggest a new line graph weighting scheme, a normalized Jaccard index. Computational results show that community detection using the weighted line graphs generated with the normalized Jaccard index leads to a more accurate community structure.

physics.soc-ph

Influence of Nanoparticle Size and Shape on Oligomer Formation of an Amyloidogenic Peptide

Understanding the influence of macromolecular crowding and nanoparticles on the formation of in-register $β$-sheets, the primary structural component of amyloid fibrils, is a first step towards describing \emph{in vivo} protein aggregation and interactions between synthetic materials and proteins. Using all atom molecular simulations in implicit solvent we illustrate the effects of nanoparticle size, shape, and volume fraction on oligomer formation of an amyloidogenic peptide from the transthyretin protein. Surprisingly, we find that inert spherical crowding particles destabilize in-register $β$-sheets formed by dimers while stabilizing $β$-sheets comprised of trimers and tetramers. As the radius of the nanoparticle increases crowding effects decrease, implying smaller crowding particles have the largest influence on the earliest amyloid species. We explain these results using a theory based on the depletion effect. Finally, we show that spherocylindrical crowders destabilize the ordered $β$-sheet dimer to a greater extent than spherical crowders, which underscores the influence of nanoparticle shape on protein aggregation.

cond-mat.soft

Molecular origin of constant m-values, denatured state collapse, and residue-dependent transition midpoints in globular proteins

Experiments show that for many two state folders the free energy of the native state DG_ND([C]) changes linearly as the denaturant concentration [C] is varied. The slope, m = d DG_ND([C])/d[C], is nearly constant. The m-value is associated with the difference in the surface area between the native (N) and the denatured (D) state, which should be a function of DR_g^2, the difference in the square of the radius of gyration between the D and N states. Single molecule experiments show that the denatured state undergoes an equilibrium collapse transition as [C] decreases, which implies m also should be [C]-dependent. We resolve the conundrum between constant m-values and [C]-dependent changes in Rg using molecular simulations of a coarse-grained representation of protein L, and the Molecular Transfer Model, for which the equilibrium folding can be accurately calculated as a function of denaturant concentration. We find that over a large range of denaturant concentration (> 3 M) the m-value is a constant, whereas under strongly renaturing conditions (< 3 M) it depends on [C]. The m-value is a constant above [C]> 3 M because the [C]-dependent changes in the surface area of the backbone groups, which make the largest contribution to m, is relatively narrow in the denatured state. The burial of the backbone gives rise to substantial surface area changes below [C]< 3 M, leading to collapse in the denatured state. The midpoint of transition of individual residues vary significantly even though global folding can be described as an all-or-none transition. Collapse driven by the loss of favorable residue-solvent interactions and a concomitant increase in the strength of intrapeptide interactions with decreasing [C]. These interactions are non-uniformly distributed throughout the native structure of protein L.

q-bio.BM