SearcharxivSearch

arXiv subjects

Marco Giulini

Publications and source records attributed to Marco Giulini.

7 recordsLinked to original sources

Integrative modelling of protein-glycan interactions with HADDOCK3

Glycans are structurally diverse and flexible biomolecules that play key roles in many biological processes. Their conformational variability makes the modeling of their interactions with proteins particularly challenging. This chapter presents a step-by-step protocol for modeling protein-glycan interactions using HADDOCK3, an integrative modeling platform that supports the inclusion of experimental or predicted interaction restraints and allows for flexible refinement of the solutions. The workflow is illustrated using the interaction between a linear homopolymer glycan, 4-beta-glucopyranose, and the catalytic domain of the Humicola grisea Cel12A enzyme, for which an experimental X-ray structure is available as a reference. Detailed instructions are provided for input structure preparation, restraint definition, docking setup, execution, and result analysis. Application of the protocol starting from unbound structures yields models of acceptable to medium quality, with interface-ligand RMSD values below 3 angstroms. Although illustrated on a specific system, the protocol has been optimized and benchmarked on multiple protein-glycan complexes and is broadly applicable to similar systems, providing a framework for integrative modeling of protein-glycan interactions.

q-bio.BM

EXCOGITO, an extensible coarse-graining toolbox for the investigation of biomolecules by means of low-resolution representation

Bottom-up coarse-grained (CG) models proved to be essential to complement and sometimes even replace all-atom representations of soft matter systems and biological macromolecules. The development of low-resolution models takes the moves from the reduction of the degrees of freedom employed, that is, the definition of a mapping between a system's high-resolution description and its simplified counterpart. Even in the absence of an explicit parametrisation and simulation of a CG model, the observation of the atomistic system in simpler terms can be informative: this idea is leveraged by the mapping entropy, a measure of the information loss inherent to the process of coarsening. Mapping entropy lies at the heart of the extensible coarse-graining toolbox, or EXCOGITO, developed to perform a number of operations and analyses on molecular systems pivoting around the properties of mappings. EXCOGITO can process an all-atom trajectory to compute the mapping entropy, identify the mapping that minimizes it, and establish quantitative relations between a low-resolution representation and the geometrical, structural, and energetic features of the system. Here, the software, which is available free of charge under an open-source licence, is presented and showcased to introduce potential users to its capabilities and usage. Published on the J. Chem. Inf. Model. on June 11, 2024. DOI: https://doi.org/10.1021/acs.jcim.4c00490

cond-mat.soft

Making sense of complex systems through resolution, relevance, and mapping entropy

Complex systems are characterised by a tight, nontrivial interplay of their constituents, which gives rise to a multi-scale spectrum of emergent properties. In this scenario, it is practically and conceptually difficult to identify those degrees of freedom that mostly determine the behaviour of the system and separate them from less prominent players. Here, we tackle this problem making use of three measures of statistical information: resolution, relevance, and mapping entropy. We address the links existing among them, taking the moves from the established relation between resolution and relevance and further developing novel connections between resolution and mapping entropy; by these means we can identify, in a quantitative manner, the number and selection of degrees of freedom of the system that preserve the largest information content about the generative process that underlies an empirical dataset. The method, which is implemented in a freely available software, is fully general, as it is shown through the application to three very diverse systems, namely a toy model of independent binary spins, a coarse-grained representation of the financial stock market, and a fully atomistic simulation of a protein.

cond-mat.stat-mech

A mapping space Odyssey: characterising the statistical and metric properties of reduced representations of macromolecules

Simplified representations of macromolecules help in rationalising and understanding the outcome of atomistic simulations, and serve to the construction of effective, coarse-grained models. The number and distribution of coarse-grained sites bears a strict relation with the amount of information conveyed by the representation and the accuracy of the associated effective model; in this work, we investigate this relationship from the very basics: specifically, we propose a rigorous notion of scalar product among mappings, which implies a distance and a metric space of simplified representations. Making use of a Wang-Landau enhanced sampling algorithm, we exhaustively explore the space of mappings, quantifying their qualitative features in terms of their squared norm and relating them with thermodynamical properties of the underlying macromolecule. A one-to-one correspondence with an interacting lattice gas on a finite volume leads to the emergence of discontinuous phase transitions in mapping space that mark the boundaries between qualitatively different representations of the same molecule.

cond-mat.soft

Accelerating the identification of informative reduced representations of proteins with deep learning for graphs

The limits of molecular dynamics (MD) simulations of macromolecules are steadily pushed forward by the relentless developments of computer architectures and algorithms. This explosion in the number and extent (in size and time) of MD trajectories induces the need of automated and transferable methods to rationalise the raw data and make quantitative sense out of them. Recently, an algorithmic approach was developed by some of us to identify the subset of a protein's atoms, or mapping, that enables the most informative description of it. This method relies on the computation, for a given reduced representation, of the associated mapping entropy, that is, a measure of the information loss due to the simplification. Albeit relatively straightforward, this calculation can be time consuming. Here, we describe the implementation of a deep learning approach aimed at accelerating the calculation of the mapping entropy. The method relies on deep graph networks, which provide extreme flexibility in the input format. We show that deep graph networks are accurate and remarkably efficient, with a speedup factor as large as $10^5$ with respect to the algorithmic computation of the mapping entropy. Applications of this method, which entails a great potential in the study of biomolecules when used to reconstruct its mapping entropy landscape, reach much farther than this, being the scheme easily transferable to the computation of arbitrary functions of a molecule's structure.

physics.comp-ph

An information theory-based approach for optimal model reduction of biomolecules

In the theoretical modelling of a physical system a crucial step consists in the identification of those degrees of freedom that enable a synthetic, yet informative representation of it. While in some cases this selection can be carried out on the basis of intuition and experience, a straightforward discrimination of the important features from the negligible ones is difficult for many complex systems, most notably heteropolymers and large biomolecules. We here present a thermodynamics-based theoretical framework to gauge the effectiveness of a given simplified representation by measuring its information content. We employ this method to identify those reduced descriptions of proteins, in terms of a subset of their atoms, that retain the largest amount of information from the original model; we show that these highly informative representations share common features that are intrinsically related to the biological properties of the proteins under examination, thereby establishing a bridge between protein structure, energetics, and function.

cond-mat.stat-mech

A deep learning approach to the structural analysis of proteins

Deep Learning (DL) algorithms hold great promise for applications in the field of computational biophysics. In fact, the vast amount of available molecular structures, as well as their notable complexity, constitutes an ideal context in which DL-based approaches can be profitably employed. To express the full potential of these techniques, though, it is a prerequisite to express the information contained in the molecule's atomic positions and distances in a set of input quantities that the network can process. Many of the molecular descriptors devised insofar are effective and manageable for relatively small structures, but become complex and cumbersome for larger ones. Furthermore, most of them are defined locally, a feature that could represent a limit for those applications where global properties are of interest. Here, we build a deep learning architecture capable of predicting non-trivial and intrinsically global quantities, that is, the eigenvalues of a protein's lowest-energy fluctuation modes. This application represents a first, relatively simple test bed for the development of a neural network approach to the quantitative analysis of protein structures, and demonstrates unexpected use in the identification of mechanically relevant regions of the molecule.

cond-mat.soft