Searcharxiv⌕ Search

arXiv subjects

Vince Grolmusz

Publications and source records attributed to Vince Grolmusz.

At least 19 recordsLinked to original sources

On the Increased and Decreased Connectivity of the Demented Human Brain

With major advances in cerebral imaging techniques, a large amount of data is available for studying the aging and demented brain. In this contribution, we apply the OASIS-3 dataset to identify small areas of human gray matter with higher or lower structural connectivity in dementia. As was anticipated, we have found that finer structures of the hippocampus and the temporal lobe have decreased connectivity in dementia. More surprisingly, the precuneus, the cuneus, and some finer structures of the insula, the paracentral lobule, and the precentral and paracentral gyri showed higher connectivity in dementia than in healthy subjects.

q-bio.NC↗

New Graphs at the braingraph.org Website for Studying the Aging Brain Circuitry

Human braingraphs or connectomes are widely studied in the last decade to understand the structural and functional properties of our brain. In the last several years our research group has computed and deposited thousands of human braingraphs to the braingraph.org site, by applying public structural (diffusion) MRI data from young and healthy subjects. Here we describe a recent addition to the {\tt braingraph.org} site, which contains connectomes from healthy and demented subjects between 42 and 95 years of age, based on the public release of the OASIS-3 dataset. The diffusion MRI data was processed with the Connectome Mapper Toolkit v.3.1. We believe that the new addition to the braingraph.org site will become a useful resource for enlightening the aging circuitry of the human brain in healthy and diseased subjects, including those with Alzheimer's disease in several stages.

q-bio.NC↗

The Length and the Width of the Human Brain Circuit Connections are Strongly Correlated

The correlations of several fundamental properties of human brain connections are investigated in a consensus connectome, constructed from 1064 braingraphs, each on 1015 vertices, corresponding to 1015 anatomical brain areas. The properties examined include the edge length, the fiber number, or edge width, meaning the number of discovered axon bundles forming the edge and the occurrence number of the edge, meaning the number of individual braingraphs where the edge exists. By using our previously published robust braingraphs at \url{https://braingraph.org}, we have prepared a single consensus graph from the data and compared the statistical similarity of the edge occurrence numbers, edge lengths, and fiber counts of the edges. We have found a strong positive Spearman correlation between the edge occurrence numbers and the fiber count numbers, showing that statistically, the most frequent cerebral connections have the largest widths, i.e., the fiber number. We have found a negative Spearman correlation between the fiber lengths and fiber counts, showing that, typically, the shortest edges are the widest or strongest by their fiber counts. We have also found a negative Spearman correlation between the occurrence numbers and the edge lengths: it shows that typically, the long edges are infrequent, and the frequent edges are short.

q-bio.NC↗

Gluing GAP to RAS Mutants: A New Approach to an Old Problem in Cancer Drug Development

Mutated genes may lead to cancer development in numerous tissues. While more than 600 cancer-causing genes are known today, some of the most widespread mutations are connected to the RAS gene: RAS mutations are found in approximately 25% of all human tumors. Specifically, KRAS mutations are involved in the three most lethal cancers in U.S.: pancreatic ductal adenocarcinoma, colorectal adenocarcinoma, and lung adenocarcinoma. These cancers are among the most difficult to treat, and they are frequently excluded from chemotherapeutic attacks as hopeless cases. The mutated KRAS proteins have specific 3-dimensional conformations, which perturb functional interaction with the GAP protein on the GAP:RAS complex surface leading to a signaling cascade and uncontrolled cell growth. Here we describe a gluing docking method for finding small molecules that bind to both the GAP and the mutated KRAS molecules. These small molecules glue together the GAP and the mutated KRAS molecules and may serve as new cancer drugs for the most lethal, most difficult-to-treat carcinomas. As a proof of concept, we identify two new, drug-like small molecules with the new method: these compounds specifically inhibit the growth of PANC-1 cell line with KRAS mutation G12D in vitro and in vivo. Importantly, the two new compounds show significantly lower IC50 and higher specificity against the G12D KRAS mutant as compared to the recently described MRTX-1133 inhibitor against the G12D KRAS mutant.

q-bio.BM↗

A Note on the LogRank Conjecture in Communication Complexity

The LogRank conjecture of Lovász and Saks from 1988 is the most famous open problem in the communication complexity theory. The statement is as follows: Suppose that two players intend to compute a Boolean function $f(x,y)$ when $x$ is known for the first and $y$ for the second player, and they may send and receive messages encoded with bits, then they can compute $f(x,y)$ with exchanging $(\log \rank (M_f))^c $ bits, where $M_f$ is a Boolean matrix, determined by function $f$. The problem is widely open and very popular, and it has resisted numerous attacks in the last 35 years. The best upper bound is still exponential in the bound of the conjecture. Unfortunately, we cannot prove the conjecture, but we present a communication protocol with $(\log \rank (M_f))^c $ bits, which computes a -- somewhat -- related quantity to $f(x,y)$. The relation is characterized by a representation of low-degree, multi-linear polynomials modulo composite numbers. This result of ours may help to settle this long-time open conjecture.

cs.CC↗

Navigating Homogeneous Paths through Amyloidogenic and Non-Amyloidogenic Hexapeptides

Hexapeptides are increasingly applied as model systems for studying the amyloidogenecity properties of oligo- and polypeptides. It is possible to construct 64 million different hexapeptides from the twenty proteinogenic amino acid residues. Today's experimental amyloid databases contain only a fraction of these annotated hexapeptides. For labeling all the possible hexapeptides as "amyloidogenic" or "non-amyloidogenic" there exist several computational predictors with good accuracies. It may be of interest to define and study a simple graph structure on the 64 million hexapeptides as nodes when two hexapeptides are connected by an edge if they differ by only a single residue. For example, in this graph, HIKKLM is connected to AIKKLM, or HIKKNM, or HIKKLC, but it is not connected with an edge to VVKKLM or HIKNPM. In the present contribution, we consider our previously published artificial intelligence-based tool, the Budapest Amyloid Predictor (BAP for short), and demonstrate a spectacular property of this predictor in the graph defined above. We show that for any two hexapeptides predicted to be "amyloidogenic" by the BAP predictor, there exists an easily constructible path of length at most 6 that passes through neighboring hexapeptides all predicted to be "amyloidogenic" by BAP. For example, the predicted amyloidogenic ILVWIW and FWLCYL hexapeptides can be connected through the length-6 path ILVWIW-IWVWIW-IWVCIW-IWVCIL-FWVCIL-FWLCIL-FWLCYL in such a way that the neighbors differ in exactly one residue, and all hexapeptides on the path are predicted to be amyloidogenic by BAP. The symmetric statement also holds for non-amyloidogenic hexapeptides. It is noted that the mentioned property of the Budapest Amyloid Predictor \url{https://pitgroup.org/bap} is not proprietary; it is also true for any linear Support Vector Machine (SVM)-based predictors.

q-bio.BM↗

Novel enzymes for biodegradation of polycyclic aromatic hydrocarbons: metagenomics-linked identification followed by functional analysis

Polycyclic aromatic hydrocarbons (PAHs) are highly toxic, carcinogenic substances. On soils contaminated with PAHs, crop cultivation, animal husbandry and even the survival of microflora in the soil are greatly perturbed, depending on the degree of contamination. Most microorganisms cannot tolerate PAH-contaminated soils, however, some microbial strains can adapt to these harsh conditions and survive on contaminated soils. Analysis of the metagenomes of contaminated environmental samples may lead to discovery of PAH-degrading enzymes suitable for green biotechnology methodologies ranging from biocatalysis to pollution control. In the present study, our goal was to apply a metagenomic data search to identify efficient novel enzymes in remediation of PAH-contaminated soils. The metagenomic hits were further analyzed using a set of bioinformatics tools to select protein sequences predicted to encode well-folded soluble enzymes. Three novel enzymes (two dioxygenases and one peroxidase) were cloned and used in soil remediation microcosms experiments. The novel enzymes were found to be efficient for degradation of naphthalene and phenanthrene. Adding the inorganic oxidant CaO2 further increased the degrading potential of the novel enzymes for anthracene and pyrene. We conclude that metagenome mining paired with bioinformatic predictions, structural modelling and functional assays constitutes a powerful approach towards novel enzymes for soil remediation.

q-bio.BM↗

Robust Circuitry-Based Scores of Structural Importance of Human Brain Areas

We consider the 1015-vertex human consensus connectome computed from the diffusion MRI data of 1064 subjects. We define seven different orders on these 1015 graph vertices, where the orders depend on parameters derived from the brain circuitry, that is, from the properties of the edges (or connections) incident to the vertices ordered. We order the vertices according to their degree, the sum, the maximum, and the average of the fiber counts on the incident edges, and the sum, the maximum and the average length of the fibers in the incident edges. We analyze the similarities of these seven orders by the Spearman correlation coefficient and by their inversion numbers and have found that all of these seven orders have great similarities. In other words, if we interpret the orders as scoring of the importance of the vertices in the consensus connectome, then the scores of the vertices will be similar in all seven orderings. That is, important vertices of the human connectome typically have many neighbors, connected with long and thick axonal fibers (where thickness is measured by fiber numbers), and their incident edges have high maximum and average values of length and fiber-number parameters, too. Therefore, these parameters may yield robust ways of deciding which vertices are more important in the anatomy of our brain circuitry than the others.

q-bio.NC↗

Opening Amyloid-Windows to the Secondary Structure of Proteins: The Amyloidogenecity Increases Tenfold Inside Beta-Sheets

Methods from artificial intelligence (AI), in general, and machine learning, in particular, have kept conquering new territories in numerous areas of science. Most of the applications of these techniques are restricted to the classification of large data sets, but new scientific knowledge can seldom be inferred from these tools. Here we show that an AI-based amyloidogenecity predictor can strongly differentiate the border- and the internal hexamers of $β$-pleated sheets when screening all the Protein Data Bank-deposited homology-filtered protein structures. Our main result shows that more than 30\% of internal hexamers of $β$ sheets are predicted to be amyloidogenic, while just outside the border regions, only 3\% are predicted as such. This result may elucidate a general protection mechanism of proteins against turning into amyloids: if the borders of $β$-sheets were amyloidogenic, then the whole $β$ sheet could turn more easily into an insoluble amyloid-structure, characterized by periodically repeated parallel $β$-sheets. We also present that no analogous phenomenon exists on the borders of $α$-helices or randomly chosen subsequences of the studied protein structures.

q-bio.BM↗

Succinct Amyloid and Non-Amyloid Patterns in Hexapeptides

Hexapeptides are widely applied as a model system for studying amyloid-forming properties of polypeptides, including proteins. Recently, large experimental databases have become publicly available with amyloidogenic labels. Using these datasets for training and testing purposes, one may build artificial intelligence (AI)-based classifiers for predicting the amyloid state of peptides. In our previous work (Biomolecules, 11(4) 500, (2021)) we described the Support Vector Machine (SVM)-based Budapest Amyloid Predictor (\url{https://pitgroup.org/bap}). Here we apply the Budapest Amyloid Predictor for discovering numerous amyloidogenic and non-amyloidogenic hexapeptide patterns with accuracy between 80\% and 84\%, as surprising and succinct novel rules for further understanding the amyloid state of peptides. For example, we have shown that for any independently mutated residue (position marked by ``x''), the patterns CxFLWx, FxFLFx, or xxIVIV are predicted to be amyloidogenic, while those of PxDxxx, xxKxEx, and xxPQxx non-amyloidogenic at all. We note that each amyloidogenic pattern with two x's (e.g.,CxFLWx) describes succinctly $20^2=400$ hexapeptides, while the non-amyloidogenic patterns comprising four point mutations (e.g.,PxDxxx) gives $20^4=160,000$ hexapeptides in total. To our knowledge, no similar applications of artificial intelligence tools or succinct amyloid patterns were described before the present work.

q-bio.BM↗

Discovering Sex and Age Implicator Edges in the Human Connectome

Determining important vertices in large graphs (e.g., Google's PageRank in the case of the graph of the World Wide Web) facilitated the construction of excellent web search engines, returning the most important hits corresponding to the submitted user queries. Interestingly, finding important edges -- instead of vertices -- in large graphs has received much less attention until now. Here we examine the human structural braingraph (or connectome), identified by diffusion magnetic resonance imaging (dMRI) methods, with edges connecting cortical and subcortical gray matter areas and weighted by fiber strengths, measured by the number of the discovered fiber tracts along the edge. We identify several "single" important edges in these braingraphs, whose high or low weights imply the sex or the age of the subject observed. We call these edges implicator edges since solely from their weight, one can infer the sex of the subject with more than 67 \% accuracy or their age group with more than 62\% accuracy. We argue that these brain connections are the most important ones characterizing the sex or the age of the subjects. Surprisingly, the edges implying the male sex are mostly located in the anterior parts of the brain, while those implying the female sex are mostly in the posterior regions. Additionally, most of the inter-hemispheric implicator edges are male ones, while the intra-hemispheric ones are predominantly female edges. Our pioneering method for finding the sex- or age implicator edges can also be applied for characterizing other biological and medical properties, including neurodegenerative- and psychiatric diseases besides the sex or the age of the subject, if large and high-quality neuroimaging datasets become available.

q-bio.NC↗

The Budapest Amyloid Predictor and its Applications

The amyloid state of proteins is widely studied with relevancy in neurology, biochemistry, and biotechnology. In contrast with amorphous aggregation, the amyloid state has a well-defined structure, consisting of parallel and anti-parallel $β$-sheets in a periodically repeated formation. The understanding of the amyloid state is growing with the development of novel molecular imaging tools, like cryogenic electron microscopy. Sequence-based amyloid predictors were developed by using mostly artificial neural networks (ANNs) as the underlying computational techniques. From a good neural network-based predictor, it is a very difficult task to identify those attributes of the input amino acid sequence, which implied the decision of the network. Here we present a Support Vector Machine (SVM)-based predictor for hexapeptides with correctness higher than 84\%, i.e., it is at least as good as the published ANN-based tools. Unlike the artificial neural networks, the decision of the SVMs are much easier to analyze, and from a good predictor, we can infer rich biochemical knowledge. Availability and Implementation: The Budapest Amyloid Predictor webserver is freely available at https://pitgroup.org/bap.

q-bio.BM↗

Introducing and Applying Newtonian Blurring: An Augmented Dataset of 126,000 Human Connectomes at braingraph.org

Gaussian blurring is a well-established method for image data augmentation: it may generate a large set of images from a small set of pictures for training and testing purposes for Artificial Intelligence (AI) applications. When we apply AI for non-imagelike biological data, hardly any related method exists. Here we introduce the "Newtonian blurring" in human braingraph (or connectome) augmentation: Started from a dataset of 1053 subjects, we first repeat a probabilistic weighted braingraph construction algorithm 10 times for describing the connections of distinct cerebral areas, then take 7 repetitions in every possible way, delete the lower and upper extremes, and average the remaining 7-2=5 edge-weights for the data of each subject. This way we augment the 1053 graph-set to 120 x 1053 = 126,360 graphs. In augmentation techniques, it is an important requirement that no artificial additions should be introduced into the dataset. Gaussian blurring and also this Newtonian blurring satisfy this goal. The resulting dataset of 126,360 graphs, each in 5 resolutions (i.e., 631,800 graphs in total), is freely available at the site https://braingraph.org/cms/download-pit-group-connectomes/. Augmenting with Newtonian blurring may also be applicable in other non-image related fields, where probabilistic processing and data averaging are implemented.

q-bio.NC↗

The braingraph.org Database with more than 1000 Robust Human Structural Connectomes in Five Resolutions

The human brain is the most complex object of study we encounter today. Mapping the neuronal-level connections between the more than 80 billion neurons in the brain is a hopeless task for science. By the recent advancement of magnetic resonance imaging (MRI), we are able to map the macroscopic connections between about 1000 brain areas. The MRI data acquisition and the subsequent algorithmic workflow contain several complex steps, where errors can occur. In the present contribution, we describe and publish 1064 human connectomes, computed from the public release of the Human Connectome Project. Each connectome is available in 5 resolutions, with 83, 129, 234, 463, and 1015 anatomically labeled nodes. For error correction, we follow an averaging and extreme value deleting strategy for each edge and for each connectome. The resulting 5320 braingraphs can be downloaded from the \url{https://braingraph.org} site. This dataset makes possible the access to these graphs for scientists unfamiliar with neuroimaging- and connectome-related tools: mathematicians, physicists, and engineers can use their expertize and ideas in the analysis of the connections of the human brain. Brain scientists also have a robust and large, multi-resolution set for connectomical studies.

q-bio.NC↗

The Graph of Our Mind

Graph theory in the last two decades penetrated sociology, molecular biology, genetics, chemistry, computer engineering, and numerous other fields of science. One of the more recent areas of its applications is the study of the connections of the human brain. By the development of diffusion magnetic resonance imaging (diffusion MRI), it is possible today to map the connections between the 1-1.5 cm$^2$ regions of the gray matter of the human brain. These connections can be viewed as a graph: the vertices are the anatomically identified regions of the gray matter, and two vertices are connected by an edge if the diffusion MRI-based workflow finds neuronal fiber tracts between these areas. This way we can compute 1015-vertex graphs with tens of thousands of edges. In a previous work, we have analyzed the male and female braingraphs graph-theoretically, and we have found statistically significant differences in numerous parameters between the sexes: the female braingraphs are better expanders, have more edges, larger bipartition widths, and larger vertex cover than the braingraphs of the male subjects. Our previous study has applied the data of 96 subjects; here we present a much larger study of 426 subjects. Our data source is an NIH-founded project, the "Human Connectome Project (HCP)" public data release. As a service to the community, we have also made all of the braingraphs computed by us from the HCP data publicly available at the \url{http://braingraph.org} for independent validation and further investigations.

q-bio.NC↗

On the Border of the Amyloidogenic Sequences: Prefix Analysis of the Parallel Beta Sheets in the PDB\_Amyloid Collection

The Protein Data Bank (PDB) today contains more than 153,000 entries with the 3-dimensional structures of biological macromolecules. Using the rich resources of this repository, it is possible identifying subsets with specific, interesting properties for different applications. Our research group prepared an automatically updated list of amyloid- and probably amyloidogenic molecules, the PDB\_Amyloid collection, which is freely available at the address \url{http://pitgroup.org/amyloid}. This resource applies exclusively the geometric properties of the steric structures for identifying amyloids. In the present contribution, we analyze the starting (i.e., prefix) subsequences of the characteristic, parallel beta-sheets of the structures in the PDB\_Amyloid collection, and identify further appearances of these length-5 prefix subsequences in the whole PDB data set. We have identified this way numerous proteins, whose normal or irregular functions involve amyloid formation, structural misfolding, or anti-coagulant properties, simply by containing these prefixes: including the T-cell receptor (TCR), bound with the major histocompatibility complexes MHC-1 and MHC-2; the p53 tumor suppressor protein; a mycobacterial RNA polymerase transcription initialization complex; the human bridging integrator protein BIN-1; and the tick anti-coagulant peptide TAP.

q-bio.BM↗

Identifying Super-Feminine, Super-Masculine and Sex-Defining Connections in the Human Braingraph

For more than a decade now, we can discover and study thousands of cerebral connections with the application of diffusion magnetic resonance imaging (dMRI) techniques and the accompanying algorithmic workflow. While numerous connectomical results were published enlightening the relation between the braingraph and certain biological, medical, and psychological properties, it is still a great challenge to identify a small number of brain connections, closely related to those conditions. In the present contribution, by applying the 1200 Subjects Release of the Human Connectome Project (HCP), we identify just 102 connections out of the total number of 1950 connections in the 83-vertex graphs of 1065 subjects, which -- by a simple linear test -- precisely, without any error determine the sex of the subject. Very surprisingly, we were able to identify two graph edges out of these 102, if, whose weights, measured in fiber numbers, are all high, then the connectome always belongs to a female subject, independently of the other edges. Similarly, we have identified 3 edges from these 102, whose weights, if two of them are high and one is low, imply that the graph belongs to a male subject -- again, independently of the other edges. We call the former 2 edges superfeminine and the first two of the 3 edges supermasculine edges of the human connectome. Even more interestingly, one of the edges, connecting the right Pars Triangularis and the right Superior Parietal areas, is one of the 2 superfeminine edges, and it is also the third edge, accompanying the two supermasculine connections, if its weight is low; therefore it is also a "switching" connection.

q-bio.NC↗

Good Neighbors, Bad Neighbors: The Frequent Network Neighborhood Mapping of the Hippocampus Enlightens Several Structural Factors of the Human Intelligence on a 414-Subject Cohort

The human connectome has become the very frequent subject of study of brain-scientists, psychologists, and imaging experts in the last decade. With diffusion magnetic resonance imaging techniques, unified with advanced data processing algorithms, today we are able to compute braingraphs with several hundred, anatomically identified nodes and thousands of edges, corresponding to the anatomical connections of the brain. The analysis of these graphs without refined mathematical tools is hopeless. These tools need to address the high error rate of the MRI processing workflow, and need to find structural causes or at least correlations of psychological properties and cerebral connections. Until now, structural connectomics was only rarely able identifying such causes or correlations. In the present work, we study the frequent neighbor sets of the most deeply investigated brain area, the hippocampus. By applying the Frequent Network Neighborhood mapping method, we identified frequent neighbor-sets of the hippocampus, which may influence numerous psychological parameters, including intelligence-related ones. We have found neighbor sets, which have significantly higher frequency in subjects with high-scored Penn Matrix tests, and with low-scored Penn Word Memory tests. Our study utilizes the braingraphs, computed from the imaging data of the Human Connectome Project's 414 subjects, each with 463 anatomically identified nodes.

q-bio.NC↗