SearcharxivSearch

arXiv subjects

Jeffrey Skolnick

Publications and source records attributed to Jeffrey Skolnick.

3 recordsLinked to original sources

HiGen: Hierarchy-Aware Sequence Generation for Hierarchical Text Classification

Hierarchical text classification (HTC) is a complex subtask under multi-label text classification, characterized by a hierarchical label taxonomy and data imbalance. The best-performing models aim to learn a static representation by combining document and hierarchical label information. However, the relevance of document sections can vary based on the hierarchy level, necessitating a dynamic document representation. To address this, we propose HiGen, a text-generation-based framework utilizing language models to encode dynamic text representations. We introduce a level-guided loss function to capture the relationship between text and label name semantics. Our approach incorporates a task-specific pretraining strategy, adapting the language model to in-domain knowledge and significantly enhancing performance for classes with limited examples. Furthermore, we present a new and valuable dataset called ENZYME, designed for HTC, which comprises articles from PubMed with the goal of predicting Enzyme Commission (EC) numbers. Through extensive experiments on the ENZYME dataset and the widely recognized WOS and NYT datasets, our methodology demonstrates superior performance, surpassing existing approaches while efficiently handling data and mitigating class imbalance. The data and code will be released publicly.

cs.CL

All-atom ab initio folding of a diverse set of proteins

Natural proteins fold to a unique, thermodynamically dominant state. Modeling of the folding process and prediction of the native fold of proteins are two major unsolved problems in biophysics. Here, we show successful all-atom ab initio folding of a representative diverse set of proteins, using a minimalist transferable energy model that consists of two-body atom-atom interactions, hydrogen-bonding, and a local sequence energy term that models sequence-specific chain stiffness. Starting from a random coil, the native-like structure was observed during replica exchange Monte Carlo (REMC) simulation for most proteins regardless of their structural classes; the lowest energy structure was close to native- in the range of 2-6 A root-mean-square deviation (RMSD). Our results demonstrate that the successful all-atom folding of a protein chain to its native state is governed by only a few crucial energetic terms.

q-bio.BM

Numerical estimation of entropy loss on dimerization: improved prediction of the quaternary structure of the GCN4 leucine zipper

A lattice based model of a protein is used to study the dimerization equilibrium of the GCN4 leucine zipper. Replica exchange Monte Carlo is used to determine the free energy of both the monomeric and dimeric forms as a function of temperature. The method of coincidences is then introduced to explicitly calculate the entropy loss associated with dimerization, and from it the free energy difference between monomer and dimer, as well as the corresponding equilibrium reaction constant. We find that the entropy loss of dimerization is a strong function of energy (or temperature), and that it is much larger than previously estimated, especially for high energy states. The results confirm that it is possible to study the dimerization equilibrium of GCN4 at physiological concentrations within the reduced representation of the protein employed.

physics.bio-ph