SearcharxivSearch

arXiv subjects

Yi-Kuo Yu

Publications and source records attributed to Yi-Kuo Yu.

At least 19 recordsLinked to original sources

Mass spectrometry based protein identification with accurate statistical significance assignment

Motivation: Assigning statistical significance accurately has become increasingly important as meta data of many types, often assembled in hierarchies, are constructed and combined for further biological analyses. Statistical inaccuracy of meta data at any level may propagate to downstream analyses, undermining the validity of scientific conclusions thus drawn. From the perspective of mass spectrometry based proteomics, even though accurate statistics for peptide identification can now be achieved, accurate protein level statistics remain challenging. Results: We have constructed a protein ID method that combines peptide evidences of a candidate protein based on a rigorous formula derived earlier; in this formula the database $P$-value of every peptide is weighted, prior to the final combination, according to the number of proteins it maps to. We have also shown that this protein ID method provides accurate protein level $E$-value, eliminating the need of using empirical post-processing methods for type-I error control. Using a known protein mixture, we find that this protein ID method, when combined with the Soric formula, yields accurate values for the proportion of false discoveries. In terms of retrieval efficacy, the results from our method are comparable with other methods tested. Availability: The source code, implemented in C++ on a linux system, is available for download at ftp://ftp.ncbi.nlm.nih.gov/pub/qmbp/qmbp_ms/RAId/RAId_Linux_64Bit

q-bio.QM

Transition from one-dimensional antiferromagnetism to three-dimensional antiferromagnetic order in single-crystalline CuSb$_{2}$O$_{6}$

Measurements of magnetic susceptibility, heat capacity and thermal expansion are reported for single crystalline CuSb$_{2}$O$_{6}$ in the temperature range $5<T<350$ K. The magnetic susceptibility exhibits a broad peak centered near 60 K that is typical of one-dimensional antiferromagnetic compounds. Long-range antiferromagnetic order at $T_N$ = 8.7 K is accompanied by an energy gap ($\Delta$ = 17.48(6) K). This transition represents a crossover from one- to three-dimensional antiferromagnetic behavior. Both heat capacity and the thermal expansion coefficients exhibit distinct jumps at $T_N$, which are similar to those observed at the normal-superconducting phase transition in a superconductor. This behavior is quite unusual, and is presumably associated with a Spin-Peierls transition occurring as a result of three-dimensional phonons coupling with {\it Jordan-Wigner-transformed} Fermions.

cond-mat.str-el

Information flow in interaction networks II: channels, path lengths and potentials

In our previous publication, a framework for information flow in interaction networks based on random walks with damping was formulated with two fundamental modes: emitting and absorbing. While many other network analysis methods based on random walks or equivalent notions have been developed before and after our earlier work, one can show that they can all be mapped to one of the two modes. In addition to these two fundamental modes, a major strength of our earlier formalism was its accommodation of context-specific directed information flow that yielded plausible and meaningful biological interpretation of protein functions and pathways. However, the directed flow from origins to destinations was induced via a potential function that was heuristic. Here, with a theoretically sound approach called the channel mode, we extend our earlier work for directed information flow. This is achieved by constructing a potential function facilitating a purely probabilistic interpretation of the channel mode. For each network node, the channel mode combines the solutions of emitting and absorbing modes in the same context, producing what we call a channel tensor. The entries of the channel tensor at each node can be interpreted as the amount of flow passing through that node from an origin to a destination. Similarly to our earlier model, the channel mode encompasses damping as a free parameter that controls the locality of information flow. Through examples involving the yeast pheromone response pathway, we illustrate the versatility and stability of our new framework.

q-bio.MN

CytoITMprobe: a network information flow plugin for Cytoscape

To provide the Cytoscape users the possibility of integrating ITM Probe into their workflows, we developed CytoITMprobe, a new Cytoscape plugin. CytoITMprobe maintains all the desirable features of ITM Probe and adds additional flexibility not achievable through its web service version. It provides access to ITM Probe either through a web server or locally. The input, consisting of a Cytoscape network, together with the desired origins and/or destinations of information and a dissipation coefficient, is specified through a query form. The results are shown as a subnetwork of significant nodes and several summary tables. Users can control the composition and appearance of the subnetwork and interchange their ITM Probe results with other software tools through tab-delimited files. The main strength of CytoITMprobe is its flexibility. It allows the user to specify as input any Cytoscape network, rather than being restricted to the pre-compiled protein-protein interaction networks available through the ITM Probe web service. Users may supply their own edge weights and directionalities. Consequently, as opposed to ITM Probe web service, CytoITMprobe can be applied to many other domains of network-based research beyond protein-networks. It also enables seamless integration of ITM Probe results with other Cytoscape plugins having complementary functionality for data analysis.

q-bio.QM

Information Flow in Interaction Networks

Interaction networks, consisting of agents linked by their interactions, are ubiquitous across many disciplines of modern science. Many methods of analysis of interaction networks have been proposed, mainly concentrating on node degree distribution or aiming to discover clusters of agents that are very strongly connected between themselves. These methods are principally based on graph-theory or machine learning. We present a mathematically simple formalism for modelling context-specific information propagation in interaction networks based on random walks. The context is provided by selection of sources and destinations of information and by use of potential functions that direct the flow towards the destinations. We also use the concept of dissipation to model the aging of information as it diffuses from its source. Using examples from yeast protein-protein interaction networks and some of the histone acetyltransferases involved in control of transcription, we demonstrate the utility of the concepts and the mathematical constructs introduced in this paper.

q-bio.MN

CytoSaddleSum: a functional enrichment analysis plugin for Cytoscape based on sum-of-weights scores

Summary: CytoSaddleSum provides Cytoscape users with access to the functionality of SaddleSum, a functional enrichment tool based on sum-of-weight scores. It operates by querying SaddleSum locally (using the standalone version) or remotely (through an HTTP request to a web server). The functional enrichment results are shown as a term relationship network, where nodes represent terms and edges show term relationships. Furthermore, query results are written as Cytoscape attributes allowing easy saving, retrieval and integration into network-based data analysis workflows. Availability: www.ncbi.nlm.nih.gov/CBBresearch/Yu/downloads The source code is placed in Public Domain.

q-bio.QM

ppiTrim: Constructing non-redundant and up-to-date interactomes

Robust advances in interactome analysis demand comprehensive, non-redundant and consistently annotated datasets. By non-redundant, we mean that the accounting of evidence for every interaction should be faithful: each independent experimental support is counted exactly once, no more, no less. While many interactions are shared among public repositories, none of them contains the complete known interactome for any model organism. In addition, the annotations of the same experimental result by different repositories often disagree. This brings up the issue of which annotation to keep while consolidating evidences that are the same. The iRefIndex database, including interactions from most popular repositories with a standardized protein nomenclature, represents a significant advance in all aspects, especially in comprehensiveness. However, iRefIndex aims to maintain all information/annotation from original sources and requires users to perform additional processing to fully achieve the aforementioned goals. To address issues with iRefIndex and to achieve our goals, we have developed ppiTrim, a script that processes iRefIndex to produce non-redundant, consistently annotated datasets of physical interactions. Our script proceeds in three stages: mapping all interactants to gene identifiers and removing all undesired raw interactions, deflating potentially expanded complexes, and reconciling for each interaction the annotation labels among different source databases. As an illustration, we have processed the three largest organismal datasets: yeast, human and fruitfly. While ppiTrim can resolve most apparent conflicts between different labelings, we also discovered some unresolvable disagreements mostly resulting from different annotation policies among repositories. URL: http://www.ncbi.nlm.nih.gov/CBBresearch/Yu/downloads/ppiTrim.html

q-bio.MN

Combining independent, arbitrarily weighted P-values: a new solution to an old problem using a novel expansion with controllable accuracy

Good's formula and Fisher's method are frequently used for combining independent P-values. Interestingly, the equivalent of Good's formula already emerged in 1910 and mathematical expressions relevant to even more general situations have been repeatedly derived, albeit in different context. We provide here a novel derivation and show how the analytic formula obtained reduces to the two aforementioned ones as special cases. The main novelty of this paper, however, is the explicit treatment of nearly degenerate weights, which are known to cause numerical instabilities. We derive a controlled expansion, in powers of differences in inverse weights, that provides both accurate statistics and stable numerics.

math.ST

RAId_aPS: MS/MS analysis with multiple scoring functions and spectrum-specific statistics

Statistically meaningful comparison/combination of peptide identification results from various search methods is impeded by the lack of a universal statistical standard. Providing an E-value calibration protocol, we demonstrated earlier the feasibility of translating either the score or heuristic E-value reported by any method into the textbook-defined E-value, which may serve as the universal statistical standard. This protocol, although robust, may lose spectrum-specific statistics and might require a new calibration when changes in experimental setup occur. To mitigate these issues, we developed a new MS/MS search tool, RAId_aPS, that is able to provide spectrum-specific E-values for additive scoring functions. Given a selection of scoring functions out of RAId score, K-score, Hyperscore and XCorr, RAId_aPS generates the corresponding score histograms of all possible peptides using dynamic programming. Using these score histograms to assign E-values enables a calibration-free protocol for accurate significance assignment for each scoring function. RAId_aPS features four different modes: (i) compute the total number of possible peptides for a given molecular mass range, (ii) generate the score histogram given a MS/MS spectrum and a scoring function, (iii) reassign E-values for a list of candidate peptides given a MS/MS spectrum and the scoring functions chosen, and (iv) perform database searches using selected scoring functions. In modes (iii) and (iv), RAId_aPS is also capable of combining results from different scoring functions using spectrum-specific statistics. The web link is http://www.ncbi.nlm.nih.gov/CBBresearch/Yu/raid_aps/index.html. Relevant binaries for Linux, Windows, and Mac OS X are available from the same page.

q-bio.QM

Derivation of the Density Functional via Effective Action

A rigorous derivation of the density functional in the Hohenberg-Kohn theory is presented. With no assumption regarding the magnitude of the electric coupling constant $e^2$ (or correlation), this work provides a firm basis for first-principles calculations. Using the auxiliary field method, in which $e^2$ need not be small, we show that the bosonic loop expansion of the exchange-correlation functional can be reorganized so as to be expressed entirely in terms of the Kohn-Sham single-particle orbitals and energies. The excitations of the many-particle system can be obtained within the same formalism. We also explicitly demonstrate at zero-temperature the single-particle limit, the weak-coupling limit of the energy functional, and its application to homogeneous electron gas.

cond-mat.other

Robust and accurate data enrichment statistics via distribution function of sum of weights

Term enrichment analysis facilitates biological interpretation by assigning to experimentally/computationally obtained data annotation associated with terms from controlled vocabularies. This process usually involves obtaining statistical significance for each vocabulary term and using the most significant terms to describe a given set of biological entities, often associated with weights. Many existing enrichment methods require selections of (arbitrary number of) the most significant entities and/or do not account for weights of entities. Others either mandate extensive simulations to obtain statistics or assume normal weight distribution. In addition, most methods have difficulty assigning correct statistical significance to terms with few entities. Implementing the well-known Lugananni-Rice formula, we have developed a novel approach, called SaddleSum, that is free from all the aforementioned constraints and evaluated it against several existing methods. With entity weights properly taken into account, SaddleSum is internally consistent and stable with respect to the choice of number of most significant entities selected. Making few assumptions on the input data, the proposed method is universal and can thus be applied to areas beyond analysis of microarrays. Employing asymptotic approximation, SaddleSum provides a term-size dependent score distribution function that gives rise to accurate statistical significance even for terms with few entities. As a consequence, SaddleSum enables researchers to place confidence in its significance assignments to small terms that are often biologically most specific.

q-bio.QM

A simple electrostatic model applicable to biomolecular recognition

An exact, analytic solution for a simple electrostatic model applicable to biomolecular recognition is presented. In the model, a layer of high dielectric constant material (representative of the solvent, water) whose thickness may vary separates two regions of low dielectric constant material (representative of proteins, DNA, RNA, or similar materials), in each of which is embedded a point charge. For identical charges, the presence of the screening layer always lowers the energy compared to the case of point charges in an infinite medium of low dielectric constant. Somewhat surprisingly, the presence of a sufficiently thick screening layer also lowers the energy compared to the case of point charges in an infinite medium of high dielectric constant. For charges of opposite sign, the screening layer always lowers the energy compared to the case of point charges in an infinite medium of either high or low dielectric constant. The behavior of the energy leads to a substantially increased repulsive force between charges of the same sign. The repulsive force between charges of opposite signs is weaker than in an infinite medium of low dielectric constant material but stronger than in an infinite medium of high dielectric constant material. The presence of this behavior, which we name asymmetric screening, in the simple system presented here confirms the generality of the behavior that was established in a more complicated system of an arbitrary number of charged dielectric spheres in an infinite solvent.

physics.class-ph

The Density Functional via Effective Action

A rigorous derivation of the density functional via the effective action in the Hohenberg-Kohn theory is outlined. Using the auxiliary field method, in which the electric coupling constant $e^2$ need not be small, we show that the loop expansion of the exchange-correlation functional can be reorganized so as to be expressed entirely in terms of the Kohn-Sham single-particle orbitals and energies.

cond-mat.other

ITM Probe: analyzing information flow in protein networks

Summary: Founded upon diffusion with damping, ITM Probe is an application for modeling information flow in protein interaction networks without prior restriction to the sub-network of interest. Given a context consisting of desired origins and destinations of information, ITM Probe returns the set of most relevant proteins with weights and a graphical representation of the corresponding sub-network. With a click, the user may send the resulting protein list for enrichment analysis to facilitate hypothesis formation or confirmation. Availability: ITM Probe web service and documentation can be found at www.ncbi.nlm.nih.gov/CBBresearch/qmbp/mn/itm_probe

q-bio.MN

Rigorous treatment of electrostatics for spatially varying dielectrics based on energy minimization

A novel energy minimization formulation of electrostatics that allows computation of the electrostatic energy and forces to any desired accuracy in a system with arbitrary dielectric properties is presented. An integral equation for the scalar charge density is derived from an energy functional of the polarization vector field. This energy functional represents the true energy of the system even in non-equilibrium states. Arbitrary accuracy is achieved by solving the integral equation for the charge density via a series expansion in terms of the equation's kernel, which depends only on the geometry of the dielectrics. The streamlined formalism operates with volume charge distributions only, not resorting to introducing surface charges by hand. Therefore, it can be applied to any spatial variation of the dielectric susceptibility, which is of particular importance in applications to biomolecular systems. The simplicity of application of the formalism to real problems is shown with analytical and numerical examples.

physics.class-ph

Statistical Characterization of a 1D Random Potential Problem - with applications in score statistics of MS-based peptide sequencing

We provide a complete thermodynamic solution of a 1D hopping model in the presence of a random potential by obtaining the density of states. Since the partition function is related to the density of states by a Laplace transform, the density of states determines completely the thermodynamic behavior of the system. We have also shown that the transfer matrix technique, or the so-called dynamic programming, used to obtain the density of states in the 1D hopping model may be generalized to tackle a long-standing problem in statistical significance assessment for one of the most important proteomic tasks - peptide sequencing using tandem mass spectrometry data.

q-bio.QM

RAId DbS: A Mass-Spectrometry Based Peptide Identification Web Server with Knowledge Integration

Summary: In anticipation of the individualized proteomics era and the need to integrate knowledge from disease studies, we have augmented our peptide identification software RAId DbS to take into account annotated single amino acid polymorphisms, post-translational modifications, and their documented disease associations while analyzing a tandem mass spectrum. To facilitate new discoveries, RAId DbS allows users to conduct searches permitting novel polymorphisms. Availability: The webserver link is http://www.ncbi.nlm.nih.gov/ /CBBResearch/qmbp/raid dbs/index.html. The relevant databases and binaries of RAId DbS for Linux, Windows, and Mac OS X are available from the same web page. Contact: yyu@ncbi.nlm.nih.gov

q-bio.QM

Heat Conduction Process on Community Networks as a Recommendation Model

Using heat conduction mechanism on a social network we develop a systematic method to predict missing values as recommendations. This method can treat very large matrices that are typical of internet communities. In particular, with an innovative, exact formulation that accommodates arbitrary boundary condition, our method is easy to use in real applications. The performance is assessed by comparing with traditional recommendation methods using real data.

physics.soc-ph