SearcharxivSearch

arXiv subjects

Arjun Krishnan

Publications and source records attributed to Arjun Krishnan.

15 recordsLinked to original sources

MetaHQ: Harmonized, high-quality metadata annotations of public omics samples and studies

Public omics databases like the Gene Expression Omnibus and the Sequence Read Archive offer substantial opportunities for data reuse to address novel biomedical questions. However, it is still difficult to find samples and studies of interest since they are described by free-text metadata and lack standardized annotations. To address this issue, multiple research groups have undertaken curation efforts to add standardized annotations to large collections of these data, but these annotations are fragmented across online resources and are stored in different formats subject to varying standardization criteria, hindering the integration of annotations across sources. We developed MetaHQ to harmonize and distribute standardized metadata for public omics samples. MetaHQ comprises a database with nearly 200,000 annotations from 13 sources and a user-friendly command-line interface (CLI) to query the database and retrieve annotations. The MetaHQ CLI is deployed as a Python Package on PyPI at https://pypi.org/project/metahq-cli that accesses the MetaHQ database available at https://doi.org/10.5281/zenodo.18462463. Project source code and documentation are available at https://github.com/krishnanlab/meta-hq.

q-bio.GN

Improving Biomedical Knowledge Graph Quality: A Community Approach

Biomedical knowledge graphs (KGs) are widely used across research and translational settings, yet their design decisions and implementation are often opaque. Unlike ontologies that more frequently adhere to established creation principles, biomedical KGs lack consistent practices for construction, documentation, and dissemination. To address this gap, we introduce a set of evaluation criteria grounded in widely accepted data standards and principles from related fields. We apply these criteria to 16 biomedical KGs, revealing that even those that appear to align with best practices often obscure essential information required for external reuse. Moreover, biomedical KGs, despite pursuing similar goals and ingesting the same sources in some cases, display substantial variation in models, source integration, and terminology for node types. Reaping the potential benefits of knowledge graphs for biomedical research while reducing wasted effort requires community-wide adoption of shared criteria and maturation of standards such as BioLink and KGX. Such improvements in transparency and standardization are essential for creating long-term reusability, improving comparability across resources, and enhancing the overall utility of KGs within biomedicine.

q-bio.OT

Computational strategies for cross-species knowledge transfer

Research organisms provide invaluable insights into human biology and diseases, serving as essential tools for functional experiments, disease modeling, and drug testing. However, evolutionary divergence between humans and research organisms hinders effective knowledge transfer across species. Here, we review state-of-the-art methods for computationally transferring knowledge across species, primarily focusing on methods that utilize transcriptome data and/or molecular networks. Our review addresses four key areas: (1) transferring disease and gene annotation knowledge across species, (2) identifying functionally equivalent molecular components, (3) inferring equivalent perturbed genes or gene sets, and (4) identifying equivalent cell types. We conclude with an outlook on future directions and several key challenges that remain in cross-species knowledge transfer, including introducing the concept of "agnology" to describe functional equivalence of biological entities, regardless of their evolutionary origins. This concept is becoming pervasive in integrative data-driven models where evolutionary origins of functions can remain unresolved.

q-bio.GN

On the phase diagram of the polymer model

In dimensions 3 or larger, it is a classical fact that the directed polymer model has two phases: Brownian behavior at high temperature, and non-Brownian behavior at low temperature. We consider the response of the polymer to an external field or tilt, and show that at fixed temperature, the polymer has Brownian behavior for some fields and non-Brownian behavior for others. In other words, the external field can induce the phase transition in the directed polymer model.

math.PR

Current and future directions in network biology

Network biology is an interdisciplinary field bridging computational and biological sciences that has proved pivotal in advancing the understanding of cellular functions and diseases across biological systems and scales. Although the field has been around for two decades, it remains nascent. It has witnessed rapid evolution, accompanied by emerging challenges. These challenges stem from various factors, notably the growing complexity and volume of data together with the increased diversity of data types describing different tiers of biological organization. We discuss prevailing research directions in network biology and highlight areas of inference and comparison of biological networks, multimodal data integration and heterogeneous networks, higher-order network analysis, machine learning on networks, and network-based personalized medicine. Following the overview of recent breakthroughs across these five areas, we offer a perspective on the future directions of network biology. Additionally, we offer insights into scientific communities, educational initiatives, and the importance of fostering diversity within the field. This paper establishes a roadmap for an immediate and long-term vision for network biology.

q-bio.MN

Accurately Modeling Biased Random Walks on Weighted Graphs Using $\textit{Node2vec+}$

Node embedding is a powerful approach for representing the structural role of each node in a graph. $\textit{Node2vec}$ is a widely used method for node embedding that works by exploring the local neighborhoods via biased random walks on the graph. However, $\textit{node2vec}$ does not consider edge weights when computing walk biases. This intrinsic limitation prevents $\textit{node2vec}$ from leveraging all the information in weighted graphs and, in turn, limits its application to many real-world networks that are weighted and dense. Here, we naturally extend $\textit{node2vec}$ to $\textit{node2vec+}$ in a way that accounts for edge weights when calculating walk biases, but which reduces to $\textit{node2vec}$ in the cases of unweighted graphs or unbiased walks. We empirically show that $\textit{node2vec+}$ is more robust to additive noise than $\textit{node2vec}$ in weighted graphs using two synthetic datasets. We also demonstrate that $\textit{node2vec+}$ significantly outperforms $\textit{node2vec}$ on a commonly benchmarked multi-label dataset (Wikipedia). Furthermore, we test $\textit{node2vec+}$ against GCN and GraphSAGE using various challenging gene classification tasks on two protein-protein interaction networks. Despite some clear advantages of GCN and GraphSAGE, they show comparable performance with $\textit{node2vec+}$. Finally, $\textit{node2vec+}$ can be used as a general approach for generating biased random walks, benefiting all existing methods built on top of $\textit{node2vec}$. $\textit{Node2vec+}$ is implemented as part of $\texttt{PecanPy}$, which is available at https://github.com/krishnanlab/PecanPy .

cs.SI

Negative correlation of adjacent Busemann increments

We consider i.i.d. last-passage percolation on $\mathbb{Z}^2$ with weights having distribution $F$ and time-constant $g_F$. We provide an explicit condition on the large deviation rate function for independent sums of $F$ that determines when some adjacent Busemann function increments are negatively correlated. As an example, we prove that $\operatorname{Bernoulli}(p)$ weights for $p > p^* \approx 0.6504$ satisfy this condition. We prove this condition by establishing a direct relationship between the negative correlations of adjacent Busemann increments and the dominance of the time-constant $g_F$ by the function describing the time-constant of last-passage percolation with exponential or geometric weights.

math.PR

Geodesic length and shifted weights in first-passage percolation

We study first-passage percolation through related optimization problems over paths of restricted length. The path length variable is in duality with a shift of the weights. This puts into a convex duality framework old observations about the convergence of the normalized Euclidean length of geodesics due to Hammersley and Welsh, Smythe and Wierman, and Kesten, and leads to new results about geodesic length and the regularity of the shape function as a function of the weight shift. For points far enough away from the origin, the ratio of the geodesic length and the $\ell^1$ distance to the endpoint is uniformly bounded away from one. The shape function is a strictly concave function of the weight shift. Atoms of the weight distribution generate singularities, that is, points of nondifferentiability, in this function. We generalize to all distributions, directions and dimensions an old singularity result of Steele and Zhang for the planar Bernoulli case. When the weight distribution has two or more atoms, a dense set of shifts produce singularities. The results come from a combination of the convex duality, the shape theorems of the different first-passage optimization problems, and modification arguments.

math.PR

Reconciling Multiple Connectivity Scores for Drug Repurposing

The basis of several recent methods for drug repurposing is the key principle that an efficacious drug will reverse the disease molecular 'signature' with minimal side-effects. This principle was defined and popularized by the influential 'connectivity map' study in 2006 regarding reversal relationships between disease- and drug-induced gene expression profiles, quantified by a disease-drug 'connectivity score.' Over the past 15 years, several studies have proposed variations in calculating connectivity scores towards improving accuracy and robustness in light of massive growth in reference drug profiles. However, these variations have been formulated inconsistently using various notations and terminologies even though they are based on a common set of conceptual and statistical ideas. Therefore, we present a systematic reconciliation of multiple disease-drug similarity metrics (ES, css, Sum, Cosine, XSum, XCor, XSpe, XCos, EWCos) and connectivity scores (CS, RGES, NCS, WCS, Tau, CSS, EMUDRA) by defining them using consistent notation and terminology. In addition to providing clarity and deeper insights, this coherent definition of connectivity scores and their relationships provides a unified scheme that newer methods can adopt, enabling the computational drug-development community to compare and investigate different approaches easily. To facilitate the continuous and transparent integration of newer methods, this article will be available as a live document (https://jravilab.github.io/connectivity_scores) coupled with a GitHub repository (https://github.com/jravilab/connectivity_scores) that any researcher can build on and push changes to.

q-bio.QM

Stationary coalescing walks on the lattice II: Entropy

This paper is a sequel to Chaika and Krishnan [arXiv:1612.00434]. We again consider translation invariant measures on families of nearest-neighbor semi-infinite walks on the integer lattice Z^d. We assume that once walks meet, they coalesce. We consider various entropic properties of these systems. We show that in systems with completely positive entropy, bi-infinite trajectories must carry entropy. In the case of directed walks in dimension 2 we show that positive entropy guarantees that all trajectories cannot be bi-infinite. To show that our theorems are proper, we construct a stationary discrete-time symmetric exclusion process whose particle trajectories form bi-infinite trajectories carrying entropy.

math.PR

Kostka Numbers and Longest Increasing Subsequences

A classical bijection relates certain Kostka numbers, the Catalan numbers, and permutations of length $n$ with longest increasing subsequence (LIS) of length at most $2.$ We generalize this bijection and find Kostka numbers which count the number of permutations of $n$ with LIS length at most $w,$ the number of permutations with $(1, \cdots, w)$ as a LIS, and other similar subsets of permutations.

math.CO

Stationary coalescing walks on the lattice

We consider translation invariant measures on families of nearest-neighbor semi-infinite walks on the integer lattice. We assume that once walks meet, they coalesce. In $2d$, we classify the collective behavior of these walks under mild assumptions: they either coalesce almost surely or form bi-infinite trajectories. Bi-infinite trajectories form measure-preserving dynamical systems, have a common asymptotic direction in $2d$, and possess other nice properties. We use our theory to classify the behavior of compatible families of semi-infinite geodesics in stationary first- and last-passage percolation. We also partially answer a question raised by C. Hoffman about the limiting empirical measure of weights seen by geodesics. We construct several examples: our main example is a standard first-passage percolation model where geodesics coalesce almost surely, but have no asymptotic direction or average weight.

math.PR

Tracy-Widom fluctuations for perturbations of the log-gamma polymer in intermediate disorder

The free-energy fluctuations of the discrete directed polymer in 1+1 dimensions is conjecturally in the Tracy-Widom universality class at all finite temperatures and in the intermediate disorder regime. Sepp\"al\"ainen's log-gamma polymer was proven to have GUE Tracy-Widom fluctuations in a restricted temperature range by Borodin et. al. (2013). We remove this restriction, and extend this result into the intermediate disorder regime. This result also identifies the scale of fluctuations of the log-gamma polymer in the intermediate disorder regime, and thus verifies a conjecture of Alberts et. al. (2010). Using a perturbation argument, we show that any polymer that matches a certain number of moments with the log-gamma polymer also has Tracy-Widom fluctuations in intermediate disorder.

math.PR

Variational formula for the time-constant of first-passage percolation

We consider first-passage percolation with positive, stationary-ergodic weights on the square lattice $\mathbb{Z}^d$. Let $T(x)$ be the first-passage time from the origin to a point $x$ in $\mathbb{Z}^d$. The convergence of the scaled first-passage time $T([nx])/n$ to the time-constant as $n$ tends to infinity can be viewed as a problem of homogenization for a discrete Hamilton-Jacobi-Bellman (HJB) equation. By borrowing several tools from the continuum theory of stochastic homogenization for HJB equations, we derive an exact variational formula for the time-constant. We then construct an explicit iteration that produces the minimizer of the variational formula (under a symmetry assumption), thereby computing the time-constant. The variational formula may also be seen as a duality principle, and we discuss some aspects of this duality.

math.PR

Variational formula for the time-constant of first-passage percolation

We consider first-passage percolation with positive, stationary-ergodic weights on the square lattice $\mathbb{Z}^d$. Let $T(x)$ be the first-passage time from the origin to a point $x$ in $\mathbb{Z}^d$. The convergence of the scaled first-passage time $T([nx])/n$ to the time-constant as $n \to \infty$ can be viewed as a problem of homogenization for a discrete Hamilton-Jacobi-Bellman (HJB) equation. We derive an exact variational formula for the time-constant, and construct an explicit iteration that produces a minimizer of the variational formula (under a symmetry assumption). We explicitly identify when the iteration produces correctors.

math.PR