SearcharxivSearch

arXiv subjects

Yaning Yang

Publications and source records attributed to Yaning Yang.

15 recordsLinked to original sources

DeepTaxon: An Interpretable Retrieval-Augmented Multimodal Framework for Unified Species Identification and Discovery

Identifying species in biology among tens of thousands of visually similar taxa while discovering unknown species in open-world environments remains a fundamental challenge in biodiversity research. Current methods treat identification and discovery as separate problems, with classification models assuming closed sets and discovery relying on threshold-based rejection. Here we present DeepTaxon, a retrieval-augmented multimodal framework that unifies species identification and discovery through interpretable reasoning over retrieved visual evidence. Given a query image, DeepTaxon retrieves the top-$k$ candidate species with $n$ exemplar images each from a retrieval index and performs chain-of-thought comparative reasoning. Critically, we redefine discovery as an explicit, retrieval-based decision problem rather than an implicit parametric memory problem. A sample is novel if and only if the retrieval index lacks sufficient evidence for identification, so each retrieval naturally yields a classification or discovery label without manual annotation, thereby providing automatic supervision for both tasks. We train the framework via supervised fine-tuning on synthetic retrieval-augmented data, followed by reinforcement learning on hard samples, converting high-recall retrieval into high-precision decisions that scale to massive taxonomic vocabularies. Extensive experiments on a large-scale in-distribution benchmark and six out-of-distribution datasets demonstrate consistent improvements in both identification and discovery. Ablation studies further reveal effective test-time scaling with candidate count $k$ and exemplar count $n$, strong zero-shot transfer to unseen domains, and consistent performance across retrieval encoders, establishing an interpretable solution for biodiversity research.

cs.CV

STAMP: Multi-pattern Attention-aware Multiple Instance Learning for STAS Diagnosis in Multi-center Histopathology Images

Spread through air spaces (STAS) constitutes a novel invasive pattern in lung adenocarcinoma (LUAD), associated with tumor recurrence and diminished survival rates. However, large-scale STAS diagnosis in LUAD remains a labor-intensive endeavor, compounded by the propensity for oversight and misdiagnosis due to its distinctive pathological characteristics and morphological features. Consequently, there is a pressing clinical imperative to leverage deep learning models for STAS diagnosis. This study initially assembled histopathological images from STAS patients at the Second Xiangya Hospital and the Third Xiangya Hospital of Central South University, alongside the TCGA-LUAD cohort. Three senior pathologists conducted cross-verification annotations to construct the STAS-SXY, STAS-TXY, and STAS-TCGA datasets. We then propose a multi-pattern attention-aware multiple instance learning framework, named STAMP, to analyze and diagnose the presence of STAS across multi-center histopathology images. Specifically, the dual-branch architecture guides the model to learn STAS-associated pathological features from distinct semantic spaces. Transformer-based instance encoding and a multi-pattern attention aggregation modules dynamically selects regions closely associated with STAS pathology, suppressing irrelevant noise and enhancing the discriminative power of global representations. Moreover, a similarity regularization constraint prevents feature redundancy across branches, thereby improving overall diagnostic accuracy. Extensive experiments demonstrated that STAMP achieved competitive diagnostic results on STAS-SXY, STAS-TXY and STAS-TCGA, with AUCs of 0.8058, 0.8017, and 0.7928, respectively, surpassing the clinical level.

cs.CV

Situation-Dependent Causal Influence-Based Cooperative Multi-agent Reinforcement Learning

Learning to collaborate has witnessed significant progress in multi-agent reinforcement learning (MARL). However, promoting coordination among agents and enhancing exploration capabilities remain challenges. In multi-agent environments, interactions between agents are limited in specific situations. Effective collaboration between agents thus requires a nuanced understanding of when and how agents' actions influence others. To this end, in this paper, we propose a novel MARL algorithm named Situation-Dependent Causal Influence-Based Cooperative Multi-agent Reinforcement Learning (SCIC), which incorporates a novel Intrinsic reward mechanism based on a new cooperation criterion measured by situation-dependent causal influence among agents. Our approach aims to detect inter-agent causal influences in specific situations based on the criterion using causal intervention and conditional mutual information. This effectively assists agents in exploring states that can positively impact other agents, thus promoting cooperation between agents. The resulting update links coordinated exploration and intrinsic reward distribution, which enhance overall collaboration and performance. Experimental results on various MARL benchmarks demonstrate the superiority of our method compared to state-of-the-art approaches.

cs.AI

Likelihood ratio tests in random graph models with increasing dimensions

We explore the Wilks phenomena in two random graph models: the $\beta$-model and the Bradley-Terry model. For two increasing dimensional null hypotheses, including a specified null $H_0: \beta_i=\beta_i^0$ for $i=1,\ldots, r$ and a homogenous null $H_0: \beta_1=\cdots=\beta_r$, we reveal high dimensional Wilks' phenomena that the normalized log-likelihood ratio statistic, $[2\{\ell(\widehat{\mathbf{\beta}}) - \ell(\widehat{\mathbf{\beta}}^0)\} - r]/(2r)^{1/2}$, converges in distribution to the standard normal distribution as $r$ goes to infinity. Here, $\ell( \mathbf{\beta})$ is the log-likelihood function on the model parameter $\mathbf{\beta}=(\beta_1, \ldots, \beta_n)^\top$, $\widehat{\mathbf{\beta}}$ is its maximum likelihood estimator (MLE) under the full parameter space, and $\widehat{\mathbf{\beta}}^0$ is the restricted MLE under the null parameter space. For the homogenous null with a fixed $r$, we establish Wilks-type theorems that $2\{\ell(\widehat{\mathbf{\beta}}) - \ell(\widehat{\mathbf{\beta}}^0)\}$ converges in distribution to a chi-square distribution with $r-1$ degrees of freedom, as the total number of parameters, $n$, goes to infinity. When testing the fixed dimensional specified null, we find that its asymptotic null distribution is a chi-square distribution in the $\beta$-model. However, unexpectedly, this is not true in the Bradley-Terry model. By developing several novel technical methods for asymptotic expansion, we explore Wilks type results in a principled manner; these principled methods should be applicable to a class of random graph models beyond the $\beta$-model and the Bradley-Terry model. Simulation studies and real network data applications further demonstrate the theoretical results.

math.ST

Wilks' theorems in the $\beta$-model

Likelihood ratio tests and the Wilks theorems have been pivotal in statistics but have rarely been explored in network models with an increasing dimension. We are concerned here with likelihood ratio tests in the $\beta$-model for undirected graphs. For two growing dimensional null hypotheses including a specified null $H_0: \beta_i=\beta_i^0$ for $i=1,\ldots, r$ and a homogenous null $H_0: \beta_1=\cdots=\beta_r$, we reveal high dimensional Wilks' phenomena that the normalized log-likelihood ratio statistic, $[2\{\ell(\widehat{\boldsymbol{\beta}}) - \ell(\widehat{\boldsymbol{\beta}}^0)\} - r]/(2r)^{1/2}$, converges in distribution to the standard normal distribution as $r$ goes to infinity. Here, $\ell( \boldsymbol{\beta})$ is the log-likelihood function on the vector parameter $\boldsymbol{\beta}=(\beta_1, \ldots, \beta_n)^\top$, $\widehat{\boldsymbol{\beta}}$ is its maximum likelihood estimator (MLE) under the full parameter space, and $\widehat{\boldsymbol{\beta}}^0$ is the restricted MLE under the null parameter space. For the corresponding fixed dimensional null $H_0: \beta_i=\beta_i^0$ for $i=1,\ldots, r$ and the homogenous null $H_0: \beta_1=\cdots=\beta_r$ with a fixed $r$, we establish Wilks type of results that $2\{\ell(\widehat{\boldsymbol{\beta}}) - \ell(\widehat{\boldsymbol{\beta}}^0)\}$ converges in distribution to a Chi-square distribution with respective $r$ and $r-1$ degrees of freedom, as the total number of parameters, $n$, goes to infinity. The Wilks type of results are further extended into a closely related Bradley--Terry model for paired comparisons, where we discover a different phenomenon that the log-likelihood ratio statistic under the fixed dimensional specified null asymptotically follows neither a Chi-square nor a rescaled Chi-square distribution. Simulation studies and an application to NBA data illustrate the theoretical results.

math.ST

Multi-task Joint Strategies of Self-supervised Representation Learning on Biomedical Networks for Drug Discovery

Self-supervised representation learning (SSL) on biomedical networks provides new opportunities for drug discovery. However, how to effectively combine multiple SSL models is still challenging and has been rarely explored. Therefore, we propose multi-task joint strategies of self-supervised representation learning on biomedical networks for drug discovery, named MSSL2drug. We design six basic SSL tasks inspired by various modality features including structures, semantics, and attributes in heterogeneous biomedical networks. Importantly, fifteen combinations of multiple tasks are evaluated by a graph attention-based multi-task adversarial learning framework in two drug discovery scenarios. The results suggest two important findings. (1) Combinations of multimodal tasks achieve the best performance compared to other multi-task joint models. (2) The local-global combination models yield higher performance than random two-task combinations when there are the same size of modalities. Therefore, we conjecture that the multimodal and local-global combination strategies can be treated as the guideline of multi-task SSL for drug discovery.

cs.LG

Beyond Zipf's Law: The Lavalette Rank Function and its Properties

Although Zipf's law is widespread in natural and social data, one often encounters situations where one or both ends of the ranked data deviate from the power-law function. Previously we proposed the Beta rank function to improve the fitting of data which does not follow a perfect Zipf's law. Here we show that when the two parameters in the Beta rank function have the same value, the Lavalette rank function, the probability density function can be derived analytically. We also show both computationally and analytically that Lavalette distribution is approximately equal, though not identical, to the lognormal distribution. We illustrate the utility of Lavalette rank function in several datasets. We also address three analysis issues on the statistical testing of Lavalette fitting function, comparison between Zipf's law and lognormal distribution through Lavalette function, and comparison between lognormal distribution and Lavalette distribution.

physics.data-an

Using Volcano Plots and Regularized-Chi Statistics in Genetic Association Studies

Labor intensive experiments are typically required to identify the causal disease variants from a list of disease associated variants in the genome. For designing such experiments, candidate variants are ranked by their strength of genetic association with the disease. However, the two commonly used measures of genetic association, the odds-ratio (OR) and p-value, may rank variants in different order. To integrate these two measures into a single analysis, here we transfer the volcano plot methodology from gene expression analysis to genetic association studies. In its original setting, volcano plots are scatter plots of fold-change and t-test statistic (or -log of the p-value), with the latter being more sensitive to sample size. In genetic association studies, the OR and Pearson's chi-square statistic (or equivalently its square root, chi; or the standardized log(OR)) can be analogously used in a volcano plot, allowing for their visual inspection. Moreover, the geometric interpretation of these plots leads to an intuitive method for filtering results by a combination of both OR and chi-square statistic, which we term "regularized-chi". This method selects associated markers by a smooth curve in the volcano plot instead of the right-angled lines which corresponds to independent cutoffs for OR and chi-square statistic. The regularized-chi incorporates relatively more signals from variants with lower minor-allele-frequencies than chi-square test statistic. As rare variants tend to have stronger functional effects, regularized-chi is better suited to the task of prioritization of candidate genes.

q-bio.QM

Grouped sparse paired comparisons in the Bradley-Terry model

In a wide class of paired comparisons, especially in the sports games, in which all subjects are divided into several groups, the intragroup comparisons are dense and the intergroup comparisons are sparse. Typical examples include the NFL regular season. Motivated by these situations, we propose group sparsity for paired comparisons and show the consistency and asymptotical normality of the maximum likelihood estimate in the Bradley-Terry model when the number of parameters goes to infinity in this paper. Simulations are carried out to illustrate the group sparsity and asymptotical results.

stat.OT

Wilks' theorems in some exponential random graph models

We are concerned here with the likelihood ratio statistics in two exponential random graph models -- the $\beta$-model and the Bradley-Terry model, in which the degree sequence on an undirected graph and the out-degree sequence on a weighted directed graph are the exclusively sufficient statistics in the exponential-family distributions on graphs, respectively. We prove the Wilks type of theorems for some fixed and growing dimensional hypothesis testing problems. More specifically, under two fixed dimensional null hypotheses $H_0: \beta_i=\beta_i^0$ for $i=1,\ldots, r$ and $H_0: \beta_1=\ldots=\beta_r$, we show that $2[\ell(\widehat{\boldsymbol{\beta}}) - \ell(\widehat{\boldsymbol{\beta}}^0)]$ converges in distribution to a Chi-square distribution with the respective degrees of freedoms, $r$ and $r-1$, as the dimension $n$ of the full parameter space goes to infinity. Here, $\ell(\boldsymbol{\beta})$ is the log-likelihood function on the parameter $\boldsymbol{\beta}$, $\widehat{\boldsymbol{\beta}}$ is the MLE under the full parameter space, and $\widehat{\boldsymbol{\beta}}^0$ is the restricted MLE under the null parameter space. For two increasing dimensional null hypotheses $H_0: \beta_i = \beta_i^0$ for $i=1, \ldots, n$ and $H_0: \beta_1=\ldots=\beta_r$ with $r/n \ge c$, we show that the normalized log-likelihood ratio statistics, $(2[\ell(\widehat{\boldsymbol{\beta}}) - \ell(\boldsymbol{\beta}^0)] -n)/(2n)^{1/2}$ and $(2[\ell(\widehat{\boldsymbol{\beta}}) - \ell(\widehat{\boldsymbol{\beta}}^0)] -r)/(2r)^{1/2}$, both converge in distribution to the standard normal distribution. Simulation studies and an application to NBA data illustrate the theoretical results.

math.ST

Effective Sample Size: Quick Estimation of the Effect of Related Samples in Genetic Case-Control Association Analyses

Affected relatives are essential for pedigree linkage analysis, however, they cause a violation of the independent sample assumption in case-control association studies. To avoid the correlation between samples, a common practice is to take only one affected sample per pedigree in association analysis. Although several methods exist in handling correlated samples, they are still not widely used in part because these are not easily implemented, or because they are not widely known. We advocate the effective sample size method as a simple and accessible approach for case-control association analysis with correlated samples. This method modifies the chi-square test statistic, p-value, and 95% confidence interval of the odds-ratio by replacing the apparent number of allele or genotype counts with the effective ones in the standard formula, without the need for specialized computer programs. We present a simple formula for calculating effective sample size for many types of relative pairs and relative sets. For allele frequency estimation, the effective sample size method captures the variance inflation exactly. For genotype frequency, simulations showed that effective sample size provides a satisfactory approximation. A gene which is previously identified as a type 1 diabetes susceptibility locus, the interferon-induced helicase gene (IFIH1), is shown to be significantly associated with rheumatoid arthritis when the effective sample size method is applied. This significant association is not established if only one affected sib per pedigree were used in the association analysis. Relationship between the effective sample size method and other methods -- the generalized estimation equation, variance of eigenvalues for correlation matrices, and genomic controls -- are discussed.

q-bio.QM

Partial correlation analysis indicates causal relationships between GC-content, exon density and recombination rate in the human genome

{\bf Background}: Several features are known to correlate with the GC-content in the human genome, including recombination rate, gene density and distance to telomere. However, by testing for pairwise correlation only, it is impossible to distinguish direct associations from indirect ones and to distinguish between causes and effects. {\bf Results}: We use partial correlations to construct partially directed graphs for the following four variables: GC-content, recombination rate, exon density and distance-to-telomere. Recombination rate and exon density are unconditionally uncorrelated, but become inversely correlated by conditioning on GC-content. This pattern indicates a model where recombination rate and exon density are two independent causes of GC-content variation. {\bf Conclusions}: Causal inference and graphical models are useful methods to understand genome evolution and the mechanisms of isochore evolution in the human genome.

q-bio.GN

Exploring Case-Control Genetic Association Tests Using Phase Diagrams

Background: By a new concept called "phase diagram", we compare two commonly used genotype-based tests for case-control genetic analysis, one is a Cochran-Armitage trend test (CAT test at $x=0.5$, or CAT0.5) and another (called MAX2) is the maximization of two chi-square test results: one from the two-by-two genotype count table that combines the baseline homozygotes and heterozygotes, and another from the table that combines heterozygotes with risk homozygotes. CAT0.5 is more suitable for multiplicative disease models and MAX2 is better for dominant/recessive models. Methods: We define the CAT0.5-MAX2 phase diagram on the disease model space such that regions where MAX2 is more powerful than CAT0.5 are separated from regions where the CAT0.5 is more powerful, and the task is to choose the appropriate parameterization to make the separation possible. Results: We find that using the difference of allele frequencies ($δ_p$) and the difference of Hardy-Weinberg disequilibrium coefficients ($δ_ε$) can separate the two phases well, and the phase boundaries are determined by the angle $tan^{-1}(δ_p/δ_ε)$, which is an improvement over the disease model selection using $δ_ε$ only. Conclusions: We argue that phase diagrams similar to the one for CAT0.5-MAX2 have graphical appeals in understanding power performance of various tests, clarifying simulation schemes, summarizing case-control datasets, and guessing the possible mode of inheritance.

q-bio.PE

Fractal Characterizations of MAX Statistical Distribution in Genetic Association Studies

Two non-integer parameters are defined for MAX statistics, which are maxima of $d$ simpler test statistics. The first parameter, $d_{MAX}$, is the fractional number of tests, representing the equivalent numbers of independent tests in MAX. If the $d$ tests are dependent, $d_{MAX} < d$. The second parameter is the fractional degrees of freedom $k$ of the chi-square distribution $χ^2_k$ that fits the MAX null distribution. These two parameters, $d_{MAX}$ and $k$, can be independently defined, and $k$ can be non-integer even if $d_{MAX}$ is an integer. We illustrate these two parameters using the example of MAX2 and MAX3 statistics in genetic case-control studies. We speculate that $k$ is related to the amount of ambiguity of the model inferred by the test. In the case-control genetic association, tests with low $k$ (e.g. $k=1$) are able to provide definitive information about the disease model, as versus tests with high $k$ (e.g. $k=2$) that are completely uncertain about the disease model. Similar to Heisenberg's uncertain principle, the ability to infer disease model and the ability to detect significant association may not be simultaneously optimized, and $k$ seems to measure the level of their balance.

q-bio.QM

How Many Genes Are Needed for a Discriminant Microarray Data Analysis ?

The analysis of the leukemia data from Whitehead/MIT group is a discriminant analysis (also called a supervised learning). Among thousands of genes whose expression levels are measured, not all are needed for discriminant analysis: a gene may either not contribute to the separation of two types of tissues/cancers, or it may be redundant because it is highly correlated with other genes. There are two theoretical frameworks in which variable selection (or gene selection in our case) can be addressed. The first is model selection, and the second is model averaging. We have carried out model selection using Akaike information criterion and Bayesian information criterion with logistic regression (discrimination, prediction, or classification) to determine the number of genes that provide the best model. These model selection criteria set upper limits of 22-25 and 12-13 genes for this data set with 38 samples, and the best model consists of only one (no.4847, zyxin) or two genes. We have also carried out model averaging over the best single-gene logistic predictors using three different weights: maximized likelihood, prediction rate on training set, and equal weight. We have observed that the performance of most of these weighted predictors on the testing set is gradually reduced as more genes are included, but a clear cutoff that separates good and bad prediction performance is not found.

physics.bio-ph