SearcharxivSearch

arXiv subjects

Alexey Sergushichev

Publications and source records attributed to Alexey Sergushichev.

4 recordsLinked to original sources

Hash-augmented adaptive multilevel splitting Monte Carlo algorithm for accurate estimation of two-sample permutation test p-values

Nonparametric permutation tests are widely used for statistical analysis. However, exact computation of test p-values can be algorithmically challenging, particularly for custom tests with complex test statistics. In contrast, Monte Carlo sampling can be easily applied to any test statistic, but it suffers from poor relative accuracy when estimating small p-values, interfering with multiple hypothesis testing correction and leading to other issues. In this work, we present a hash-augmented adaptive multilevel splitting Monte Carlo algorithm that enables accurate estimation of arbitrarily small p-values in two-sample permutation tests. Using the Kolmogorov-Smirnov and the Mann-Whitney U tests as examples, we highlight potential pitfalls related to the discreteness of the test statistic distribution and show how to address them. By comparing with an exact algorithm, we demonstrate the accuracy of the p-value estimates provided by the proposed algorithm and the validity of the associated confidence intervals. We provide a reference implementation of the proposed algorithm in the Python package hamstest, which allows p-value estimation for a user-defined statistic.

stat.ME

Digital Modeling of Spatial Pathway Activity from Histology Reveals Tumor Microenvironment Heterogeneity

Spatial transcriptomics (ST) enables simultaneous mapping of tissue morphology and spatially resolved gene expression, offering unique opportunities to study tumor microenvironment heterogeneity. Here, we introduce a computational framework that predicts spatial pathway activity directly from hematoxylin-and-eosin-stained histology images at microscale resolution 55 and 100 um. Using image features derived from a computational pathology foundation model, we found that TGFb signaling was the most accurately predicted pathway across three independent breast and lung cancer ST datasets. In 87-88% of reliably predicted cases, the resulting spatial TGFb activity maps reflected the expected contrast between tumor and adjacent non-tumor regions, consistent with the known role of TGFb in regulating interactions within the tumor microenvironment. Notably, linear and nonlinear predictive models performed similarly, suggesting that image features may relate to pathway activity in a predominantly linear fashion or that nonlinear structure is small relative to measurement noise. These findings demonstrate that features extracted from routine histopathology may recover spatially coherent and biologically interpretable pathway patterns, offering a scalable strategy for integrating image-based inference with ST information in tumor microenvironment studies.

q-bio.QM

Transcriptome signature for the identification of bevacizumab responders in ovarian cancer

The standard of care for ovarian cancer comprises cytoreductive surgery, followed by adjuvant platinum-based chemotherapy plus taxane therapy and maintenance therapy with the antiangiogenic compound bevacizumab and/or a PARP inhibitor. Nevertheless, there is currently no clear clinical indication for the use of bevacizumab, highlighting the urgent need for biomarkers to assess the response to bevacizumab. In the present study, based on a novel RNA-seq dataset (n=181) and a previously published microarray-based dataset (n=377), we have identified an expression signature potentially associated with benefit from bevacizumab addition and assumed to reflect cancer stemness acquisition driven by activation of CTCFL. Patients with this signature demonstrated improved overall survival when bevacizumab was added to standard chemotherapy in both novel (HR=0.41(0.23-0.74), adj.p-value=7.70e-03) and previously published cohorts (HR=0.51(0.34-0.75), adj.p-value=3.25e-03), while no significant differences in survival explained by treatment were observed in patients negative for this signature. In addition to the CTCFL signature, we found several other reproducible expression signatures which may also represent biomarker candidates not related to established molecular subtypes of ovarian cancer and require further validation studies based on additional RNA-seq data.

q-bio.GN

Assemblathon 2: evaluating de novo methods of genome assembly in three vertebrate species

Background - The process of generating raw genome sequence data continues to become cheaper, faster, and more accurate. However, assembly of such data into high-quality, finished genome sequences remains challenging. Many genome assembly tools are available, but they differ greatly in terms of their performance (speed, scalability, hardware requirements, acceptance of newer read technologies) and in their final output (composition of assembled sequence). More importantly, it remains largely unclear how to best assess the quality of assembled genome sequences. The Assemblathon competitions are intended to assess current state-of-the-art methods in genome assembly. Results - In Assemblathon 2, we provided a variety of sequence data to be assembled for three vertebrate species (a bird, a fish, and snake). This resulted in a total of 43 submitted assemblies from 21 participating teams. We evaluated these assemblies using a combination of optical map data, Fosmid sequences, and several statistical methods. From over 100 different metrics, we chose ten key measures by which to assess the overall quality of the assemblies. Conclusions - Many current genome assemblers produced useful assemblies, containing a significant representation of their genes, regulatory sequences, and overall genome structure. However, the high degree of variability between the entries suggests that there is still much room for improvement in the field of genome assembly and that approaches which work well in assembling the genome of one species may not necessarily work well for another.

q-bio.GN