Searcharxiv⌕ Search

arXiv subjects

Xinjie Mao

Publications and source records attributed to Xinjie Mao.

7 recordsLinked to original sources

SCALE:Scalable Conditional Atlas-Level Endpoint transport for virtual cell perturbation prediction

Virtual-cell models aim to predict how cell populations respond to perturbations, but control and treated cells are measured as unpaired populations, complicating the learning of perturbation-specific effects. We present SCALE, a conditional transport model that represents cells as unordered sets and predicts treated populations without cell-level matching. A shared set-aware encoder and conditional DiT backbone learn latent transport, making endpoint supervision directly delta-aligned without an auxiliary delta objective. Across genetic, chemical, developmental and immune perturbations, SCALE recovered gene-expression changes, response directions and population structure. In CRISPR data with dominant cell-line effects, SCALE outperformed competing methods across seven metrics and maintained separation among gene-target representations rather than collapsing them into a shared region. SCALE further prioritized cytokines predicted to produce distinct immune activation and inflammatory responses. Experiments using matched PBMC samples from three donors confirmed these predicted differences. Together, SCALE enables perturbation-specific prediction from unpaired populations and supports experimental prioritization.

cs.LG↗

AblateCell: A Reproduce-then-Ablate Agent for Virtual Cell Repositories

Systematic ablations are essential to attribute performance gains in AI Virtual Cells, yet they are rarely performed because biological repositories are under-standardized and tightly coupled to domain-specific data and formats. While recent coding agents can translate ideas into implementations, they typically stop at producing code and lack a verifier that can reproduce strong baselines and rigorously test which components truly matter. We introduce AblateCell, a reproduce-then-ablate agent for virtual cell repositories that closes this verification gap. AblateCell first reproduces reported baselines end-to-end by auto-configuring environments, resolving dependency and data issues, and rerunning official evaluations while emitting verifiable artifacts. It then conducts closed-loop ablation by generating a graph of isolated repository mutations and adaptively selecting experiments under a reward that trades off performance impact and execution cost. Evaluated on three single-cell perturbation prediction repositories (CPA, GEARS, BioLORD), AblateCell achieves 88.9% (+29.9% to human expert) end-to-end workflow success and 93.3% (+53.3% to heuristic) accuracy in recovering ground-truth critical components. These results enable scalable, repository-grounded verification and attribution directly on biological codebases.

cs.AI↗

Benchmarking virtual cell models for in-the-wild perturbation response

Virtual cell (VC) models aim to predict cellular responses to any perturbations in silico and have emerged as a promising approach for drug discovery and precision medicine. Yet, a clear gap still remains: while models routinely reported impressive results on standard benchmarks, it is unclear whether their predictions are truly meaningful in practice. This is mainly due to limitations in current evaluation setups, which are often overly simplified or inconsistent, and do not reflect the complexity and variability of real biological systems. Here, we introduce a standardized and modular benchmarking framework for virtual cell prediction. Our framework evaluates diverse models under in-the-wild challenging scenarios, including unseen cell contexts, unseen perturbations, and cross-dataset generalization, which better reflect practical applications. Our analysis shows that model performance is highly context-dependent and shaped by task design and evaluation criteria. In commonly used setups, performance is often overestimated, and naive dataset aggregation can even reduce performance. When evaluated under more strict conditions, model performance drops markedly, indicating limited robustness to shifts across cellular contexts. In unseen perturbation settings, models including simple linear approaches capture global transcriptional trends but fail to recover fine-grained perturbation-specific effects. In addition, different evaluation metrics focus on different biological properties, leading to substantially different model rankings. Together, our framework provides a more reliable and biologically grounded evaluation, offering clearer guidance for applying virtual cell models in real scenarios.

q-bio.CB↗

HarmonyCell: Automating Single-Cell Perturbation Modeling under Semantic and Distribution Shifts

Single-cell perturbation studies face dual heterogeneity bottlenecks: (i) semantic heterogeneity--identical biological concepts encoded under incompatible metadata schemas across datasets; and (ii) statistical heterogeneity--distribution shifts from biological variation demanding dataset-specific inductive biases. We propose HarmonyCell, an end-to-end agent framework resolving each challenge through a dedicated mechanism: an LLM-driven Semantic Unifier autonomously maps disparate metadata into a canonical interface without manual intervention; and an adaptive Monte Carlo Tree Search engine operates over a hierarchical action space to synthesize architectures with optimal statistical inductive biases for distribution shifts. Evaluated across diverse perturbation tasks under both semantic and distribution shifts, HarmonyCell achieves a 95% valid execution rate on heterogeneous input datasets (versus 0% for general agents) while matching or even exceeding expert-designed baselines in rigorous out-of-distribution evaluations. This dual-track orchestration enables scalable automatic virtual cell modeling without dataset-specific engineering.

cs.AI↗

Accurate de novo sequencing of the modified proteome with OmniNovo

Post-translational modifications (PTMs) serve as a dynamic chemical language regulating protein function, yet current proteomic methods remain blind to a vast portion of the modified proteome. Standard database search algorithms suffer from a combinatorial explosion of search spaces, limiting the identification of uncharacterized or complex modifications. Here we introduce OmniNovo, a unified deep learning framework for reference-free sequencing of unmodified and modified peptides directly from tandem mass spectra. Unlike existing tools restricted to specific modification types, OmniNovo learns universal fragmentation rules to decipher diverse PTMs within a single coherent model. By integrating a mass-constrained decoding algorithm with rigorous false discovery rate estimation, OmniNovo achieves state-of-the-art accuracy, identifying 51\% more peptides than standard approaches at a 1\% false discovery rate. Crucially, the model generalizes to biological sites unseen during training, illuminating the dark matter of the proteome and enabling unbiased comprehensive analysis of cellular regulation.

q-bio.QM↗

Observations of Running Penumbral Waves emerging in a Sunspot

We present results from the investigation of 5-min umbral oscillations in a single-polarity sunspot of active region NOAA 12132. The spectra of TiO, H$α$, and 304 Å are used for corresponding atmospheric heights from the photosphere to lower corona. Power spectrum analysis at the formation height of H$α$ - 0.6 Å to H$α$ center resulted in the detection of 5-min oscillation signals in intensity interpreted as running waves outside the umbral center, mostly with vertical magnetic field inclination $>15°$. A phase-speed filter is used to extract the running wave signals with speed $v_{ph}> 4$ km s$^{-1}$, from the time series of H$α$ - 0.4 Å images, and found twenty-four 3-min umbral oscillatory events in a duration of one hour. Interestingly, the initial emergence of the 3-min umbral oscillatory events are noticed closer to or at umbral boundaries. These 3-min umbral oscillatory events are observed for the first time as propagating from a fraction of preceding Running Penumbral Waves (RPWs). These fractional wavefronts rapidly separates from RPWs and move towards umbral center, wherein they expand radially outwards suggesting the beginning of a new umbral oscillatory event. We found that most of these umbral oscillatory events develop further into RPWs. We speculate that the waveguides of running waves are twisted in spiral structures and hence the wavefronts are first seen at high latitudes of umbral boundaries and later at lower latitudes of the umbral center.

astro-ph.SR↗

Improved magnetogram calibration of SMFT and its comparison with the HMI

In this paper, we try to improve the magnetogram calibration method of the Solar Magnetic Field Telescope (SMFT). The improved calibration process fits the observed full Stokes information, using six points on the profile of Fe ı 5324.18 Å line, and the analytical Stokes profiles under the Milne-Eddington atmosphere model, adopting the Levenberg-Marquardt least-square fitting algorithm. In Comparison with the linear calibration methods, which employs one point, there is large difference in the strength of longitudinal field $B_l$ and tranverse field $B_t$, caused by the non-linear relationship, but the discrepancy is little in the case of inclination and azimuth. We conclude that it is better to deal with the non-linear effects in the calibration of $B_l$ and $B_t$ using six points. Moreover, in comparison with SDO/HMI, SMFT has larger stray light and acquires less magnetic field strength. For vector magnetic fields in two sunspot regions, the magnetic field strength, inclination and azimuth angles between SMFT and HMI are roughly in agrement, with the linear fitted slope of 0.73/0.7, 0.95/1.04 and 0.99/1.1. In the case of pores and quiet regions ($B_l$ $<$ 600 G), the fitted slopes of the longitudinal magnetic field strength are 0.78 and 0.87 respectively.

astro-ph.SR↗