Searcharxiv⌕ Search

arXiv subjects

Ping Xu

Publications and source records attributed to Ping Xu.

At least 37 records · Page 2Linked to original sources

Formal geometry and Tamarkin--Tsygan calculi of dg manifolds

The main goal of this paper is to study the formal geometry of dg manifolds à la Fedosov. For any dg manifold $(\mathcal{M}, Q)$, we construct a Fedosov dg foliation (or dg Lie algebroid) $\mathcal{F}_Q \to \mathcal{N}_Q$. We establish homotopy contractions between their respective spaces of polyvector fields, differential forms, polydifferential operators, and polyjets. As a consequence, we prove that their respective Cartan calculi and noncommutative calculi, in the sense of Tamarkin--Tsygan, are isomorphic.

math.DG↗

scUnified: An AI-Ready Standardized Resource for Single-Cell RNA Sequencing Analysis

Single-cell RNA sequencing (scRNA-seq) technology enables systematic delineation of cellular states and interactions, providing crucial insights into cellular heterogeneity. Building on this potential, numerous computational methods have been developed for tasks such as cell clustering, cell type annotation, and marker gene identification. To fully assess and compare these methods, standardized, analysis-ready datasets are essential. However, such datasets remain scarce, and variations in data formats, preprocessing workflows, and annotation strategies hinder reproducibility and complicate systematic evaluation of existing methods. To address these challenges, we present scUnified, an AI-ready standardized resource for single-cell RNA sequencing data that consolidates 13 high-quality datasets spanning two species (human and mouse) and nine tissue types. All datasets undergo standardized quality control and preprocessing and are stored in a uniform format to enable direct application in diverse computational analyses without additional data cleaning. We further demonstrate the utility of scUnified through experimental analyses of representative biological tasks, providing a reproducible foundation for the standardized evaluation of computational methods on a unified dataset.

q-bio.GN↗

scSiameseClu: A Siamese Clustering Framework for Interpreting single-cell RNA Sequencing Data

Single-cell RNA sequencing (scRNA-seq) reveals cell heterogeneity, with cell clustering playing a key role in identifying cell types and marker genes. Recent advances, especially graph neural networks (GNNs)-based methods, have significantly improved clustering performance. However, the analysis of scRNA-seq data remains challenging due to noise, sparsity, and high dimensionality. Compounding these challenges, GNNs often suffer from over-smoothing, limiting their ability to capture complex biological information. In response, we propose scSiameseClu, a novel Siamese Clustering framework for interpreting single-cell RNA-seq data, comprising of 3 key steps: (1) Dual Augmentation Module, which applies biologically informed perturbations to the gene expression matrix and cell graph relationships to enhance representation robustness; (2) Siamese Fusion Module, which combines cross-correlation refinement and adaptive information fusion to capture complex cellular relationships while mitigating over-smoothing; and (3) Optimal Transport Clustering, which utilizes Sinkhorn distance to efficiently align cluster assignments with predefined proportions while maintaining balance. Comprehensive evaluations on seven real-world datasets demonstrate that scSiameseClu outperforms state-of-the-art methods in single-cell clustering, cell type annotation, and cell type classification, providing a powerful tool for scRNA-seq data interpretation.

q-bio.GN↗

scCDCG: Efficient Deep Structural Clustering for single-cell RNA-seq via Deep Cut-informed Graph Embedding

Single-cell RNA sequencing (scRNA-seq) is essential for unraveling cellular heterogeneity and diversity, offering invaluable insights for bioinformatics advancements. Despite its potential, traditional clustering methods in scRNA-seq data analysis often neglect the structural information embedded in gene expression profiles, crucial for understanding cellular correlations and dependencies. Existing strategies, including graph neural networks, face challenges in handling the inefficiency due to scRNA-seq data's intrinsic high-dimension and high-sparsity. Addressing these limitations, we introduce scCDCG (single-cell RNA-seq Clustering via Deep Cut-informed Graph), a novel framework designed for efficient and accurate clustering of scRNA-seq data that simultaneously utilizes intercellular high-order structural information. scCDCG comprises three main components: (i) A graph embedding module utilizing deep cut-informed techniques, which effectively captures intercellular high-order structural information, overcoming the over-smoothing and inefficiency issues prevalent in prior graph neural network methods. (ii) A self-supervised learning module guided by optimal transport, tailored to accommodate the unique complexities of scRNA-seq data, specifically its high-dimension and high-sparsity. (iii) An autoencoder-based feature learning module that simplifies model complexity through effective dimension reduction and feature extraction. Our extensive experiments on 6 datasets demonstrate scCDCG's superior performance and efficiency compared to 7 established models, underscoring scCDCG's potential as a transformative tool in scRNA-seq data analysis. Our code is available at: https://github.com/XPgogogo/scCDCG.

cs.LG↗

Kapranov $L_{\infty}[1]$ algebras

Given any Kähler manifold $X$, Kapranov discovered an $L_\infty[1]$ algebra structure on $Ω^{0,\bullet}_X(T^{1,0}_X)$. Motivated by this result, we introduce, as a generalization of $L_\infty[1]$ algebras, a notion of $L_\infty[1]$ $\mathfrak{R}$-algebra, where $\mathfrak{R}$ is a differential graded commutative algebra with unit. We show that standard notions (such as quasi-isomorphism and linearization) and results (including homotopy transfer theorems) can be extended to this context. For instance, we provide a linearization theorem. As an application, we prove that, given any DG Lie algebroid $(\mathcal{L},Q_{\mathcal{L}})$ over a DG manifold $(\mathcal{M},Q)$, there exists an induced $L_\infty[1]$ $\mathfrak{R}$-algebra structure on $Γ(\mathcal{L})$, where $\mathfrak{R}$ is the DG commutative algebra $(C^\infty(\mathcal{M}),Q)$ -- its unary bracket is $Q_{\mathcal{L}}$ while its binary bracket is a cocycle representative of the Atiyah class of the DG Lie algebroid. This $L_\infty[1]$ $\mathfrak{R}$-algebra $Γ(\mathcal{L})$ is linearizable if and only if the Atiyah class of the DG Lie algebroid vanishes. However, the $L_\infty[1]$ ($\mathbb{K}$-)algebra $Γ(\mathcal{L})$ induced by this $L_\infty[1]$ $\mathfrak{R}$-algebra is necessarily homotopy abelian. As a special case, we prove that, given any complex manifold $X$, the Kapranov $L_\infty[1]$ $\mathfrak{R}$-algebra $Ω^{0,\bullet}_X(T^{1,0}_X)$, where $\mathfrak{R}$ is the DG commutative algebra $(Ω^{0,\bullet}_X,\bar{\partial})$, is linearizable if and only if the Atiyah class of the holomorphic tangent bundle $T_X$ vanishes. Nevertheless, the induced $L_\infty[1]$ $\mathbb{C}$-algebra structure on $Ω^{0,\bullet}_X(T^{1,0}_X)$ is necessarily homotopy abelian.

math.DG↗

A zero-dead-time strontium lattice clock with a stability at $10^{-19}$ level

Optical atomic clocks play a crucial role in fundamental physics, relativistic geodesy, and the future redefinition of the SI second. Standard operation relies on cyclic interrogation sequences, which alternate between atomic interrogation and dead time used for state preparation and readout. This approach introduces the Dick effect, where laser frequency noise aliases onto the atomic transition frequency. Although reducing laser noise improves clock stability, the Dick effect remains a key limitation. In this work, we demonstrate a zero-dead-time optical clock based on two interleaved ensembles of cold $^{87}\text{Sr}$ atoms. Our system significantly suppresses this noise and achieves a fractional frequency instability at the $10^{-19}$ level between 10,000 and 20,000 seconds over repeated measurements, with a best value of $2.9 \times 10^{-19}$ at $τ= 20,000$ seconds. The estimated long-term stability based on the combined data of these measurements reaches $2.5 \times 10^{-19}$ at one day. These results represent a more than ninefold improvement over a conventional single-ensemble clock, highlighting its potential for next-generation timekeeping applications.

physics.atom-ph↗

Improved systematic evaluation of a strontium optical clock with uncertainty below $1\times 10^{-18}$

We report a systematic uncertainty of $9.2\times 10^{-19}$ for the USTC Sr1 optical lattice clock, achieving accuracy at the level required for the roadmap of the redefinition of the SI second. A finite-element model with {\it in situ}-validated, spatially-resolved chamber emissivity reduced blackbody radiation shift uncertainty to $6.3\times 10^{-19}$. Concurrently, an externally mounted lattice cavity combined with a larger beam waist suppressed density shifts. Enhanced lattice depth modulation consolidated lattice light shift uncertainty to $6.3\times 10^{-19}$ by enabling simultaneous determination of key polarizabilities and magic wavelength. Magnetic shifts were resolved below $10^{-18}$ via precision characterization of the second-order Zeeman coefficient. Supported by a crystalline-coated ultra-low-expansion cavity-stabilized laser and refined temperature control suppressing BBR fluctuations, the clock also achieves a frequency stability better than $1\times10^{-18}$ at 30,000-s averaging time. These developments collectively establish a new benchmark in USTC Sr1 clock performance and pave the way for high-accuracy applications in metrology and fundamental physics.

physics.atom-ph↗

GHZ-W Genuinely Entangled Subspace Verification with Adaptive Local Measurements

Genuinely entangled subspaces (GESs) are valuable resources in quantum information science. Among these, the three-qubit GHZ-W GES, spanned by the three-qubit Greenberger-Horne-Zeilinger (GHZ) and W states, is a universal and crucial entangled subspace resource for three-qubit systems. In this work, we develop two adaptive verification strategies, the XZ strategy and the rotation strategy, for the three-qubit GHZ-W GES using local measurements and one-way classical communication. These strategies are experimentally feasible, efficient and possess a concise analytical expression for the sample complexity of the rotation strategy, which scales approximately as $2.248ε^{-1}\lnδ^{-1}$, where $ε$ is the infidelity and $1-δ$ is the confidence level. Furthermore, we comprehensively analyze the two-dimensional two-qubit subspaces and classify them into three distinct types, which include unverifiable entangled subspaces, revealing intrinsic limitations in local verification of entangled subspaces.

quant-ph↗

Soft Graph Clustering for single-cell RNA Sequencing Data

Clustering analysis is fundamental in single-cell RNA sequencing (scRNA-seq) data analysis for elucidating cellular heterogeneity and diversity. Recent graph-based scRNA-seq clustering methods, particularly graph neural networks (GNNs), have significantly improved in tackling the challenges of high-dimension, high-sparsity, and frequent dropout events that lead to ambiguous cell population boundaries. However, their reliance on hard graph constructions derived from thresholded similarity matrices presents challenges:(i) The simplification of intercellular relationships into binary edges (0 or 1) by applying thresholds, which restricts the capture of continuous similarity features among cells and leads to significant information loss.(ii) The presence of significant inter-cluster connections within hard graphs, which can confuse GNN methods that rely heavily on graph structures, potentially causing erroneous message propagation and biased clustering outcomes. To tackle these challenges, we introduce scSGC, a Soft Graph Clustering for single-cell RNA sequencing data, which aims to more accurately characterize continuous similarities among cells through non-binary edge weights, thereby mitigating the limitations of rigid data structures. The scSGC framework comprises three core components: (i) a zero-inflated negative binomial (ZINB)-based feature autoencoder; (ii) a dual-channel cut-informed soft graph embedding module; and (iii) an optimal transport-based clustering optimization module. Extensive experiments across ten datasets demonstrate that scSGC outperforms 13 state-of-the-art clustering models in clustering accuracy, cell type annotation, and computational efficiency. These results highlight its substantial potential to advance scRNA-seq data analysis and deepen our understanding of cellular heterogeneity.

cs.LG↗

Approximating Discrimination Within Models When Faced With Several Non-Binary Sensitive Attributes

Discrimination mitigation within machine learning (ML) models could be complicated because multiple factors may be interwoven hierarchically and historically. Yet few existing fairness measures can capture the discrimination level within ML models in the face of multiple sensitive attributes (SAs). To bridge this gap, we propose a fairness measure based on distances between sets from a manifold perspective, named as 'Harmonic Fairness measure via Manifolds (HFM)' with two optional versions, which can deal with a fine-grained discrimination evaluation for several SAs of multiple values. Because directly computing HFM may be costly, to accelerate its subprocedure -- the computation of distances of sets, we further propose two approximation algorithms named 'Approximation of distance between sets for one sensitive attribute with multiple values (ApproxDist)' and 'Approximation of extended distance between sets for several sensitive attributes with multiple values (ExtendDist)' to respectively resolve bias evaluation of one single SA with multiple values and that of several SAs with multiple values. Moreover, we provide an algorithmic effectiveness analysis for ApproxDist under certain assumptions to explain how well it could work. The empirical results demonstrate that our proposed fairness measure HFM is valid and approximation algorithms (i.e. ApproxDist and ExtendDist) are effective and efficient.

cs.LG↗

Deep Cut-informed Graph Embedding and Clustering

Graph clustering aims to divide the graph into different clusters. The recently emerging deep graph clustering approaches are largely built on graph neural networks (GNN). However, GNN is designed for general graph encoding and there is a common issue of representation collapse in existing GNN-based deep graph clustering algorithms. We attribute two main reasons for such issues: (i) the inductive bias of GNN models: GNNs tend to generate similar representations for proximal nodes. Since graphs often contain a non-negligible amount of inter-cluster links, the bias results in error message passing and leads to biased clustering; (ii) the clustering guided loss function: most traditional approaches strive to make all samples closer to pre-learned cluster centers, which causes a degenerate solution assigning all data points to a single label thus making all samples similar and less discriminative. To address these challenges, we investigate graph clustering from a graph cut perspective and propose an innovative and non-GNN-based Deep Cut-informed Graph embedding and Clustering framework, namely DCGC. This framework includes two modules: (i) cut-informed graph encoding; (ii) self-supervised graph clustering via optimal transport. For the encoding module, we derive a cut-informed graph embedding objective to fuse graph structure and attributes by minimizing their joint normalized cut. For the clustering module, we utilize the optimal transport theory to obtain the clustering assignments, which can balance the guidance of "proximity to the pre-learned cluster center". With the above two tailored designs, DCGC is more suitable for the graph clustering task, which can effectively alleviate the problem of representation collapse and achieve better performance. We conduct extensive experiments to demonstrate that our method is simple but effective compared with benchmarks.

cs.LG↗

Towards Trustworthy Federated Learning

This paper develops a comprehensive framework to address three critical trustworthy challenges in federated learning (FL): robustness against Byzantine attacks, fairness, and privacy preservation. To improve the system's defense against Byzantine attacks that send malicious information to bias the system's performance, we develop a Two-sided Norm Based Screening (TNBS) mechanism, which allows the central server to crop the gradients that have the l lowest norms and h highest norms. TNBS functions as a screening tool to filter out potential malicious participants whose gradients are far from the honest ones. To promote egalitarian fairness, we adopt the q-fair federated learning (q-FFL). Furthermore, we adopt a differential privacy-based scheme to prevent raw data at local clients from being inferred by curious parties. Convergence guarantees are provided for the proposed framework under different scenarios. Experimental results on real datasets demonstrate that the proposed framework effectively improves robustness and fairness while managing the trade-off between privacy and accuracy. This work appears to be the first study that experimentally and theoretically addresses fairness, privacy, and robustness in trustworthy FL.

cs.LG↗

Strategic priorities for transformative progress in advancing biology with proteomics and artificial intelligence

Artificial intelligence (AI) is transforming scientific research, including proteomics. Advances in mass spectrometry (MS)-based proteomics data quality, diversity, and scale, combined with groundbreaking AI techniques, are unlocking new challenges and opportunities in biological discovery. Here, we highlight key areas where AI is driving innovation, from data analysis to new biological insights. These include developing an AI-friendly ecosystem for proteomics data generation, sharing, and analysis; improving peptide and protein identification and quantification; characterizing protein-protein interactions and protein complexes; advancing spatial and perturbation proteomics; integrating multi-omics data; and ultimately enabling AI-empowered virtual cells.

q-bio.OT↗

$A_\infty$-Algebras from Lie Pairs

Given an inclusion $A\hookrightarrow L$ of Lie algebroids sharing the same base manifold $M$, i.e. a Lie pair, we prove that the space $Γ(Λ^\bullet A^\vee)\otimes_{R} \frac{U(L)}{U(L)\cdotΓ(A)}$, where $R=C^\infty(M)$, admits an $A_\infty$-algebra structure, unique up to $A_\infty$-isomorphisms. As a consequence, the Chevalley-Eilenberg cohomology $H^\bullet_{CE} \big( A, \frac{U(L)}{U(L)\cdotΓ(A)} \big)$ admits a canonical associative algebra structure. This $A_\infty$-algebra can be considered as the universal enveloping algebra of the $L_\infty$-algebroid $A[1]\times_M L/A$. Our construction is based on the homotopy equivalence of the $L_\infty$-algebroid $A[1]\times_M L/A$ and the dg Lie algebroid corresponding to the comma double Lie algebroid of Jotz-Mackenzie.

math.DG↗

Efficient Verification of Stabilizer Code Subspaces with Local Measurements

We address the task of verifying whether a quantum computer, designed to be protected by a specific stabilizer code, correctly encodes the corresponding logical qubits. To achieve this, we develop a general framework for subspace verification and explore several stabilizer code subspaces of practical significance. First, we present two efficient verification strategies for general stabilizer code subspaces, utilizing measurements of their stabilizer generators and stabilizer groups, respectively. Then, building on the observation that certain tests can be conducted in parallel when the subspace exhibits specific structural properties, we propose a coloring strategy tailored to graph code subspaces and an XZ strategy tailored to Calderbank-Shor-Steane (CSS) code subspaces. Compared to stabilizer-based strategies, these new strategies require significantly fewer measurement settings and consume fewer state copies, approaching near-global optimality. Notably, all the strategies employ a limited number of Pauli measurements, are non-adaptive, and work on mixed states, enabling efficient experimental certification of both logical qubits and logical operations in noisy quantum computers. This work contributes to the first systematic study of efficient verification of stabilizer code subspaces with local measurements.

quant-ph↗

Quantum memory assisted entangled state verification with local measurements

We consider the quantum memory assisted quantum state verification task, where an adversary prepare independent multipartite entangled states and send to the local verifiers, who then store several copies in the quantum memory and measure them collectively to make decision. We establish an exact analytic formula for optimizing two-copy state verification, where the verifiers store two copies, and give a globally optimal two-copy strategy for multi-qubit graph states involving only Bell measurements. When the verifiers can store arbitrarily many copies, we present a dimension expansion technique that designs efficient verification strategies for this case, showcasing its application to efficiently verifying GHZ-like states. These strategies become increasingly advantageous with growing memory resources, ultimately approaching the theoretical limit of efficiency. Our findings demonstrate that quantum memories enhance state verification efficiency, sheding light on error-resistant strategies and practical applications of large-scale quantum memory-assisted verification.

quant-ph↗

Long cycles and spectral radii in planar graphs

There is a rich history of studying the existence of cycles in planar graphs. The famous Tutte theorem on the Hamilton cycle states that every 4-connected planar graph contains a Hamilton cycle. Later on, Thomassen (1983), Thomas and Yu (1994) and Sanders (1996) respectively proved that every 4-connected planar graph contains a cycle of length $n-1, n-2$ and $n-3$. Chen, Fan and Yu (2004) further conjectured that every 4-connected planar graph contains a cycle of length $\ell$ for $\ell\in\{n,n-1,\ldots,n-25\}$ and they verified that $\ell\in \{n-4, n-5, n-6\}$. When we remove the ``4-connected" condition, how to guarantee the existence of a long cycle in a planar graph? A natural question asks by adding a spectral radius condition: What is the smallest constant $C$ such that for sufficiently large $n$, every graph $G$ of order $n$ with spectral radius greater than $C$ contains a long cycle in a planar graph? In this paper, we give a stronger answer to the above question. Let $G$ be a planar graph with order $n\geq 1.8\times 10^{17}$ and $k\leq \lfloor\log_2(n-3)\rfloor-8$ be a non-negative integer, we show that if $ρ(G)\geq ρ(K_2\vee(P_{n-2k-4}\cup 2P_{k+1}))$ then $G$ contains a cycle of length $\ell$ for every $\ell\in \{n-k, n-k-1, \ldots, 3\}$ unless $G\cong K_2\vee(P_{n-2k-4}\cup 2P_{k+1})$.

math.CO↗

Differential graded manifolds of finite positive amplitude

We prove that dg manifolds of finite positive amplitude, i.e. bundles of positively graded curved $L_\infty[1]$-algebras, form a category of fibrant objects. As a main step in the proof, we obtain a factorization theorem using path spaces. First we construct an infinite-dimensional factorization of a diagonal morphism using actual path spaces motivated by the AKSZ construction. Then we cut down to finite dimensions using the Fiorenza-Manetti method. The main ingredient in our method is the homotopy transfer theorem for curved $L_\infty[1]$-algebras. As an application, we study the derived intersections of manifolds.

math.DG↗