SearcharxivSearch

arXiv subjects

Jian Jiang

Publications and source records attributed to Jian Jiang.

At least 19 recordsLinked to original sources

Electric-field control of hydrogen bonding via interfacial charge at atomic resolution

Hydrogen-bond networks govern molecular structure and function across chemistry, biology and materials science, yet their deterministic control at the atomic scale remains a central challenge (1-9).Here, we directly visualize how an external electric field enables reversible control of a hydrogen-bond network in monolayer ice on graphite through interfacial charge redistribution. Low-temperature scanning tunnelling microscopy reveals a field-driven transition from a mobile, physisorbed, non-wetting water phase to an ordered hexagonal monolayer, enabling deterministic nucleation, growth and complete wetting on an otherwise inert surface. Systematic variation of the field induces continuous lattice strain coexisting with discrete conductance states, revealing coupled structural and electronic responses. Reversal of the field polarity drives collective dipolar inversion, enabling switching between symmetry-equivalent configurations without disrupting the lattice. Supported by first-principles theory and bias-dependent imaging, these effects arise from field-induced modification of the interfacial electronic structure rather than purely geometric or orientational effects. These results establish interfacial charge redistribution as a general mechanism for electrically programming hydrogen-bond networks, providing a route to control molecular organization, electronic properties and collective dipolar order at interfaces.

cond-mat.mtrl-sci

A multi-ion optical clock with $\mathbf{5 \times 10^{-19}}$ uncertainty

Today's most accurate clocks are based on laser spectroscopy of electronic transitions in single trapped ions and feature fractional frequency uncertainties below $1\times10^{-18}$. Scaling these systems to multiple, simultaneously interrogated ions reduces measurement times, driving recent advances in multi-ion clocks. However, maintaining state-of-the-art systematic uncertainties while increasing the number of ions remains a central challenge. Here, we report on a multi-ion optical atomic clock with a fractional frequency uncertainty of $5.3\times10^{-19}$ and up to 10 \Sr ions. Ion-resolved state detection enables minimization of position-dependent shifts, with residual effects suppressed below the $10^{-20}$-level. Clock operation with eight to ten ions reduces the measurement time by a factor of 4.8 compared to single-ion operation. A comparison with an established \Yb single-ion clock yields an unperturbed frequency ratio of $0.6926711632159660405(20)$, with a statistical uncertainty of $0.9\times10^{-18}$ and a combined uncertainty of $2.9\times 10^{-18}$. These results demonstrate robust multi-ion clock operation with reduced averaging time and state-of-the-art accuracy.

physics.atom-ph

SCISSR: Scribble-Conditioned Interactive Surgical Segmentation and Refinement

Accurate segmentation of tissues and instruments in surgical scenes is annotation-intensive due to irregular shapes, thin structures, specularities, and frequent occlusions. While SAM models support point, box, and mask prompts, points are often too sparse and boxes too coarse to localize such challenging targets. We present SCISSR, a scribble-promptable framework for interactive surgical scene segmentation. It introduces a lightweight Scribble Encoder that converts freehand scribbles into dense prompt embeddings compatible with the mask decoder, enabling iterative refinement for a target object by drawing corrective strokes on error regions. Because all added modules (the Scribble Encoder, Spatial Gated Fusion, and LoRA adapters) interact with the backbone only through its standard embedding interfaces, the framework is not tied to a single model: we build on SAM 2 in this work, yet the same components transfer to other prompt-driven segmentation architectures such as SAM 3 without structural modification. To preserve pre-trained capabilities, we train only these lightweight additions while keeping the remaining backbone frozen. Experiments on EndoVis 2018 demonstrate strong in-domain performance, while evaluation on the out-of-distribution CholecSeg8k further confirms robustness across surgical domains. SCISSR achieves 95.41% Dice on EndoVis 2018 with five interaction rounds and 96.30% Dice on CholecSeg8k with three interaction rounds, outperforming iterative point prompting on both benchmarks.

eess.IV

Surg$\Sigma$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence

Surgical intelligence has the potential to improve the safety and consistency of surgical care, yet most existing surgical AI frameworks remain task-specific and struggle to generalize across procedures and institutions. Although multimodal foundation models, particularly multimodal large language models, have demonstrated strong cross-task capabilities across various medical domains, their advancement in surgery remains constrained by the lack of large-scale, systematically curated multimodal data. To address this challenge, we introduce Surg$\Sigma$, a spectrum of large-scale multimodal data and foundation models for surgical intelligence. At the core of this framework lies Surg$\Sigma$-DB, a large-scale multimodal data foundation designed to support diverse surgical tasks. Surg$\Sigma$-DB consolidates heterogeneous surgical data sources (including open-source datasets, curated in-house clinical collections and web-source data) into a unified schema, aiming to improve label consistency and data standardization across heterogeneous datasets. Surg$\Sigma$-DB spans 6 clinical specialties and diverse surgical types, providing rich image- and video-level annotations across 18 practical surgical tasks covering understanding, reasoning, planning, and generation, at an unprecedented scale (over 5.98M conversations). Beyond conventional multimodal conversations, Surg$\Sigma$-DB incorporates hierarchical reasoning annotations, providing richer semantic cues to support deeper contextual understanding in complex surgical scenarios. We further provide empirical evidence through recently developed surgical foundation models built upon Surg$\Sigma$-DB, illustrating the practical benefits of large-scale multimodal annotations, unified semantic design, and structured reasoning annotations for improving cross-task generalization and interpretability.

cs.AI

Surg-R1: A Hierarchical Reasoning Foundation Model for Scalable and Interpretable Surgical Decision Support with Multi-Center Clinical Validation

Surgical scene understanding demands not only accurate predictions but also interpretable reasoning that surgeons can verify against clinical expertise. However, existing surgical vision-language models generate predictions without reasoning chains, and general-purpose reasoning models fail on compositional surgical tasks without domain-specific knowledge. We present Surg-R1, a surgical Vision-Language Model that addresses this gap through hierarchical reasoning trained via a four-stage pipeline. Our approach introduces three key contributions: (1) a three-level reasoning hierarchy decomposing surgical interpretation into perceptual grounding, relational understanding, and contextual reasoning; (2) the largest surgical chain-of-thought dataset with 320,000 reasoning pairs; and (3) a four-stage training pipeline progressing from supervised fine-tuning to group relative policy optimization and iterative self-improvement. Evaluation on SurgBench, comprising six public benchmarks and six multi-center external validation datasets from five institutions, demonstrates that Surg-R1 achieves the highest Arena Score (64.9%) on public benchmarks versus Gemini 3.0 Pro (46.1%) and GPT-5.1 (37.9%), outperforming both proprietary reasoning models and specialized surgical VLMs on the majority of tasks spanning instrument localization, triplet recognition, phase recognition, action recognition, and critical view of safety assessment, with a 15.2 percentage point improvement over the strongest surgical baseline on external validation.

cs.CV

Innovative Tooth Segmentation Using Hierarchical Features and Bidirectional Sequence Modeling

Tooth image segmentation is a cornerstone of dental digitization. However, traditional image encoders relying on fixed-resolution feature maps often lead to discontinuous segmentation and poor discrimination between target regions and background, due to insufficient modeling of environmental and global context. Moreover, transformer-based self-attention introduces substantial computational overhead because of its quadratic complexity (O(n^2)), making it inefficient for high-resolution dental images. To address these challenges, we introduce a three-stage encoder with hierarchical feature representation to capture scale-adaptive information in dental images. By jointly leveraging low-level details and high-level semantics through cross-scale feature fusion, the model effectively preserves fine structural information while maintaining strong contextual awareness. Furthermore, a bidirectional sequence modeling strategy is incorporated to enhance global spatial context understanding without incurring high computational cost. We validate our method on two dental datasets, with experimental results demonstrating its superiority over existing approaches. On the OralVision dataset, our model achieves a 1.1% improvement in mean intersection over union (mIoU).

cs.CV

HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference

The linear memory growth of the KV cache poses a significant bottleneck for LLM inference in long-context tasks. Existing static compression methods often fail to preserve globally important information. Although recent dynamic retrieval approaches attempt to address this issue, they typically suffer from coarse-grained caching strategies and incur high I/O overhead. To overcome these limitations, we propose HeteroCache, a training-free dynamic compression framework. Our method is built on two key insights: attention heads exhibit diverse temporal heterogeneity, and there is significant spatial redundancy among heads within the same layer. Guided by these insights, HeteroCache categorizes heads based on stability and similarity, applying a fine-grained weighting strategy that allocates larger cache budgets to heads with rapidly shifting attention to capture context changes. Furthermore, it features a hierarchical storage mechanism where representative heads monitor attention drift to trigger asynchronous, on-demand context retrieval, thereby hiding I/O latency. Experiments demonstrate that HeteroCache achieves state-of-the-art performance on long-context benchmarks and accelerates decoding by up to $3\times$ compared to the original model with a 224K context. Our code is available at https://github.com/ponytaill/HeteroCache.

cs.CL

Systematic Evaluation and Guidelines for Segment Anything Model in Surgical Video Analysis

Surgical video segmentation is critical for AI to interpret spatial-temporal dynamics in surgery, yet model performance is constrained by limited annotated data. The SAM2 model, pretrained on natural videos, offers potential for zero-shot surgical segmentation, but its applicability in complex surgical environments, with challenges like tissue deformation and instrument variability, remains unexplored. We present the first comprehensive evaluation of the zero-shot capability of SAM2 in 9 surgical datasets (17 surgery types), covering laparoscopic, endoscopic, and robotic procedures. We analyze various prompting (points, boxes, mask) and {finetuning (dense, sparse) strategies}, robustness to surgical challenges, and generalization across procedures and anatomies. Key findings reveal that while SAM2 demonstrates notable zero-shot adaptability in structured scenarios (e.g., instrument segmentation, {multi-organ segmentation}, and scene segmentation), its performance varies under dynamic surgical conditions, highlighting gaps in handling temporal coherence and domain-specific artifacts. These results highlight future pathways to adaptive data-efficient solutions for the surgical data science field.

cs.CV

MolCluster: Integrating Graph Neural Network with Community Detection for Coarse-Grained Mapping

Coarse-grained (CG) modeling simplifies molecular systems by mapping groups of atoms into representative units. However, traditional CG approaches rely on fixed mapping rules, which limit their ability to handle diverse chemical systems and require extensive manual intervention. Thus, supervised learning-based CG methods have been proposed, enabling more automated and adaptable mapping. Nevertheless, these methods suffer from limited labeled datasets and the inability to control mapping resolution, which is essential for multiscale modeling. To overcome these limitations, we propose MolCluster, an unsupervised model that integrates a graph neural network and a community detection algorithm to extract CG representations. Additionally, a predefined group pair loss ensures the preservation of target groups, and a bisection strategy enables precise, customizable resolution across different molecular systems. In the case of the downstream task, evaluations on the MARTINI2 dataset demonstrate that MolCluster, benefiting from its label-free pretraining strategy, outperforms both traditional clustering and supervised models. Overall, these results highlight the potential of MolCluster as a core model for customizable and chemically consistent CG mapping.

physics.comp-ph

Meta-analysis and Topological Perturbation in Interactomic Network for Anti-opioid Addiction Drug Repurposing

The ongoing opioid crisis highlights the urgent need for novel therapeutic strategies that can be rapidly deployed. This study presents a novel approach to identify potential repurposable drugs for the treatment of opioid addiction, aiming to bridge the gap between transcriptomic data analysis and drug discovery. Speciffcally, we perform a meta-analysis of seven transcriptomic datasets related to opioid addiction by differential gene expression (DGE) analysis, and propose a novel multiscale topological differentiation to identify key genes from a protein-protein interaction (PPI) network derived from DEGs. This method uses persistent Laplacians to accurately single out important nodes within the PPI network through a multiscale manner to ensure high reliability. Subsequent functional validation by pathway enrichment and rigorous data curation yield 1,865 high-conffdence targets implicated in opioid addiction, which are cross-referenced with DrugBank to compile a repurposing candidate list. To evaluate drug-target interactions, we construct predictive models utilizing two natural language processing-derived molecular embeddings and a conventional molecular ffngerprint. Based on these models, we prioritize compounds with favorable binding afffnity proffles, and select candidates that are further assessed through molecular docking simulations to elucidate their receptor-level interactions. Additionally, pharmacokinetic and toxicological evaluations are performed via ADMET (absorption, distribution, metabolism, excretion, and toxicity) proffling, providing a multidimensional assessment of druggability and safety. This study offers a generalizable approach for drug repurposing in other complex diseases beyond opioid addiction. Keywords: Opioid addiction; Interactomic network; Topological perturbation; Differentially expressed gene; Drug repurposin

q-bio.MN

A review of topological data analysis and topological deep learning in molecular sciences

Topological Data Analysis (TDA) has emerged as a powerful framework for extracting robust, multiscale, and interpretable features from complex molecular data for artificial intelligence (AI) modeling and topological deep learning (TDL). This review provides a comprehensive overview of the development, methodologies, and applications of TDA in molecular sciences. We trace the evolution of TDA from early qualitative tools to advanced quantitative and predictive models, highlighting innovations such as persistent homology, persistent Laplacians, and topological machine learning. The paper explores TDA's transformative impact across diverse domains, including biomolecular stability, protein-ligand interactions, drug discovery, materials science, and viral evolution. Special attention is given to recent advances in integrating TDA with machine learning and AI, enabling breakthroughs in protein engineering, solubility and toxicity prediction, and the discovery of novel materials and therapeutics. We also discuss the limitations of current TDA approaches and outline future directions, including the integration of TDA with advanced AI models and the development of new topological invariants. This review aims to serve as a foundational reference for researchers seeking to harness the power of topology in molecular science.

q-bio.BM

Physical embedding machine learning force fields for organic systems

Machine learning force fields possess unprecedented potential in achieving both accuracy and efficiency in molecular simulations. Nevertheless, their application in organic systems is often hindered by structural collapse during simulation and significant deviations in the prediction of macroscopic properties. Here, two physics-embedded strategies are introduced to overcome these limitations. First, a physics-inspired self-adaptive bond-length sampling method achieves long-timescale stable simulations by requiring only several tens of single-molecule data sets, and has been validated across molecular systems, including engineering fluids, polypeptides, and pharmaceuticals. Second, a top-down intermolecular correction strategy based on a physical equation is introduced. This strategy requires only a small amount of simulation data and completes the optimization of tunable parameters within a few hours on a single RTX 4090 GPU, significantly reducing errors in density and viscosity, as validated in systems including ethylene carbonate, ethyl acetate, and dimethyl carbonate. Together, these approaches directly integrate physical insights into the machine learning models, thereby enhancing robustness and generalizability, and providing a scalable pathway for physics-embedded machine learning force fields.

physics.comp-ph

Privacy-enhancing Sclera Segmentation Benchmarking Competition: SSBC 2025

This paper presents a summary of the 2025 Sclera Segmentation Benchmarking Competition (SSBC), which focused on the development of privacy-preserving sclera-segmentation models trained using synthetically generated ocular images. The goal of the competition was to evaluate how well models trained on synthetic data perform in comparison to those trained on real-world datasets. The competition featured two tracks: $(i)$ one relying solely on synthetic data for model development, and $(ii)$ one combining/mixing synthetic with (a limited amount of) real-world data. A total of nine research groups submitted diverse segmentation models, employing a variety of architectural designs, including transformer-based solutions, lightweight models, and segmentation networks guided by generative frameworks. Experiments were conducted across three evaluation datasets containing both synthetic and real-world images, collected under diverse conditions. Results show that models trained entirely on synthetic data can achieve competitive performance, particularly when dedicated training strategies are employed, as evidenced by the top performing models that achieved $F_1$ scores of over $0.8$ in the synthetic data track. Moreover, performance gains in the mixed track were often driven more by methodological choices rather than by the inclusion of real data, highlighting the promise of synthetic data for privacy-aware biometric development. The code and data for the competition is available at: https://github.com/dariant/SSBC_2025.

cs.CV

UniCA: Unified Covariate Adaptation for Time Series Foundation Model

Time Series Foundation Models (TSFMs) have achieved remarkable success through large-scale pretraining. However, their design primarily targets real-valued series, limiting their ability to handle general forecasting tasks involving diverse and often heterogeneous covariates -- such as categorical variables and multimodal data (e.g., images, text) -- which are typically task-specific and difficult to leverage during pretraining. To address this gap, we propose Unified Covariate Adaptation (UniCA), a framework to bridge TSFMs with general covariate-aware forecasting. UniCA first performs covariate homogenization to transform heterogeneous covariates into high-level homogeneous series representations and then fuses them via a unified attention-based fusion mechanism. UniCA is compatible and universal for adaptation with both homogeneous and heterogeneous covariates, incorporating extra covariate information while preserving the generalization ability of TSFMs.Extensive experiments on multiple unimodal and multimodal covariate-aware forecasting benchmarks demonstrate the superiority of UniCA, highlighting the promise of covariate-aware TSFM adaptation in real-world forecasting scenarios.Code: https://github.com/hanlu-nju/UniCA.

cs.LG

SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

Foundation models have achieved transformative success across biomedical domains by enabling holistic understanding of multimodal data. However, their application in surgery remains underexplored. Surgical intelligence presents unique challenges - requiring surgical visual perception, temporal analysis, and reasoning. Existing general-purpose vision-language models fail to address these needs due to insufficient domain-specific supervision and the lack of a large-scale high-quality surgical database. To bridge this gap, we propose SurgVLM, one of the first large vision-language foundation models for surgical intelligence, where this single universal model can tackle versatile surgical tasks. To enable this, we construct a large-scale multimodal surgical database, SurgVLM-DB, comprising over 1.81 million frames with 7.79 million conversations, spanning more than 16 surgical types and 18 anatomical structures. We unify and reorganize 23 public datasets across 10 surgical tasks, followed by standardizing labels and doing hierarchical vision-language alignment to facilitate comprehensive coverage of gradually finer-grained surgical tasks, from visual perception, temporal analysis, to high-level reasoning. Building upon this comprehensive dataset, we propose SurgVLM, which is built upon Qwen2.5-VL, and undergoes instruction tuning to 10+ surgical tasks. We further construct a surgical multimodal benchmark, SurgVLM-Bench, for method evaluation. SurgVLM-Bench consists of 6 popular and widely-used datasets in surgical domain, covering several crucial downstream tasks. Based on SurgVLM-Bench, we evaluate the performance of our SurgVLM (3 SurgVLM variants: SurgVLM-7B, SurgVLM-32B, and SurgVLM-72B), and conduct comprehensive comparisons with 14 mainstream commercial VLMs (e.g., GPT-4o, Gemini 2.0 Flash, Qwen2.5-Max).

cs.CV

Structured Semantics from Unstructured Notes: Language Model Approaches to EHR-Based Decision Support

The advent of large language models (LLMs) has opened new avenues for analyzing complex, unstructured data, particularly within the medical domain. Electronic Health Records (EHRs) contain a wealth of information in various formats, including free text clinical notes, structured lab results, and diagnostic codes. This paper explores the application of advanced language models to leverage these diverse data sources for improved clinical decision support. We will discuss how text-based features, often overlooked in traditional high dimensional EHR analysis, can provide semantically rich representations and aid in harmonizing data across different institutions. Furthermore, we delve into the challenges and opportunities of incorporating medical codes and ensuring the generalizability and fairness of AI models in healthcare.

cs.IR

Machine learning predictions from unpredictable chaos

Chaos is omnipresent in nature, and its understanding provides enormous social and economic benefits. However, the unpredictability of chaotic systems is a textbook concept due to their sensitivity to initial conditions, aperiodic behavior, fractal dimensions, nonlinearity, and strange attractors. In this work, we introduce, for the first time, chaotic learning, a novel multiscale topological paradigm that enables accurate predictions from chaotic systems. We show that seemingly random and unpredictable chaotic dynamics counterintuitively offer unprecedented quantitative predictions. Specifically, we devise multiscale topological Laplacians to embed real-world data into a family of interactive chaotic dynamical systems, modulate their dynamical behaviors, and enable the accurate prediction of the input data. As a proof of concept, we consider 28 datasets from four categories of realistic problems: 10 brain waves, four benchmark protein datasets, 13 single-cell RNA sequencing datasets, and an image dataset, as well as two distinct chaotic dynamical systems, namely the Lorenz and Rossler attractors. We demonstrate chaotic learning predictions of the physical properties from chaos. Our new chaotic learning paradigm profoundly changes the textbook perception of chaos and bridges topology, chaos, and learning for the first time.

nlin.CD

Artificial Intelligence Approaches for Anti-Addiction Drug Discovery

Drug addiction is a complex and pervasive global challenge that continues to pose significant public health concerns. Traditional approaches to anti-addiction drug discovery have struggled to deliver effective therapeutics, facing high attrition rates, long development timelines, and inefficiencies in processing large-scale data. Artificial intelligence (AI) has emerged as a transformative solution to address these issues. Using advanced algorithms, AI is revolutionizing drug discovery by enhancing the speed and precision of key processes. This review explores the transformative role of AI in the pipeline for anti-addiction drug discovery, including data collection, target identification, and compound optimization. By highlighting the potential of AI to overcome traditional barriers, this review systematically examines how AI addresses critical gaps in anti-addiction research, emphasizing its potential to revolutionize drug discovery and development, overcome challenges, and advance more effective therapeutic strategies.

q-bio.BM