SearcharxivSearch

arXiv subjects

Jincheng Li

Publications and source records attributed to Jincheng Li.

At least 19 recordsLinked to original sources

CORE: Common Outcome Regularities from Action-Free Visual Demonstrations for Robot Manipulation

Robot imitation learning often relies on costly robot demonstrations, while abundant action-free visual demonstrations, such as human videos, are difficult to use because they lack robot-executable actions and suffer from embodiment gaps. We propose CORE, a policy learning framework that extracts Common Outcome Regularities (CORE) from visual demonstrations. Rather than transferring explicit actions across embodiments, CORE exploits a key observation: although successful trajectories for the same task can be diverse, their terminal states often share stable object configurations, spatial relations, and contact constraints. CORE first trains a terminal outcome encoder with contrastive and auxiliary temporal objectives, then aggregates successful terminal embeddings into visual goal prototypes, and finally injects these prototypes as global goal conditions into robot policies. Compared with language instructions, visual goal prototypes provide more concrete geometric and physical constraints for task completion. Across Meta-World, RoboTwin 2.0, and real-world manipulation, CORE improves the average success rate of the corresponding policy backbones by up to +3.9, +11.1, and +17.0 percentage points, respectively, and outperforms text-conditioned variants under the evaluated settings. The project and code are available at https://logssim.github.io/CORE.github.io/.

cs.RO

Two-Stage Cross-Domain Cervical Abnormality Screening with Cytopathological Image Synthesis and Knowledge Distillation

Cross-domain diagnosis remains a major challenge in cervical cell pathology due to pronounced domain shifts across institutions and the subtle visual differences among disease stages, which jointly impair model generalization. To address these issues, this paper proposes a two-stage framework for cross-domain cervical cell detection. In the first stage, we propose the Spatially-Continuous Unpaired Neural Schrödinger Bridge (SC-UNSB), which constructs a synthetic intermediate domain to mitigate cross-domain distribution shifts by modeling image translation as an entropy-regularized optimal transport process. In the second stage, we propose a dual-level feature alignment strategy within a knowledge distillation, which progressively aligns shallow structural features and deep semantic representations to facilitate the transfer of domain-invariant knowledge from the source to the target model. Experimental results demonstrate that the proposed method effectively mitigates domain shift and category ambiguity, improving the cross-domain detection performance.

cs.CV

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

Fine-grained vision-language understanding requires precise alignment between visual content and linguistic descriptions, a capability that remains limited in current models, particularly in non-English settings. While models like CLIP perform well on global alignment, they often struggle to capture fine-grained details in object attributes, spatial relations, and linguistic expressions, with limited support for bilingual comprehension. To address these challenges, we introduce FG-CLIP 2, a bilingual vision-language model designed to advance fine-grained alignment for both English and Chinese. Our approach leverages rich fine-grained supervision, including region-text matching and long-caption modeling, alongside multiple discriminative objectives. We further introduce the Textual Intra-modal Contrastive (TIC) loss to better distinguish semantically similar captions. Trained on a carefully curated mixture of large-scale English and Chinese data, including a newly released 12M Chinese region-text dataset, FG-CLIP 2 achieves powerful bilingual performance. To enable rigorous evaluation, we present a new benchmark for Chinese multimodal understanding, featuring long-caption retrieval and bounding box classification. Extensive experiments on 29 datasets across 8 tasks show that FG-CLIP 2 outperforms existing methods, achieving state-of-the-art results in both languages. We release the model, code, and benchmark to facilitate future research on bilingual fine-grained vision-language alignment.

cs.CV

AtPatch: Debugging Transformers via Hot-Fixing Over-Attention

Transformer-based deep neural networks (DNNs) affected by backdoor attacks and unfairness typically exhibit anomalous attention patterns, leading to over-attend to backdoor triggers or protected attributes. Existing neuron-editing mitigation strategies often struggle to handle such situation and most of them lack flexibility and tend to distort feature representations. Motivated by such over-attention phenomenon and software engineering paradigms such as delta debugging and hot patching, we propose AtPatch, a hot-fix method that dynamically redistributes attention maps during model inference. Specifically, for a given input, AtPatch first extracts the attention map from the model's inference process. Then, it uses a pre-trained detector to identify anomalous columns and replace them with unified benign attention. Then, AtPatch rescales other columns to mitigate the impact of over-attention. Finally, AtPatch returns the redistributed attention map to the model for continued inference. Notably, if the detector does not report any anomalous columns, AtPatch directly returns the original attention map to the model. Unlike existing techniques, AtPatch selectively redistributes the attention map, making it better at preserving the model's original functionality. Furthermore, AtPatch's on-the-fly nature allows it to work without modifying model parameters or retraining, making it better suited for deployed models. We conducted extensive experiments to validate AtPatch. Experimental results show that, compared to existing methods, AtPatch can more effectively mitigate backdoor attacks and unfairness while better preserving the model's original functionality.

cs.SE

LGAN: An Efficient High-Order Graph Neural Network via the Line Graph Aggregation

Graph Neural Networks (GNNs) have emerged as a dominant paradigm for graph classification. Specifically, most existing GNNs mainly rely on the message passing strategy between neighbor nodes, where the expressivity is limited by the 1-dimensional Weisfeiler-Lehman (1-WL) test. Although a number of k-WL-based GNNs have been proposed to overcome this limitation, their computational cost increases rapidly with k, significantly restricting the practical applicability. Moreover, since the k-WL models mainly operate on node tuples, these k-WL-based GNNs cannot retain fine-grained node- or edge-level semantics required by attribution methods (e.g., Integrated Gradients), leading to the less interpretable problem. To overcome the above shortcomings, in this paper, we propose a novel Line Graph Aggregation Network (LGAN), that constructs a line graph from the induced subgraph centered at each node to perform the higher-order aggregation. We theoretically prove that the LGAN not only possesses the greater expressive power than the 2-WL under injective aggregation assumptions, but also has lower time complexity. Empirical evaluations on benchmarks demonstrate that the LGAN outperforms state-of-the-art k-WL-based GNNs, while offering better interpretability.

cs.LG

Perfect continuous-variable quantum microcombs

Quantum microcombs generated in high-Q microresonators provide compact, multiplexed sources of entangled modes for continuous-variable (CV) quantum information processing. While deterministic generation of CV states via Kerr-induced two-mode squeezing has been demonstrated, achieving spectrally uniform squeezing remains challenging because of asymmetry and anomalies in the dispersion profile. Here we overcome these limitations by combining a microresonator with an engineered mode spectrum and optimized pump conditions. We realize a CV quantum microcomb comprising 14 independent two-mode squeezed states, each exhibiting more than 4 dB of raw squeezing (up to 4.3 dB) across a 0.7 THz bandwidth. This uniform, high-performance quantum resource represents a key step toward scalable, integrated CV quantum technologies operating beyond classical limits.

quant-ph

Conceptual Design of the Muonium-to-Antimuonium Conversion Experiment (MACE)

The spontaneous conversion of muonium to antimuonium is one of the interesting charged lepton flavor violation phenomena offering a sensitive probe of potential new physics and serving as a tool to constrain the parameter space beyond the Standard Model. The Muonium-to-Antimuonium Conversion Experiment (MACE) is designed to utilize a high-intensity muon beam, a Michel electron magnetic spectrometer, a positron transport system, and a positron detection system, to either discover or constrain this rare process with a conversion probability of $\mathcal{O}(10^{-13})$. This article presents an overview of the theoretical framework as well as a detailed description of the experimental design for the search for muonium-to-antimuonium conversion.

hep-ex

KAPG: Adaptive Password Guessing via Knowledge-Augmented Generation

As the primary mechanism of digital authentication, user-created passwords exhibit common patterns and regularities that can be learned from leaked datasets. Password choices are profoundly shaped by external factors, including social contexts, cultural trends, and popular vocabulary. Prevailing password guessing models primarily emphasize patterns derived from leaked passwords, while neglecting these external influences -- a limitation that hampers their adaptability to emerging password trends and erodes their effectiveness over time. To address these challenges, we propose KAPG, a knowledge-augmented password guessing framework that adaptively integrates external lexical knowledge into the guessing process. KAPG couples internal statistical knowledge learned from leaked passwords with external information that reflects real-world trends. By using password prefixes as anchors for knowledge lookup, it dynamically injects relevant external cues during generation while preserving the structural regularities of authentic passwords. Experiments on twelve leaked datasets show that KnowGuess achieves average improvements of 36.5\% and 74.7\% over state-of-the-art models in intra-site and cross-site scenarios, respectively. Further analyses of password overlap and model efficiency highlight its robustness and computational efficiency. To counter these attacks, we further develop KAPSM, a trend-aware and site-specific password strength meter. Experiments demonstrate that KAPSM significantly outperforms existing tools in accuracy across diverse evaluation settings.

cs.CR

High-Precision Mixed Feature Fusion Network Using Hypergraph Computation for Cervical Abnormal Cell Detection

Automatic detection of abnormal cervical cells from Thinprep Cytologic Test (TCT) images is a critical component in the development of intelligent computer-aided diagnostic systems. However, existing algorithms typically fail to effectively model the correlations of visual features, while these spatial correlation features actually contain critical diagnostic information. Furthermore, no detection algorithm has the ability to integrate inter-correlation features of cells with intra-discriminative features of cells, lacking a fusion strategy for the end-to-end detection model. In this work, we propose a hypergraph-based cell detection network that effectively fuses different types of features, combining spatial correlation features and deep discriminative features. Specifically, we use a Multi-level Fusion Sub-network (MLF-SNet) to enhance feature extractioncapabilities. Then we introduce a Cross-level Feature Fusion Strategy with Hypergraph Computation module (CLFFS-HC), to integrate mixed features. Finally, we conducted experiments on three publicly available datasets, and the results demonstrate that our method significantly improves the performance of cervical abnormal cell detection.

cs.CV

LMM-Det: Make Large Multimodal Models Excel in Object Detection

Large multimodal models (LMMs) have garnered wide-spread attention and interest within the artificial intelligence research and industrial communities, owing to their remarkable capability in multimodal understanding, reasoning, and in-context learning, among others. While LMMs have demonstrated promising results in tackling multimodal tasks like image captioning, visual question answering, and visual grounding, the object detection capabilities of LMMs exhibit a significant gap compared to specialist detectors. To bridge the gap, we depart from the conventional methods of integrating heavy detectors with LMMs and propose LMM-Det, a simple yet effective approach that leverages a Large Multimodal Model for vanilla object Detection without relying on specialized detection modules. Specifically, we conduct a comprehensive exploratory analysis when a large multimodal model meets with object detection, revealing that the recall rate degrades significantly compared with specialist detection models. To mitigate this, we propose to increase the recall rate by introducing data distribution adjustment and inference optimization tailored for object detection. We re-organize the instruction conversations to enhance the object detection capabilities of large multimodal models. We claim that a large multimodal model possesses detection capability without any extra detection modules. Extensive experiments support our claim and show the effectiveness of the versatile LMM-Det. The datasets, models, and codes are available at https://github.com/360CVGroup/LMM-Det.

cs.CV

Integrated optomechanical ultrasonic sensors with nano-Pascal-level sensitivity

Ultrasonic sensors are widely used for object detection and localization in underwater and biological settings. The operational range and spatial resolution are inherently limited by sensor sensitivity, in which conventional piezoelectric transducers have been overwhelmed by advanced photonic sensors. Here, we demonstrate an optomechanical ultrasonic sensor integrated into a photonic platform, which comprises a suspended SiO2 membrane embedded with a high-Q Si3N4 microring resonator. By exploiting simultaneous optical and mechanical resonances, the sensor achieves a record low noise-equivalent pressure (NEP) of 218 nPa/Hz^1/2 at 289 kHz in air and 9.6 nPa/Hz^1/2 at 52 kHz in water. We demonstrate its versatility through photoacoustic gas spectroscopy in air and underwater ultrasound imaging, achieving a minimum detectable C2H2 concentration of 2.9 ppm (integration time 1 s) and an imaging resolution of 1.89 mm, respectively. Our work represents a significant advancement in compact CMOS-compatible ultrasound sensing, unlocking new possibilities in biomedical imaging, environmental monitoring, industrial testing, and underwater communications.

physics.optics

FG-CLIP: Fine-Grained Visual and Textual Alignment

Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short captions. To address this, we propose Fine-Grained CLIP (FG-CLIP), which enhances fine-grained understanding through three key innovations. First, we leverage large multimodal models to generate 1.6 billion long caption-image pairs for capturing global-level semantic details. Second, a high-quality dataset is constructed with 12 million images and 40 million region-specific bounding boxes aligned with detailed captions to ensure precise, context-rich representations. Third, 10 million hard fine-grained negative samples are incorporated to improve the model's ability to distinguish subtle semantic differences. We construct a comprehensive dataset, termed FineHARD, by integrating high-quality region-specific annotations with hard fine-grained negative samples. Corresponding training methods are meticulously designed for these data. Extensive experiments demonstrate that FG-CLIP outperforms the original CLIP and other state-of-the-art methods across various downstream tasks, including fine-grained understanding, open-vocabulary object detection, image-text retrieval, and general multimodal benchmarks. These results highlight FG-CLIP's effectiveness in capturing fine-grained image details and improving overall model performance. The data, code, and models are available at https://github.com/360CVGroup/FG-CLIP.

cs.CV

High-Order Exceptional Point-Based Rotation Sensing in Anti-Parity time Symmetric Microresonators

Exceptional points (EPs), which arise from non-Hermitian systems, have been extensively investigated for the development of high-performance gyroscopes. However, operating a non-Hermitian gyroscope at high-order EP (HOEP) to achieve extreme performance requires strict and precise control of parameters. Here, we propose the design of an anti-parity-time (anti-PT) symmetric optical gyroscope operating at a fourth-order EP, achieving both ultra-sensitivity and high robustness. Our configuration exhibits eigenfrequency splitting two orders of magnitude higher than that of anti-PT gyroscopes operating at second-order EP. Furthermore, we demonstrate a significant reduction in angular random walk (ARW) under noise limits, compared to anti-parity symmetric gyroscopes based on second-order EP. Our results provide a novel approach for developing high-sensitivity rotation detection based on HOEPs.

physics.optics

Large-scale cluster quantum microcombs

An optical frequency comb comprises a cluster of equally spaced, phase-locked spectral lines. Replacing these classical components with correlated quantum light gives rise to cluster quantum frequency combs, providing abundant quantum resources for measurement-based quantum computation and multi-user quantum networks. We propose and generate cluster quantum microcombs within an on-chip optical microresonator driven by multi-frequency lasers. Through resonantly enhanced four-wave mixing processes, continuous-variable cluster states with 60 qumodes are deterministically created. The graph structures can be programmed into one- and two-dimensional lattices by adjusting the configurations of the pump lines, which are confirmed inseparable based on the measured covariance matrices. Our work demonstrates the largest-scale cluster states with unprecedented raw squeezing levels from a photonic chip, offering a compact and scalable platform for computational and communicational tasks with quantum advantages.

physics.optics

VM-UNet: Vision Mamba UNet for Medical Image Segmentation

In the realm of medical image segmentation, both CNN-based and Transformer-based models have been extensively explored. However, CNNs exhibit limitations in long-range modeling capabilities, whereas Transformers are hampered by their quadratic computational complexity. Recently, State Space Models (SSMs), exemplified by Mamba, have emerged as a promising approach. They not only excel in modeling long-range interactions but also maintain a linear computational complexity. In this paper, leveraging state space models, we propose a U-shape architecture model for medical image segmentation, named Vision Mamba UNet (VM-UNet). Specifically, the Visual State Space (VSS) block is introduced as the foundation block to capture extensive contextual information, and an asymmetrical encoder-decoder structure is constructed with fewer convolution layers to save calculation cost. We conduct comprehensive experiments on the ISIC17, ISIC18, and Synapse datasets, and the results indicate that VM-UNet performs competitively in medical image segmentation tasks. To our best knowledge, this is the first medical image segmentation model constructed based on the pure SSM-based model. We aim to establish a baseline and provide valuable insights for the future development of more efficient and effective SSM-based segmentation systems. Our code is available at https://github.com/JCruan519/VM-UNet.

eess.IV

Picotesla-sensitivity microcavity optomechanical magnetometry

Cavity optomechanical systems have enabled precision sensing of magnetic fields, by leveraging the optical resonance-enhanced readout and mechanical resonance-enhanced response. Previous studies have successfully achieved scalable and reproducible microcavity optomechanical magnetometry (MCOM) by incorporating Terfenol-D thin films into high-quality ($Q$) factor whispering gallery mode (WGM) microcavities. However, the sensitivity was limited to 585 pT/Hz$^{1/2}$, over 20 times inferior to those using Terfenol-D particles. In this work, we propose and demonstrate a high-sensitivity and scalable MCOM approach by sputtering a FeGaB thin film onto a high-$Q$ SiO$_2$ WGM microdisk. Theoretical studies are conducted to explore the magnetic actuation constant and noise-limited sensitivity by varying the parameters of the FeGaB film and SiO$_2$ microdisk. Multiple magnetometers with different radii are fabricated and characterized. By utilizing a microdisk with a radius of 355 $μ$m and a thickness of 1 $μ$m, along with a FeGaB film with a radius of 330 $μ$m and a thickness of 1.3 $μ$m, we have achieved a remarkable peak sensitivity of 1.68 pT/Hz$^{1/2}$ at 9.52 MHz. This represents a significant improvement of over two orders of magnitude compared with previous studies employing sputtered Terfenol-D film. Notably, the magnetometer operates without a bias magnetic field, thanks to the remarkable soft magnetic properties of the FeGaB film. Furthermore, as a proof-of-concept, we have demonstrated the real-time measurement of a pulsed magnetic field simulating the corona current in a high-voltage transmission line using our developed magnetometer. These high-sensitivity magnetometers hold great potential for various applications, such as magnetic induction tomography and corona current monitoring.

physics.optics

CCMB: A Large-scale Chinese Cross-modal Benchmark

Vision-language pre-training (VLP) on large-scale datasets has shown premier performance on various downstream tasks. In contrast to plenty of available benchmarks with English corpus, large-scale pre-training datasets and downstream datasets with Chinese corpus remain largely unexplored. In this work, we build a large-scale high-quality Chinese Cross-Modal Benchmark named CCMB for the research community, which contains the currently largest public pre-training dataset Zero and five human-annotated fine-tuning datasets for downstream tasks. Zero contains 250 million images paired with 750 million text descriptions, plus two of the five fine-tuning datasets are also currently the largest ones for Chinese cross-modal downstream tasks. Along with the CCMB, we also develop a VLP framework named R2D2, applying a pre-Ranking + Ranking strategy to learn powerful vision-language representations and a two-way distillation method (i.e., target-guided Distillation and feature-guided Distillation) to further enhance the learning capability. With the Zero and the R2D2 VLP framework, we achieve state-of-the-art performance on twelve downstream datasets from five broad categories of tasks including image-text retrieval, image-text matching, image caption, text-to-image generation, and zero-shot image classification. The datasets, models, and codes are available at https://github.com/yuxie11/R2D2

cs.CV

What Makes Good Open-Vocabulary Detector: A Disassembling Perspective

Open-vocabulary detection (OVD) is a new object detection paradigm, aiming to localize and recognize unseen objects defined by an unbounded vocabulary. This is challenging since traditional detectors can only learn from pre-defined categories and thus fail to detect and localize objects out of pre-defined vocabulary. To handle the challenge, OVD leverages pre-trained cross-modal VLM, such as CLIP, ALIGN, etc. Previous works mainly focus on the open vocabulary classification part, with less attention on the localization part. We argue that for a good OVD detector, both classification and localization should be parallelly studied for the novel object categories. We show in this work that improving localization as well as cross-modal classification complement each other, and compose a good OVD detector jointly. We analyze three families of OVD methods with different design emphases. We first propose a vanilla method,i.e., cropping a bounding box obtained by a localizer and resizing it into the CLIP. We next introduce another approach, which combines a standard two-stage object detector with CLIP. A two-stage object detector includes a visual backbone, a region proposal network (RPN), and a region of interest (RoI) head. We decouple RPN and ROI head (DRR) and use RoIAlign to extract meaningful features. In this case, it avoids resizing objects. To further accelerate the training time and reduce the model parameters, we couple RPN and ROI head (CRR) as the third approach. We conduct extensive experiments on these three types of approaches in different settings. On the OVD-COCO benchmark, DRR obtains the best performance and achieves 35.8 Novel AP$_{50}$, an absolute 2.8 gain over the previous state-of-the-art (SOTA). For OVD-LVIS, DRR surpasses the previous SOTA by 1.9 AP$_{50}$ in rare categories. We also provide an object detection dataset called PID and provide a baseline on PID.

cs.CV