SearcharxivSearch

arXiv subjects

Chen Jia

Publications and source records attributed to Chen Jia.

At least 19 recordsLinked to original sources

Compass: Degradation-Simulated Reciprocal Learning with Lightweight Needle RWKV for Multimodal Crack Segmentation under Missing Modalities

In multimodal crack segmentation for industrial facilities, the key challenge is preventing missing modalities from degrading pixel-level performance while maintaining low computational cost. Existing methods struggle to address semantic degradation caused by missing modalities. We propose Compass, a lightweight network for robust crack segmentation under arbitrary missing modalities. Compass comprises Degradation Simulation Distillation (DSD), Needle Block, and Evidential Topology-Preserving Fusion (ETPF). DSD constructs a degradation simulation stream that mimics more severe missing conditions and performs reciprocal distillation with the original stream, decoupling complete perception from degradation adaptation. Within DSD, Feature-Aware Prototype Transmitter (FAPT) performs modality agnostic prototype-guided feature completion to maintain semantic integrity under incomplete modality conditions. As a lightweight backbone, Needle injects crack-direction cues into WKV modulation and combines connectivity-aware gating with anisotropic context probing for structure-aware modeling. ETPF fuses multimodal features via Dempster-Shafer evidential combination with uncertainty-gated decoding, preserving crack topology while suppressing unreliable features. Experiments on three datasets demonstrate state-of-the-art (SOTA) performance under diverse missing modality scenarios. Even with 90\% depth modality missing on CrackDepth, Compass achieves F1 of 0.8216 and mIoU of 0.8434 with only 2.58M parameters. The code is available at https://github.com/Karl1109/Compass.

cs.CV

MAC-XA: Multi-view Anatomy-Correspondence Fusion for Coronary Stenosis Reporting from X-ray Angiography

Multi-view reasoning in coronary X-ray angiography is inherently a cross-projection geometric problem, yet automated report generation in this setting remains largely unexplored. The 3D vascular topology leads to projection-dependent branch overlap and foreshortening, rendering single-view modeling fundamentally incomplete and unstable for lesion localization and stenosis grading. Although multi-view fusion appears promising, learning anatomically consistent fusion from real angiograms is impeded by a critical limitation: cross-view alignment is unobservable and cannot be explicitly supervised. Consequently, conventional fusion relies on implicit correlations rather than verified anatomical correspondence. We address this by reformulating multi-view stenosis reporting as an alignment-constrained aggregation problem. A controllable synthetic angiography generation strategy is introduced to expose geometry-derived patch-level correspondence supervision unavailable in real data. An anatomy-correspondence module learns cross-view correspondence matrices that explicitly align auxiliary features within the main-view coordinate space prior to fusion, thereby constraining evidence aggregation to anatomically consistent regions. Experiments on synthetic data and zero-shot transfer to real angiograms show that this alignment-constrained design improves correspondence consistency and structured stenosis reporting compared to single-view modeling and conventional multi-view fusion methods. The code will be publicly available upon publication.

cs.CV

Appearance-Preserving Refinement of Generated 3D Assets for Monochromatic Fabrication

Recent advances in 3D mesh generation have enabled the creation of visually realistic assets. However, much of their visual fidelity is encoded in textures rather than geometry. When such assets are fabricated using monochromatic materials, texture information is largely lost, causing visually important details to disappear even when the original geometry is faithfully preserved. A key challenge is that the geometric perturbations required to recover texture-dependent appearance cues often introduce sharp local features and high-frequency surface structures, which may increase stress concentration and fabrication risk. In this paper, we present GenMF, an appearance-oriented geometry refinement framework for monochromatic fabrication. GenMF transforms texture-dependent visual cues into geometry-induced shading effects and formulates geometry refinement as a balance between appearance preservation and fabrication-oriented robustness. To discourage structurally and narrow the gap between simulation and physical manufacturing, we further introduce a differentiable stress-aware regularization based on a learned thermal-stress predictor. Experimental results demonstrate that GenMF significantly improves appearance preservation under monochromatic rendering while reducing stress concentration under a consistent thermo-mechanical simulation setting. Physical 3D printing examples further show that the refined geometries preserve more recognizable visual details while remaining suitable for fabrication. These results suggest that appearance-aware geometry refinement provides an effective bridge between generated 3D assets and fabrication-ready monochromatic objects.

cs.GR

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding

Recent advances in Video Large Language Models (Video-LLMs) have enabled performance on long-video understanding tasks. However, existing methods still face two key limitations: evidence acquisition often relies on a single search intent, and answer generation lacks an effective visual feedback mechanism. To address these limitations, we propose \textbf{CoVER}, a Comprehensive Visual Evidence and Reflection framework for long-video understanding. CoVER enables Video-LLMs to \textbf{See More} by dynamically gathering query-expanded visual evidence, and \textbf{Think Deeper} by verifying draft answers with effective answer-specific visual feedback. Together, these mechanisms shift long-video understanding from answer-centric generation to evidence-centric and visually verifiable reasoning. Experimental results show that CoVER-7B substantially outperforms models with the same parameter scale and even surpasses state-of-the-art closed-source models on certain metrics.

cs.CV

SCRWKV: Ultra-Compact Structure-Calibrated Vision-RWKV for Topological Crack Segmentation

Achieving pixel-level accurate segmentation of structural cracks across diverse scenarios remains a formidable challenge. Existing methods face significant bottlenecks in balancing crack topology modeling with computational efficiency, often failing to reconcile high segmentation quality with low resource demands. To address these limitations, we propose the Ultra-Compact Structure-Calibrated Vision RWKV (SCRWKV), a network that achieves high-precision modeling via a novel Structure-Field Encoder (SFE) backbone while maintaining linear complexity. The SFE integrates the Adaptive Multi-scale Cascaded Modulator (AMCM) to enhance texture representation and utilizes the Structure-Calibrated Insight Unit (SCIU) as its core engine. Specifically, the SCIU employs the Geometry-guided Bidirectional Structure Transformation (GBST) to capture topological correlations and integrates the Dynamic Self-Calibrating Decay (DSCD) into Dy-WKV to suppress noise propagation. Furthermore, we introduce a lightweight Cross-Scale Harmonic Fusion (CSHF) decoder to achieve precise feature aggregation. Systematic evaluations on multiple benchmarks characterized by complex textures and severe interference demonstrate that SCRWKV, with only 1.22M parameters, significantly outperforms SOTA methods. Achieving an F1 score of 0.8428 and mIoU of 0.8512 on the TUT dataset, the model confirms its robust potential for efficient real-world deployment. The code is available at https://github.com/zhxhzy/SCRWKV.

cs.CV

High-bandwidth Coherence Cloning using Optical-Phase-Locking Feedforward

Ultra-narrow-linewidth lasers with suppressed high-frequency phase noise are critical for quantum control and precision metrology. While optical phase locking (OPL) is the standard technique for cloning the coherence of such sources, its effectiveness is often limited at high frequencies by feedback latency. We present a robust feedforward architecture that overcomes this limitation by recycling and demodulating the existing master-slave beat signal to drive a single electro-optic modulator for near-instantaneous noise cancellation. This approach eliminates the extraneous sidebands and transmission losses typical of more complex modulators. Through active stabilization of the beat amplitude and demodulation phase, we demonstrate robust suppression exceeding 30 dB from 10 kHz to 10 MHz. This hardware-efficient framework is readily compatible with standard OPL setups, offering a scalable solution for high-fidelity coherent control.

quant-ph

Distribution-informed Efficient Conformal Prediction for Full Ranking

Quantifying uncertainty is critical for the safe deployment of ranking models in real-world applications. Recent work offers a rigorous solution using conformal prediction in a full ranking scenario, which aims to construct prediction sets for the absolute ranks of test items based on the relative ranks of calibration items. However, relying on upper bounds of non-conformity scores renders the method overly conservative, resulting in substantially large prediction sets. To address this, we propose Distribution-informed Conformal Ranking (DCR), which produces efficient prediction sets by deriving the exact distribution of non-conformity scores. In particular, we find that the absolute ranks of calibration items follow Negative Hypergeometric distributions, conditional on their relative ranks. DCR thus uses the rank distribution to derive non-conformity score distribution and determine conformal thresholds. We provide theoretical guarantees that DCR achieves improved efficiency over the baseline while ensuring valid coverage under mild assumptions. Extensive experiments demonstrate the superiority of DCR, reducing average prediction set size by up to 36%, while maintaining valid coverage.

cs.LG

Probing False Vacuum Decay and Bubble Nucleation in a Rydberg Atom Array

In quantum field theory (QFT), the "vacuum" is not just empty space but the lowest-energy state of a quantum field. If the energy landscape has multiple local minima, the local ground states are the false vacuum (FV) which can tunnel towards the global ground state (true vacuum, TV). This process exhibits signature akin to classical supercooled gas transitions and many-body tunneling in discrete quantum systems. Here, we study the FV decay and bubble nucleation in a Rydberg atom ring. The $1/r^6$ van-der-Waals interactions and individual-site addressability allow us to explore physics beyond the standard Ising model. We observe that the FV decay rate decreases exponentially with the inverse of the symmetry-breaking field, directly mirroring QFT predictions. Moreover, we demonstrate that even minor deviations from the ideal metastable state can cause a stark departure from this universal scaling law. Extending beyond short-time decay dynamics, we also examine resonant bubble nucleation, a feature distinctive to systems with discrete energy spectra. Our findings and methods open avenues for future studies of many-body tunneling in higher dimensions or more complex geometries.

quant-ph

Bootstrapping LLMs via Preference-Based Policy Optimization

Bootstrapping large language models (LLMs) through preference-based policy optimization offers a promising direction for aligning model behavior with human preferences without relying on extensive manual annotations. In this work, we propose a novel preference-based policy optimization (PbPO) framework that formulates the learning process as a min-max game between the main policy and a reward model (RM). The RM is constrained within a confidence set derived from preference data to ensure reliable exploitation. Our iterative online algorithm actively collects preference data through guided exploration of the evolving policy, enabling continual self-improvement of both the policy and the RM. We provide theoretical guarantees for our method, establishing high-probability regret bounds for both settings with sequence-level RM and token-level RM, demonstrating its effectiveness in bootstrapping LLMs. Extensive experiments on five benchmarks show that our approach consistently outperforms existing state-of-the-art preference optimization techniques.

cs.AI

LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural Cracks

Achieving pixel-level segmentation with low computational cost using multimodal data remains a key challenge in crack segmentation tasks. Existing methods lack the capability for adaptive perception and efficient interactive fusion of cross-modal features. To address these challenges, we propose a Lightweight Adaptive Cue-Aware Vision Mamba network (LIDAR), which efficiently perceives and integrates morphological and textural cues from different modalities under multimodal crack scenarios, generating clear pixel-level crack segmentation maps. Specifically, LIDAR is composed of a Lightweight Adaptive Cue-Aware Visual State Space module (LacaVSS) and a Lightweight Dual Domain Dynamic Collaborative Fusion module (LD3CF). LacaVSS adaptively models crack cues through the proposed mask-guided Efficient Dynamic Guided Scanning Strategy (EDG-SS), while LD3CF leverages an Adaptive Frequency Domain Perceptron (AFDP) and a dual-pooling fusion strategy to effectively capture spatial and frequency-domain cues across modalities. Moreover, we design a Lightweight Dynamically Modulated Multi-Kernel convolution (LDMK) to perceive complex morphological structures with minimal computational overhead, replacing most convolutional operations in LIDAR. Experiments on three datasets demonstrate that our method outperforms other state-of-the-art (SOTA) methods. On the light-field depth dataset, our method achieves 0.8204 in F1 and 0.8465 in mIoU with only 5.35M parameters. Code and datasets are available at https://github.com/Karl1109/LIDAR-Mamba.

cs.CV

Online Knowledge Distillation with Reward Guidance

This work studies knowledge distillation (KD) for large language models (LLMs) through preference optimization. We propose a reward-guided imitation learning framework for sequential KD, formulating a min-max optimization problem between the policy and reward model (RM) to minimize the performance gap between the student and teacher policies. Specifically, the reward optimization is constrained to achieve near-optimality within a confidence set for preference alignment. For preference data construction, we explore both offline and online preference-based KD. Additionally, we reformulate the RM using the $Q$-value function and extend the framework to white-box KD, where the teacher policy's predicted probabilities are accessible. Theoretical analysis and empirical results demonstrate the effectiveness of the proposed framework.

cs.LG

Observation of average topological phase in disordered Rydberg atom array

Topological phases have been extensively studied over the past two decades, primarily in quantum pure states, where they are protected by exact symmetries. Recently, numerous studies have theoretically demonstrated the existence of average symmetry-protected topological (SPT) phases in mixed quantum states, which naturally arise in real systems due to decoherence or disorder. Despite extensive experimental observations of exact SPT phases in various systems, ranging from solid-state materials to synthetic matters, average SPT phases are yet to be observed until this work. Here we report direct observations of disorder-induced many-body interacting average SPT phase in an atom array at half-filling, whereby random offsets to tweezer locations forming a lattice implement structural disorder, resulting in fluctuating long-range dipolar interactions between tweezer confined single atoms. The induced topological phase is vindicated by the spatially resolved atom-atom correlation functions for different forms of dimer compositions. The ground state degeneracy in disordered configurations is detected and compared to the regular lattice without disorder. By probing the quench dynamics of a highly excited state, we observe markedly slower decay of edge spin magnetization in comparison to the bulk spin, consistent with the presence of topologically protected edge modes in disordered lattices.

cond-mat.quant-gas

SCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in Structures

Pixel-level segmentation of structural cracks across various scenarios remains a considerable challenge. Current methods encounter challenges in effectively modeling crack morphology and texture, facing challenges in balancing segmentation quality with low computational resource usage. To overcome these limitations, we propose a lightweight Structure-Aware Vision Mamba Network (SCSegamba), capable of generating high-quality pixel-level segmentation maps by leveraging both the morphological information and texture cues of crack pixels with minimal computational cost. Specifically, we developed a Structure-Aware Visual State Space module (SAVSS), which incorporates a lightweight Gated Bottleneck Convolution (GBC) and a Structure-Aware Scanning Strategy (SASS). The key insight of GBC lies in its effectiveness in modeling the morphological information of cracks, while the SASS enhances the perception of crack topology and texture by strengthening the continuity of semantic information between crack pixels. Experiments on crack benchmark datasets demonstrate that our method outperforms other state-of-the-art (SOTA) methods, achieving the highest performance with only 2.8M parameters. On the multi-scenario dataset, our method reached 0.8390 in F1 score and 0.8479 in mIoU. The code is available at https://github.com/Karl1109/SCSegamba.

cs.CV

RELexED: Retrieval-Enhanced Legal Summarization with Exemplar Diversity

This paper addresses the task of legal summarization, which involves distilling complex legal documents into concise, coherent summaries. Current approaches often struggle with content theme deviation and inconsistent writing styles due to their reliance solely on source documents. We propose RELexED, a retrieval-augmented framework that utilizes exemplar summaries along with the source document to guide the model. RELexED employs a two-stage exemplar selection strategy, leveraging a determinantal point process to balance the trade-off between similarity of exemplars to the query and diversity among exemplars, with scores computed via influence functions. Experimental results on two legal summarization datasets demonstrate that RELexED significantly outperforms models that do not utilize exemplars and those that rely solely on similarity-based exemplar selection.

cs.CL

Feedforward Cancellation of High-Frequency Phase Noise in Frequency-Doubled Lasers

The cancellation of high-frequency laser phase noise using feedforward techniques, as opposed to feedback methods, has achieved significant advancements in recent years. However, directly applying existing feedforward techniques to laser systems based on nonlinear conversion still faces substantial challenges. Here, we propose and demonstrate a feedforward scheme that suppresses phase noise in frequency-doubled light by utilizing phase noise information of its fundamental pump. This scheme is enabled by the fact that the phase jitter of the frequency-doubled light is simply twice that of the pump, except for a first-order low-pass filtering effect introduced by the SHG enhancement cavity. Testing this method on a 420-nm frequency-doubled laser system, we realize a 25-dB suppression of the servo noise bump near 1 MHz on the 420-nm light, and an average suppression of 30 dB for strong injected noise ranging from 100 kHz to 20 MHz. This scheme shows promising potential for applications requiring blue or ultraviolet light with minimal high-frequency phase noise, such as precision control of atoms and molecules.

physics.optics

Staircase Cascaded Fusion of Lightweight Local Pattern Recognition and Long-Range Dependencies for Structural Crack Segmentation

Accurately segmenting structural cracks at the pixel level remains a major hurdle, as existing methods fail to integrate local textures with pixel dependencies, often leading to fragmented and incomplete predictions. Moreover, their high parameter counts and substantial computational demands hinder practical deployment on resource-constrained edge devices. To address these challenges, we propose CrackSCF, a Lightweight Cascaded Fusion Crack Segmentation Network designed to achieve robust crack segmentation with exceptional computational efficiency. We design a lightweight convolutional block (LRDS) to replace all standard convolutions. This approach efficiently captures local patterns while operating with a minimal computational footprint. For a holistic perception of crack structures, a lightweight Long-range Dependency Extractor (LDE) captures global dependencies. These are then intelligently unified with local patterns by our Staircase Cascaded Fusion Module (SCFM), ensuring the final segmentation maps are both seamless in continuity and rich in fine-grained detail. To comprehensively evaluate our method, this paper created the challenging TUT benchmark dataset and evaluated it alongside five other public datasets. The experimental results show that the CrackSCF method consistently outperforms the existing methods, and it demonstrates greater robustness in dealing with complex background noise. On the TUT dataset, CrackSCF achieved 0.8382 on F1 score and 0.8473 on mIoU, and it only required 4.79M parameters.

cs.CV

Robust High-frequency Laser Phase Noise Suppression by Adaptive Pound-Drever-Hall Feedforward

Suppressing high-frequency laser phase noise, particularly at frequencies near and beyond typical feedback bandwidths of a few MHz, is a critical yet challenging task in many advanced applications. Feedforward-based methods generally outperform feedback in high-frequency range, but their performances are more susceptible to perturbations. In this work, we focus on the Pound-Drever-Hall (PDH)-feedforward method we demonstrated recently [Yu-Xin Chao et al., Optica 11(7), 945-950 (2024)] and analyze the factors that affect its long-term stability. By constructing a simple circuit allowing for adaptive control of the feedforward gain in response to power fluctuations of cavity transmission, we demonstrate a robust $\geq 40$ dB suppression of laser phase noise around 2 MHz and a noise suppression bandwidth up to 50 MHz. In comparison, when using normal PDH feedback, robust noise suppression of over 40 dB can only occur for frequencies below tens of kHz in most setups. Our findings may pave the way for general usage of PDH feedforward and allow for simple construction of low-noise lasers for precise quantum controls and precision metrology.

physics.optics

Parameter inference and nonequilibrium identification for Markovian systems based on coarse-grained observations

Most experiments can only detect a set of coarse-grained clusters of a molecular system, while the internal microstates are often inaccessible. Here, based on an infinitely long coarse-grained trajectory, we obtain a set of sufficient statistics which extracts all statistic information of coarse-grained observations. Based on these sufficient statistics, we set up a theoretical framework of parameter inference and nonequilibrium identification for a general Markovian system with an arbitrary number of microstates and arbitrary coarse-grained partitioning. Our framework can identify whether the sufficient statistics are enough for empirical estimation of all unknown parameters and we can also provide a quantitative criterion that reveals nonequilibrium. Our nonequilibrium criterion generalizes the one obtained [J. Chem. Phys. 132:041102 (2010)] for a three-state system with two coarse-grained clusters, and is capable of detecting a larger nonequilibrium region compared to the classical criterion based on autocorrelation functions.

cond-mat.stat-mech