SearcharxivSearch

arXiv subjects

Gang Xu

Publications and source records attributed to Gang Xu.

At least 19 recordsLinked to original sources

Momentum-Space-Engineered Spatial Photonic Ising Machine for Long-Range Interactions

Long-range Ising models (LRIMs) with dense nonlocal and competing interactions are central to statistical physics, quantum simulation, and complex networks. Although the spatial photonic Ising machine (SPIM) exploits intrinsic optical parallelism for Ising computation, its ability to faithfully encode dense long-range couplings and capture the resulting thermodynamic signatures remains underexplored. Here, we present a momentum-space-engineered SPIM framework that maps prescribed long-range coupling kernels onto momentum-space masks for parallel Hamiltonian evaluation. Based on a high-fidelity optical field propagation model, the annealing dynamics of LRIMs with power-law and Ruderman-Kittel-Kasuya-Yosida (RKKY) interactions are systematically investigated. For the power-law model, we investigate the modulation of the estimated critical temperature by the decay exponent {\sigma} and coupling cutoff radius R. For the RKKY model, we reproduce diverse ordered states induced by complex competing long-range interactions. A proof-of-principle experiment demonstrates the physical feasibility of our approach. This framework broadens the class of many-body systems accessible to the SPIM platform.

quant-ph

CIR-DDG: backbone-agnostic residual correction of antibody-antigen affinity changes with explicit cross-chain geometry

Motivation: Accurate prediction of mutation-induced protein--protein binding free-energy changes is important for antibody affinity maturation, yet scarce labels and complex interface geometry limit generalization. Heterogeneous predictors may process three-dimensional complexes without preserving the cross-chain signals most relevant to a mutation in their final scalar output. Results: We introduce CIR-DDG, a lightweight residual adapter that combines a fixed base prediction with 22 interpretable descriptors of cross-chain distance, contact density and site--partner context. In complex-level five-fold evaluation on SKEMPI 2.0 measurements from 343 complexes, CIR-DDG improved all six tested backbones on antibody--antigen interface mutations: Spearman correlation increased by 0.0346--0.1296, while RMSE decreased by 0.0074--0.0408\kcalmol. Cross-validated probing, equal-capacity controls and feature ablations support the complementarity of explicit geometry. On an independent SARS-CoV-2 RBD--ACE2 deep-mutational-scan benchmark of 3669 substitutions, the fold-specific adapters transferred without any retraining: the absolute interface Spearman correlation increased by 0.026--0.081 for all four evaluable backbones, showing that the learned geometric correction generalizes beyond SKEMPI thermodynamic measurements. Availability and implementation: CIR-DDG is available at https://github.com/ecnuabmlab/CIR-ddG.

q-bio.QM

Robust Global Structure-from-Motion via View Graph Pruning

Structure-from-Motion (SfM) aims to estimate camera poses and reconstruct 3D structures from a collection of unordered images. Compared with incremental SfM, global SfM achieves better scalability by jointly estimating camera poses based on a view graph constructed from pairwise correspondences. However, its performance is highly sensitive to erroneous edges caused by visually ambiguous matches, which may lead to incorrect camera registration and reconstruction artifacts. In this work, we propose a subgraph-guided view graph pruning framework for robust global SfM. Our key idea is to exploit the internal consistency of reliable subgraphs to identify and remove unreliable connections. Specifically, we first partition the view graph into locally consistent subgraphs and perform global SfM within each subgraph to obtain reliable camera poses. We then apply RANSAC-based edge pruning across subgraphs to remove inconsistent edges, and finally perform global SfM on the refined view graph. Extensive experiments on ambiguous, sequential, and unordered image datasets demonstrate that our method improves the robustness of global SfM under challenging conditions. Further evaluation with neural rendering shows that the improved camera estimation leads to higher-quality novel view synthesis results.

cs.CV

Booster-based beam recycling for swap-out injection at the High Energy Photon Source

Fourth-generation synchrotron light sources employ ultralow-emittance storage rings with stringent injection requirements. On-axis swap-out injection alleviates the dependence on storage-ring dynamic aperture, but high-charge operation requires an efficient injector architecture capable of producing high-charge replacement bunches. This paper presents the accelerator physics design and performance analysis of a booster-based beam-recycling swap-out injection scheme implemented at the High Energy Photon Source (HEPS). In this approach, the full-energy booster serves as both an injector and a high-energy accumulator. An extracted storage-ring bunch is returned to the booster, merged with a low-charge bunch previously injected from the linac and accelerated to full energy. Following high-energy damping, the merged bunch is reinjected into the original storage-ring bucket. The scheme avoids the need for a dedicated accumulator ring while enabling high-charge bunch replacement. The recycling scheme was commissioned through staged machine studies. Full recycling-chain simulations, commissioning studies, and measured performance analysis are presented. The measured results characterize the recycling operation and quantify the transmission efficiency and performance limitations of the complete recycling loop. These results demonstrate the feasibility of the booster-based beam-recycling architecture and establish its operational basis for high-charge swap-out injection in future fourth-generation synchrotron light sources.

physics.acc-ph

WildShadowRemover: In-the-Wild Video Shadow Removal via Detail-Preserving Video Diffusion Models

Video shadow removal in the wild remains challenging due to complex illumination, diverse shadow appearances, and limited training data. Despite its importance to numerous vision and graphics applications, it remains largely unexplored in unconstrained real-world scenarios. To address this gap, we present WildShadowRemover, a framework that adapts a pretrained video diffusion model for robust video shadow removal via LoRA fine-tuning. To preserve fine image details while retaining the model's powerful generative prior, we augment the frozen VAE decoder with a detail injection module and introduce a shadow-mask-guided frequency-decomposed modulation module to selectively restore high-frequency textures while suppressing shadow artifacts. Monocular depth priors from Depth Anything 3 further provide geometry-aware guidance under challenging lighting conditions. We also construct WildShadow, a large-scale paired video shadow removal dataset and benchmark, covering diverse synthetic scenes. Extensive experiments demonstrate that our method outperforms existing approaches in shadow removal quality and temporal consistency, producing temporally coherent shadow-free videos with superior visual quality and strong generalization across challenging in-the-wild scenarios.

cs.CV

Robust Activation Map Rectification for Weakly Supervised Volumetric Segmentation: Temporal Coherence as a Free Lunch

Weakly supervised segmentation relies heavily on class activation maps (CAMs) to initially localize target regions. However, CAMs are often noisy and prone to catastrophic failures. Existing remedies typically introduce additional training stages or prototype learning, increasing computational cost and reducing robustness. In this paper, we propose a training-free prototype-free framework that rectifies unreliable CAMs by exploiting temporal and structural coherence in volumetric data as a free lunch. Our approach is built on two key components. First, we introduce Variance-Reduced Activation Aggregation (VRAA) which suppresses noise and amplify coherent semantic signals. We provide a theoretical justification by modeling CAMs as high-dimensional random vectors and show that aggregation yields provable variance reduction. Second, we design a Bidirectional Extremity Rectification (BER) mechanism that detects and rectifies implausible activations through bidirectional extremity checks, effectively mitigating extreme-value failures without learning additional parameters. Our method is model-agnostic and can be seamlessly integrated with existing pipelines. Extensive experiments on multiple public benchmarks demonstrate substantial improvements over state-of-the-art weakly supervised methods, achieving up to 20% Dice and 40% mIoU gains while reducing inference time by more than 5 times. These results indicate that leveraging coherence as an implicit inductive bias yields a principled and efficient approach to stabilizing weakly supervised volumetric segmentation. Our code will be available.

cs.CV

Multiband transport hierarchy and large Nernst effect in EuAuBi: Establishing a Nernst scaling for asymmetric multiband systems

In correlated materials, coexisting pockets of vastly different carrier densities raise two fundamental questions: which pocket governs the various transport coefficients, and does the conventional Nernst scaling $\nu/T \propto \mu/E_F$, originally derived for single-band systems, still hold? We address both questions in the polar semimetal EuAuBi, where a dilute electron pocket ($n_e \sim 10^{16}~\mathrm{cm}^{-3}$) coexists with a dense hole pocket ($n_h \sim 10^{21}~\mathrm{cm}^{-3}$). We find a clear hierarchy: the hole pocket dominates the longitudinal resistivity; the Hall effect crosses from electron- to hole-dominance with increasing field; the Seebeck coefficient is dominated by the electron pocket at low temperature and by both pockets at high temperature. Remarkably, the Nernst effect is governed entirely by the ultrahigh-mobility electron pocket, yielding a large low-field signal of $\sim 5~\mu\mathrm{V/K}$ near 1~T at 202~K, comparable to anomalous Nernst signals in magnetic Weyl semimetals. By analyzing the two-band thermoelectric conductivity, we show that the Nernst coefficient follows a scaling $\nu/T \propto \mu_e/{E_{F, tot}}$. This scaling originates from a compensation between the electron-to-hole conductivity ratio and the Fermi-energy ratio, establishing that the large Nernst effect is a semiclassical multiband phenomenon rather than a topological Berry-curvature contribution. This understanding advances the thermoelectric transport physics of multiband electronic systems and offers a guiding principle for low-field transverse thermoelectric design.

cond-mat.str-el

ORACLE: Anticipating Scams from Partial Trajectories in Streaming App Usage

Smartphone scams are increasingly prevalent and typically manifest as multi-stage, cross-application processes with gradually emerging intent. Effective intervention thus requires anticipating scams before the intent becomes explicit. This is inherently challenging, as decisions must rely on partial trajectories with temporally distributed evidence. In this paper, we propose \textbf{ORACLE} Online Reasoning for Anticipating Cross-temporal Latent thrEats, the first agentic framework for early scam anticipation from \textit{streaming app-usage} trajectories. To support this setting, we curate a real-world long-horizon benchmark of streaming app-usage trajectories, covering 12 scam types, spanning extended periods (15 days on average), involving diverse applications (95 apps), and interleaving normal and scam behaviors. To address fragmented evidence, we introduce a self-evolving context manager that adaptively consolidates entity-centric interactions over time, enabling more effective reconstruction of cross-temporal evidence from partial observations. To enhance sensitivity to latent early-stage signals, we propose an on-policy self-distillation scheme in which a teacher model, conditioned on summarized anti-scam reflections and clues by skills, supervises a student model without access to such reflections. This scheme thereby distills evidence-informed knowledge and improves recognition of emerging fraud patterns from partial trajectories. Experiments show that \method{} consistently improves early scam anticipation, yielding timely warnings while reducing false alerts in realistic streaming scenarios.

cs.LG

RS-WorldModel: a Unified Model for Remote Sensing Understanding and Future Sense Forecasting

Remote sensing world models aim to both explain observed changes and forecast plausible futures, two tasks that share spatiotemporal priors. Existing methods, however, typically address them separately, limiting cross-task transfer. We present RS-WorldModel, a unified world model for remote sensing that jointly handles spatiotemporal change understanding and text-guided future scene forecasting, and we build RSWBench-1.1M, a 1.1 million sample dataset with rich language annotations covering both tasks. RS-WorldModel is trained in three stages: (1) Geo-Aware Generative Pre-training (GAGP) conditions forecasting on geographic and acquisition metadata; (2) synergistic instruction tuning (SIT) jointly trains understanding and forecasting; (3) verifiable reinforcement optimization (VRO) refines outputs with verifiable, task-specific rewards. With only 2B parameters, RS-WorldModel surpasses open-source models up to 120$ \times $ larger on most spatiotemporal change question-answering metrics. It achieves an FID of 43.13 on text-guided future scene forecasting, outperforming all open-source baselines as well as the closed-source Gemini-2.5-Flash Image (Nano Banana).

cs.AI

Symmetry-Broken Cavity Solitons and Collective Polarization Conformity in Fabry-Perot Kerr Resonators

We report on the experimental generation of polarization symmetry-broken cavity solitons (CSs) in a passive, fiber-based, coherently-driven, Fabry-Perot (FP) Kerr resonator. Polarization resolved measurements reveal the spontaneous transition of initially symmetric CSs into asymmetrical vectorial states, triggered by a cross-phase modulation-induced polarization bifurcation. Most notably, due to counter-propagation of light occurring in FP resonators, we unveil a collective polarization conformity effect, whereby multiple CSs circulating in the cavity converge to the same asymmetric polarization state once their number exceeds a certain threshold. These results demonstrate that Fabry-Perot resonators support novel collective soliton dynamics that are absent in ring architectures.

physics.optics

UniE2F: A Unified Diffusion Framework for Event-to-Frame Reconstruction with Video Foundation Models

Event cameras excel at high-speed, low-power, and high-dynamic-range scene perception. However, as they fundamentally record only relative intensity changes rather than absolute intensity, the resulting data streams suffer from a significant loss of spatial information and static texture details. In this paper, we address this limitation by leveraging the generative prior of a pre-trained video diffusion model to reconstruct high-fidelity video frames from sparse event data. Specifically, we first establish a baseline model by directly applying event data as a condition to synthesize videos. Then, based on the physical correlation between the event stream and video frames, we further introduce the event-based inter-frame residual guidance to enhance the accuracy of video frame reconstruction. Furthermore, we extend our method to video frame interpolation and prediction in a zero-shot manner by modulating the reverse diffusion sampling process, thereby creating a unified event-to-frame reconstruction framework. Experimental results on real-world and synthetic datasets demonstrate that our method significantly outperforms previous approaches both quantitatively and qualitatively. We also refer the reviewers to the video demo contained in the supplementary material for video results. The code will be publicly available at https://github.com/CS-GangXu/UniE2F.

cs.CV

Uniaxial stress enhanced anisotropic magnetoresistance and superconductivity in the kagome superconductor LaRu$_{3}$Si$_{2}$

Elucidating the role of the kagome electronic structure in determining the various quantum ground states is of fundamental importance. In this work, we employ in-plane uniaxial stress as a tuning parameter to probe the electronic structure and its impact on the superconducting and normal-state properties of the kagome superconductor LaRu$_{3}$Si$_{2}$, combining magnetotransport measurements with first-principles calculations. We identify a pronounced anisotropy in both the upper critical field and the normal-state magnetoresistance, indicating strong electronic anisotropy despite the three-dimensional crystal structure. Furthermore, we find that the superconducting transition temperature $T_{\rm c}$ increases under in-plane stress applied within the kagome plane, although the enhancement is modest, reaching approximately 0.3 K at 0.6 GPa. Furthermore, the absolute magnetoresistance exhibits a pronounced increase from about 22${\%}$ at zero stress to 35${\%}$ at 0.6 GPa, indicating a substantial modification of the normal state above $T_{\rm c}$. Previous studies have reported time-reversal-symmetry (TRS) breaking below a temperature scale that coincides with the onset of magnetoresistance. The simultaneous enhancement of both $T_{\rm c}$ and magnetoresistance under stress therefore suggests a positive correlation between superconductivity and normal-state electronic and magnetic properties in LaRu$_{3}$Si$_{2}$. Detailed calculations demonstrate that stress-induced changes in $T_{\rm c}$ arise from the joint evolution of the total density of states and the flat band, whereas the large magnetoresistance enhancement is dominated by the stress-driven downward shift of the Ru $dz^{2}$ kagome flat band.

cond-mat.supr-con

Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility

Vision language models (VLMs) extend the reasoning capabilities of large language models (LLMs) to cross-modal settings, yet remain highly vulnerable to multimodal jailbreak attacks. Existing defenses predominantly rely on safety fine-tuning or aggressive token manipulations, incurring substantial training costs or significantly degrading utility. Recent research shows that LLMs inherently recognize unsafe content in text, and the incorporation of visual inputs in VLMs frequently dilutes risk-related signals. Motivated by this, we propose Risk Awareness Injection (RAI), a lightweight and training-free framework for safety calibration that restores LLM-like risk recognition by amplifying unsafe signals in VLMs. Specifically, RAI constructs an Unsafe Prototype Subspace from language embeddings and performs targeted modulation on selected high-risk visual tokens, explicitly activating safety-critical signals within the cross-modal feature space. This modulation restores the model's LLM-like ability to detect unsafe content from visual inputs, while preserving the semantic integrity of original tokens for cross-modal reasoning. Extensive experiments across multiple jailbreak and utility benchmarks demonstrate that RAI substantially reduces attack success rate without compromising task performance.

cs.AI

LongCat-Flash-Thinking-2601 Technical Report

We introduce LongCat-Flash-Thinking-2601, a 560-billion-parameter open-source Mixture-of-Experts (MoE) reasoning model with superior agentic reasoning capability. LongCat-Flash-Thinking-2601 achieves state-of-the-art performance among open-source models on a wide range of agentic benchmarks, including agentic search, agentic tool use, and tool-integrated reasoning. Beyond benchmark performance, the model demonstrates strong generalization to complex tool interactions and robust behavior under noisy real-world environments. Its advanced capability stems from a unified training framework that combines domain-parallel expert training with subsequent fusion, together with an end-to-end co-design of data construction, environments, algorithms, and infrastructure spanning from pre-training to post-training. In particular, the model's strong generalization capability in complex tool-use are driven by our in-depth exploration of environment scaling and principled task construction. To optimize long-tailed, skewed generation and multi-turn agentic interactions, and to enable stable training across over 10,000 environments spanning more than 20 domains, we systematically extend our asynchronous reinforcement learning framework, DORA, for stable and efficient large-scale multi-environment training. Furthermore, recognizing that real-world tasks are inherently noisy, we conduct a systematic analysis and decomposition of real-world noise patterns, and design targeted training procedures to explicitly incorporate such imperfections into the training process, resulting in improved robustness for real-world applications. To further enhance performance on complex reasoning tasks, we introduce a Heavy Thinking mode that enables effective test-time scaling by jointly expanding reasoning depth and width through intensive parallel thinking.

cs.AI

Unexpected type-II multiferroic phase in GdMnO3 under high magnetic fields

Perovskite manganites with small A-site ions, as the first and canonical branch of type-II multiferroics, are ideal systems to exhibit magnetism-induced ferroelectricity. Despite their established magnetoelectric phase diagrams under low magnetic fields, here an unidentified phase with a large magnetism-induced polarization (up to 1500 {\mu}C/m2) is revealed in GdMnO3 under high magnetic fields up to 60 T. Based on multiprobe experiments, a complete phase diagram is constructed with successive polar-nonpolar-polar-nonpolar transitions. Such a nonmonotonic evolution is well mimicked by model simulation, while the spin-lattice coupling is the key ingredient for the reentrant ferroelectric phase.

cond-mat.mtrl-sci

Multi-peak vector soliton families in defocusing Kerr resonators

We report the existence of multi-peaked vector soliton families in normally dispersive passive Kerr resonators. Through cross-phase modulation between two orthogonal polarization components, each peak becomes tightly interlocked, enabling robust localization of the entire wave packet in defocusing cavities. Analysis using snakes-and-ladder diagrams demonstrates the diversity of these vector soliton families, which include dark-bright multi-peak solitons, flat-topped solitons, and modulation instability patterns, among others. Furthermore, stability analysis based on the coupled Lugiato-Lefever equations reveals that specific combinations of parameters can sustain stable vector cavity solitons, whose peak numbers can be continuously tuned by adding appropriate perturbations. These findings significantly expand the scope of soliton dynamics and optical frequency comb generation in pumped-dissipative systems, independent of dispersion conditions.

physics.optics

Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion

Realistic 3D city generation is fundamental to a wide range of applications, including virtual reality and digital twins. However, most existing methods rely on training a single diffusion model, which limits their ability to generate personalized and boundless city-scale scenes. In this paper, we present Yo'City, a novel agentic framework that enables user-customized and infinitely expandable 3D city generation by leveraging the reasoning and compositional capabilities of off-the-shelf large models. Specifically, Yo'City first conceptualizes the city through a top-down planning strategy that defines a hierarchical "City-District-Grid" structure. The Global Planner determines the overall layout and potential functional districts, while the Local Designer further refines each district with detailed grid-level descriptions. Subsequently, the grid-level 3D generation is achieved through a "produce-refine-evaluate" isometric image synthesis loop, followed by image-to-3D generation. To simulate continuous city evolution, Yo'City further introduces a user-interactive, relationship-guided expansion mechanism, which performs scene graph-based distance- and semantics-aware layout optimization, ensuring spatially coherent city growth. To comprehensively evaluate our method, we construct a diverse benchmark dataset and design six multi-dimensional metrics that assess generation quality from the perspectives of semantics, geometry, texture, and layout. Extensive experiments demonstrate that Yo'City consistently outperforms existing state-of-the-art methods across all evaluation aspects.

cs.CV

LongCat-Flash-Omni Technical Report

We introduce LongCat-Flash-Omni, a state-of-the-art open-source omni-modal model with 560 billion parameters, excelling at real-time audio-visual interaction. By adopting a curriculum-inspired progressive training strategy that transitions from simpler to increasingly complex modality sequence modeling tasks, LongCat-Flash-Omni attains comprehensive multimodal capabilities while maintaining strong unimodal capability. Building upon LongCat-Flash, which adopts a high-performance Shortcut-connected Mixture-of-Experts (MoE) architecture with zero-computation experts, LongCat-Flash-Omni integrates efficient multimodal perception and speech reconstruction modules. Despite its immense size of 560B parameters (with 27B activated), LongCat-Flash-Omni achieves low-latency real-time audio-visual interaction. For training infrastructure, we developed a modality-decoupled parallelism scheme specifically designed to manage the data and model heterogeneity inherent in large-scale multimodal training. This innovative approach demonstrates exceptional efficiency by sustaining over 90% of the throughput achieved by text-only training. Extensive evaluations show that LongCat-Flash-Omni achieves state-of-the-art performance on omni-modal benchmarks among open-source models. Furthermore, it delivers highly competitive results across a wide range of modality-specific tasks, including text, image, and video understanding, as well as audio understanding and generation. We provide a comprehensive overview of the model architecture design, training procedures, and data strategies, and open-source the model to foster future research and development in the community.

cs.MM