Searcharxiv⌕ Search

arXiv subjects

Rui Li

Publications and source records attributed to Rui Li.

At least 55 records · Page 3Linked to original sources

Constructing edge-disjoint Steiner trees in Cartesian product networks

Cartesian product networks are always regarded as a tool for ``combining'' two given networks with established properties to obtain a new one that inherits properties from both. For a graph $F=(V,E)$ and a set $S\subseteq V(F)$ of at least two vertices, \emph{an $S$-Steiner tree} or \emph{a Steiner tree connecting $S$} (or simply, \emph{an $S$-tree}) is a subgraph $T=(V',E')$ of $F$ that is a tree with $S\subseteq V'$. For $S\subseteq V(F)$ and $|S|\geq 2$, the {\it generalized local edge-connectivity} $λ(S)$ is the maximum number of edge-disjoint Steiner trees connecting $S$ in $F$. For an integer $k$ with $2\leq k\leq n$, the {\it generalized $k$-edge-connectivity} $λ_k(F)$ of a graph $F$ is defined as $λ_k(F)=\min\{λ(S)\,|\,S\subseteq V(F) \ and \ |S|=k\}$.In this paper, we give sharp upper and lower bounds for $λ_k(G\Box H)$, where $\Box$ is the Cartesian product operation, and $G,H$ are two graphs.

math.CO↗

Large Scale Entanglement Structure Detection in 100-Qubit Systems via Local Joint Measurements

Identifying the entanglement structure of a many-body quantum state, namely how its constituents partition into unentangled blocks, is a central task in quantum information science, yet conventional tomography scales exponentially with system size. Here we introduce a scalable framework that recognizes large-scale entanglement structures directly from local correlation fingerprints. By choosing a representative local Pauli basis that satisfies a boundary-matching condition p_1 = p_R, the entire chain is read out in a single measurement configuration, keeping the measurement effort independent of system size. In noisy simulations, this single-basis protocol classifies GHZ-, W-, and cluster-type structures among 30 candidate partitions with a mean accuracy exceeding 95% for systems of up to 100 qubits. We further validate the protocol on a superconducting quantum processor, where it reliably classifies block structures for systems of up to 13 qubits before noise- and depth-induced degradation sets in at larger sizes. By mapping these failure modes explicitly, our results delineate the boundary of hardware-level scalability and point to a concrete strategy for characterizing entanglement structure on near-term quantum devices.

quant-ph↗

The Radioactive Background of the JUNO Calibration System

The Jiangmen Underground Neutrino Observatory (JUNO) experiment is a reactor antineutrino detector employing 20 kton of ultra-pure liquid scintillator to determine the neutrino mass ordering and to precisely measure oscillation parameters. The total singles background rate from radioactivity is required to be below 10 Hz in the energy range of 0.7-12 MeV within the fiducial volume for reactor neutrino analysis. The calibration system is designed to characterize the detector energy and position responses, while several of its components are located close to the target and may contribute to the background budget. Therefore, extensive material screening and selection are required to construct a low-background calibration system and to ensure that its contribution remains within the design requirements. In this work, a comprehensive study of the radioactive background induced by the calibration system is presented, including material radioactivity measurements using high-purity germanium detectors and neutron activation analysis techniques, detailed Monte Carlo simulations to evaluate the background, and comparisons with in-situ detector data to validate the predictions. In this data analysis, dedicated spatial selection methods are developed to isolate calibration-related contributions and to suppress the liquid scintillator background. The total radioactivity contribution from the calibration system is estimated to be less than 76 mHz, which satisfies the requirement of 200 mHz (2% of the total background budget). The results from in-situ data are found to be consistent with the expectations based on material assay and simulation within uncertainties. These results demonstrate that the calibration-induced background is well understood, in agreement between data and simulation, and negligible for reactor antineutrino measurements in JUNO.

physics.ins-det↗

AeroReformer2: Spoken-Query Referring Segmentation for Aerial Images

Spoken language offers a natural, hands-free interface for specifying an arbitrary target in dense remote-sensing imagery, yet existing referring remote-sensing image segmentation benchmarks accept only written expressions. To bridge this gap, we introduce \dataset, a spoken-query benchmark derived from RISBench that adds accent- and voice-diverse speech while preserving the original image, mask, and data splits. Its hard evaluation sets combine rotor, wind, and mixed interference with three signal-to-noise levels. We also propose \model, an efficient bilateral network that combines a boundary-preserving visual path with token-preserving speech encoding, kernel linear cross-modal attention, and a resolution refinement head. The design conditions visual features at two scales without materializing a dense speech--visual affinity matrix, then restores fine boundaries using high-resolution visual features. On the clean test split, \model with Swin-Base achieves 62.09\% mean intersection over union (mIoU) and 68.22\% overall intersection over union (oIoU), outperforming the strongest audio-adapted remote-sensing baseline by 5.38 and 2.08 percentage points, respectively. It retains the best hard-set mIoU at 54.09\%. To the best of our knowledge, this is the first benchmark and model study of full-sentence spoken-query referring segmentation for remote-sensing imagery. The code will be made publicly available.

cs.CV↗

Near-Unity Excitation and Radiative Efficiencies in Electroluminescence Without External Carrier Injection

Electroluminescence occurring without external charge injection is typically characterized by weak emission and excessive driving voltage, due to low excitation and radiative recombination efficiencies. Here, we demonstrate non-injecting electroluminescence (NI-EL) that challenges this conventional perception. To achieve this, we introduce an operational paradigm that leverages remote, state-abundant charge reservoirs, which elevates the excitation efficiency close to unity - a greater-than-20-fold improvement over the benchmark. This strategy is augmented by quantum dots (QDs) with nonmonotonically graded shells, raising the high-field radiative efficiency by approximately 7-8-fold. The resultant RGB NI-EL devices uniformly exhibit bright and efficient pulsed emissions, with key metrics including: a turn-on threshold of 3.7 Vrms for red; a luminance of 291,628 nits and a power efficiency of 302.6 lm/W for green, with light-outcoupling enhancement; and, for blue, the superior emitter stability of the first QD-based blue NI-EL over its light-emitting diode counterpart. The minimal dielectric loss, sub-100 ns response time, and external EL efficiency of up to 45.4% further reinforce the mechanism-performance causality. These results overcome the inherent mechanistic limitations of NI-EL and establish performance that rivals or surpasses injection-type EL, including AC- or DC-driven variants, positioning NI-EL as a promising platform for high-performance pulsed light sources.

physics.optics↗

GeoUniPR: A Geometry-Consistent Unified Framework for Cross-Modal Place Recognition

Cross-modal place recognition (CMPR) aims to identify the same location across heterogeneous sensing modalities, such as vision and LiDAR. Existing methods commonly bridge the modality gap using complex alignment modules, multi-stage training, or full fine-tuning of pretrained backbones. In this work, we revisit CMPR from the perspective of geometric consistency and propose GeoUniPR, a unified and concise geometry-consistent framework. GeoUniPR reduces cross-modal discrepancy at the representation level by projecting LiDAR point clouds into the camera perspective to construct Geometry-Consistent depth image views (DIV), which establish direct RGB-LiDAR correspondence. We further augment DIV with native LiDAR cues, including intensity and surface-normal information, yielding a multi-channel geometric representation that improves structural consistency. Based on this representation, GeoUniPR learns a unified embedding space using two modality-specific ViT-based encoders with identical architectures, trained through parameter-efficient adaptation without auxiliary alignment modules, multi-stage training, or full backbone fine-tuning. In addition, we introduce Spatially-Consistent InfoNCE (SC-InfoNCE), a CMPR-specific contrastive objective that suppresses distance-induced false negatives under spatial continuity. Extensive experiments on KITTI and KITTI-360 demonstrate that GeoUniPR achieves state-of-the-art (SOTA) performance in both same-modal and cross-modal place recognition, with strong cross-dataset generalization.

cs.CV↗

Levitated Milligram-scale Ferromagnetic Magnetometer at Room Temperature

Levitated mechanical oscillators are emerging ultrasensitive sensors with tremendous potential in both applied and fundamental physics. Levitated ferromagnets, with internal spin noises rapidly averaged, promise ultrahigh magnetic sensitivity. Here, we demonstrate a milligram-scale diamagnetically levitated ferromagnet system operating at room temperature. Through optimized geometry and multi-channel dissipation control, we achieve a magnetic sensitivity of 23~fT$/\sqrt{\text{Hz}}$ at frequency of 100-Hz level. We anticipate that a ferromagnetic magnetometer with subfemtotesla sensitivity is within reach, after modest technical improvements. This platform establishes a high-performance magnetometer for biomagnetic field detection and beyond-standard-model force searches.

quant-ph↗

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training

Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human preferences more efficient. However, existing latent reward models output only scalar scores. They do not estimate the uncertainty of each prediction. The generator therefore cannot determine which feedback is reliable. This can drive optimization in the wrong direction and lead to reward hacking. We propose \textsc{SURE}, a unified latent-space framework for image and video diffusion models. It learns reward distributions and directly uses their reliability to guide dense post-training. First, we propose sample-adaptive latent reward model (\textsc{SURE-LRM}). It predicts a Gaussian utility for each noisy latent. Its mean predicts the reward score. Its variance reflect the uncertainty of prediction without human annotation. The learned distribution then guides post-training through uncertainty-guided reward feedback learning (\textsc{SURE-REFL}). This method provides uncertainty-guided dense feedback along the denoising trajectory. At selected transitions, \textsc{SURE-REFL} queries the frozen \textsc{SURE-LRM}. It converts detached variance into reliability weights for samples at the same transition. Each weighted reward is backpropagated only through its local transition. The entire process remains in latent space and requires neither pixel-space decoding nor the full denoising graph. Experiments show that \textsc{SURE-LRM} improves preference prediction over strong baselines. \textsc{SURE-REFL} achieves the sota performance among various metrics and further improves optimization stability. It also achieves the highest VBench quality, semantic, and total scores among the evaluated methods.

cs.CV↗

Target-Aware Early Stage Ranking

Early Stage Ranking (ESR) in large-scale recommendation systems is dominated by ''user--item decoupling'' Two Tower architectures, which scale efficiently but cannot capture fine-grained, target-aware user--item interactions directly. We propose Target-Aware Early Stage Ranking (TESR), which augments the Two Tower with a Mixture of Attention (MoA) module trained as a request-level sequence modeling over user history. MoA combines (i) Hard Matching Attention (HMA) to capture explicit categorical-ID level overlap signals between user history and candidate item, (ii) target-aware HSTU attention for implicit affinities conditioned on the candidate, and (iii) target dependent and independent cross-attention for symmetric user-item contextualization. On top of this, a Multi-Logit Parameterized Gating (MLPG) head amplifies these signals at scoring time. To keep latency within ESR budgets, we co-design the architecture with FP8 quantization, custom kernels, and a Torch Inductor compilation path. On a production deployment, TESR delivers consistent offline NE wins and online topline gains, and is, to our knowledge, the first deployment of full target-aware attention sequence modeling in an ESR stage at this scale.

cs.LG↗

When Classes Evolve: A Benchmark and Framework for Stage-Aware Class-Incremental Learning

Class-Incremental Learning (CIL) aims to sequentially learn new classes while mitigating catastrophic forgetting of previously learned knowledge. Conventional CIL approaches implicitly assume that classes are morphologically static, focusing primarily on preserving previously learned representations as new classes are introduced. In practice, however, instances of the same semantic class may undergo substantial morphological evolution, such as a larva turning into a butterfly. Consequently, a model must both discriminate between classes and adapt to evolving appearances within a single class. To systematically address this challenge, we formalize Stage-Aware CIL (Stage-CIL), a paradigm in which each class is learned progressively through distinct morphological stages. We further introduce Stage-Bench, a 10-domain, two-stage benchmark and protocol for evaluating both inter-class forgetting and stage-level degradation within classes. Finally, we propose STAGE, an evolution-aware reference baseline that disentangles semantic identity from evolution dynamics through a fixed-size memory pool, enabling stage-aware prediction of later morphological forms from earlier representations. Extensive experiments show that conventional CIL reductions and existing continual-learning baselines remain insufficient under Stage-CIL, while STAGE consistently outperforms strong competitors, demonstrating the promise of explicit evolution-aware modeling for this new setting.

cs.LG↗

WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture

As scalability becomes increasingly important in recommendation modeling, recent architectures have advanced the modeling of two broad sources of ranking signals along separate paths: non-sequence features, including user, item, context, and cross features; and sequence features from user behavior histories. Wukong and HSTU have emerged as representative scalable backbones for these paths: Wukong scales high-order non-sequence feature-interaction modeling, while HSTU scales long user-behavior sequence modeling. Despite their complementary strengths, practical architectures that combine these two types of feature modeling remain underexplored. We present WHALE, a scalable unified recommendation architecture that jointly models non-sequence and sequence features on top of Wukong and HSTU. Each WHALE layer contains a Wukong module, an HSTU module, and an attention-based fusion module in which Wukong-derived interaction representations query HSTU-derived behavior representations. This design keeps both backbones active throughout the network and enables progressive Wukong-HSTU exchange, allowing high-order feature crosses to repeatedly retrieve fine-grained evidence from long user histories. To make WHALE practical for industrial deployment, we introduce customized Triton kernels and other model-systems co-design techniques to improve training and inference efficiency. On large-scale industrial recommendation data, WHALE achieves consistent gains in offline experiments. Additionally, it delivers positive online gains with a modest serving-throughput trade-off. The method has been deployed in production systems. Overall, WHALE provides a practical example of how these two sources of information can be scalably unified in an industrial recommendation model.

cs.IR↗

Long-range near-surface wake signatures of offshore wind farm clusters revealed by satellite observations

Wind energy has expanded rapidly in recent years, leading to increasingly dense clusters of offshore wind farms. As a result, wind farm wake effects, manifested as reduced wind speeds downstream of operating turbines, have become an important consideration for wind resource assessment and large-scale planning. Here we investigate near-surface wake signatures using wind speeds retrieved from 7122 Sentinel-1A/B synthetic aperture radar (SAR) images acquired between 2020 and 2022, covering more than 60 offshore wind farms across Europe and Asia. A consistent processing workflow and an automated wake-detection algorithm are applied to identify wake-affected regions and quantify wake-related wind speed changes at 10 m height. The results show that near-surface wake signatures can persist over distances exceeding 100 km under favorable conditions and that wake-affected regions exhibit an average wind speed reduction of 0.990 m/s (12.4%) at 10 m height. Several cases of wake interactions extending across national boundaries are also observed in densely developed offshore regions. These findings provide a large-scale, observation-based characterization of near-surface wind farm wake signatures.

physics.ao-ph↗

LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratory

The deployment of embodied agents in self-driving laboratories could accelerate scientific discovery, yet their reliability is constrained by the irreversible and safety-critical nature of chemical experiments. Progress is further hindered by scarce failure data and the lack of fine-grained evaluation protocols. To address these challenges, we introduce LabRobFail, a failure-centric framework for learning and evaluating robotic failure analysis in chemical laboratories. LabRobFail-Sim injects controllable failures at the control, physics, and semantic levels, enabling the construction of LabRobFail-Data, which contains over 20,000 trajectories across 70+ task scenarios, five failure categories, and 11 fine-grained failure types. LabRobFail-Bench evaluates six capabilities spanning task understanding, failure detection, temporal localization, severity assessment, failure classification, and actionable correction. We further develop LabRobFail-VLM, a domain-specialized vision-language model that generates structured failure diagnoses and recovery instructions. On seen environments, it achieves 90.83% failure-detection accuracy and 77.21% temporal-localization accuracy, substantially outperforming general-purpose VLMs. When integrated as a real-time supervisor, it improves downstream task success rates by 4-16 percentage points, demonstrating the value of fine-grained failure understanding for closed-loop recovery and reliable laboratory autonomy. Our code and data are available at https://github.com/Su-ISE-2001/SciRobo

cs.RO↗

Realizing Scaling Laws in Recommender Systems: A Foundation-Expert Paradigm for Hyperscale Model Deployment

Scaling laws have been established for recommender systems, yet efficiently deploying foundation model (FM) across multiple recommendation surfaces remains a major unsolved challenge. Existing methods for transfer learning face fundamental limitations in this setting: knowledge distillation suffers from transfer fidelity degradation in the large-data regime, and static user or item embeddings lack the expressiveness to capture contextualized user-item interactions. We propose the Foundation-Expert paradigm, where a central FM generates target-aware embeddings which are ingested by lightweight surface-specific expert models. Target-aware embeddings are representations that dynamically capture a user's interest in a specific item conditioned on their full interaction history. Unlike knowledge distillation, which transfers FM knowledge as soft labels, our approach provides these embeddings as input features to each expert model, enabling direct interaction with surface-specific representations. This paradigm achieves transfer ratios of 0.64--1.0 from FM to experts, substantially exceeding existing methods. Fully deployed at Meta serving tens of billions of daily requests since 2025, it delivers 0.050% statistically significant online topline metric improvement and 0.359% cumulative gains across multiple surfaces.

cs.IR↗

ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents

Reinforcement learning for multi-step LLM agents often relies on scalar rewards that indicate success but cannot explain why a trajectory is good or bad. Rubric-based rewards improve interpretability through natural-language criteria, but existing methods share two limitations: they score at the trajectory level, offering no guidance for individual steps; and their scorer is closed-source and static, so it cannot adapt as the agent evolves during training. We propose ARCO (Adaptive Rubric CO-evolution), which generates a per-step rubric and predicts a rubric-conditioned step-level reward for each action, and continually updates this rubric model on on-policy rollouts so that its criteria and scores co-evolve with the agent's improving behavior. Across HotpotQA, 2WikiMultiHopQA, and MuSiQue with two open-source backbones, ARCO achieves the highest EM in all settings over outcome-, rubric-, and process-reward baselines, and analyses show its rubrics are step-specific, robust to design choices, and useful for diagnosing agent behavior. Code and data are available at https://github.com/zihangtian/ARCO.

cs.AI↗

High-accuracy ultrasonic positioning of calibration sources in the Jiangmen Underground Neutrino Observatory

Precise source positioning is essential for detector calibration in large liquid scintillator detectors such as JUNO, particularly in regions where purely mechanical control is insufficient. An ultrasonic positioning system has been developed to reconstruct the three-dimensional coordinates of a calibration source without interfering with photon collection or contaminating the liquid scintillator. The method combines a sound-speed modeling based on dedicated laboratory measurements and in-detector temperature profiles, waveform-based arrival-time reconstruction, and an in-situ calibration of the effective receiver geometry using central-axis deployments. With six active receivers, central-axis positioning yields a mean error of 1.23 cm relative to the known deployment reference. For off-axis operation in the Cable Loop System calibration plane, a detector-realistic simulation that includes timing resolution, sound-speed variation, and receiver-coordinate smearing predicts a positioning uncertainty of 2.40 cm. These results demonstrate that ultrasonic positioning can provide centimetre-level source accuracy for large liquid scintillator detectors and can support off-axis calibration in JUNO-like experiments.

physics.ins-det↗

Point Ladder Tuning: Parameter-Efficient Hierarchical Adaptation for 3D Point Cloud Understanding

Fine-tuning pre-trained point-cloud backbones typically updates all parameters, resulting in substantial computation and memory overhead. More importantly, modern point backbones rely on aggressive tokenization and downsampling, which yields compact global tokens but irreversibly discards fine-grained local geometry, an inherent bottleneck for parameter-efficient adaptation. Consequently, existing PEFT methods that operate only on these coarsened tokens can modulate global semantics but struggle to recover the missing multi-scale locality. We present Point Ladder Tuning (PLT), a locality-aware PEFT framework that performs hierarchical, instance-conditioned adaptation while keeping the backbone frozen. PLT forms a lightweight closed loop: (i) a Hierarchical Ladder Network (HLN) constructs a multi-resolution local feature pyramid directly from raw points; (ii) a Local-Global Fusion (LGF) aligns and fuses local pyramids with intermediate backbone semantics; and (iii) a Dynamic Prompt Generator produces instance-aware multi-scale prompts to modulate the frozen backbone effectively. For dense prediction, we further introduce a lightweight segmentation head that progressively upsamples fused features and leverages backbone priors to refine fine structures. Extensive experiments on classification and dense prediction show that PLT consistently surpasses prior PEFT baselines with minimal tunable parameters. PLT achieves state-of-the-art performance using only 2.71% trainable parameters for classification and 7.69% for dense prediction, and scales favorably to larger backbones, requiring merely 0.36% parameters on PointGPT-L. The code is released at https://github.com/JunLinChang/ECCV2026-PLT.

cs.CV↗

Rethinking Reward Signals in Video GRPO: When Scores Become Targets

Group Relative Policy Optimization (GRPO) enables stable and preference-oriented updates via group-wise comparisons for post-training video generation. However, GRPO directly optimizes reward-induced advantages. Under sustained optimization, the reward score can lose fidelity as a proxy for true video quality, consistent with the phenomenon described by Goodhart's Law. This leads to two recurring issues: (i) shortcut-driven optimization under composite objectives and (ii) reward saturation within prompt groups. To address these issues, we introduce TaRoS, a Target-Robust Reward Signaling framework for Video generation GRPO. TaRoS leverages component level performance assessment together with intra-group sparsity to organize multi-aspect rewards towards optimization objectives. In addition, it adaptively downweights components that exhibit saturation, thereby preserving effective optimization directions and mitigating redundancy. This maintains meaningful optimization directions and preserves within-group ranking separation, thereby preventing reward hacking and leading to more reliable policy updates. Extensive experiments show consistent improvements in visual fidelity, motion coherence, and text-video alignment over strong baselines.

cs.CV↗