Searcharxiv⌕ Search

arXiv subjects

Jun Luo

Publications and source records attributed to Jun Luo.

At least 37 records · Page 2Linked to original sources

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data

Recent advances in language models have established reinforcement learning as the primary paradigm for eliciting self-correction and long-chain reasoning. While group relative policy optimization (GRPO) offers superior scalability by eliminating the critic network, deploying it on a central infrastructure entails collecting a large volume of data from distributed owners, which poses significant privacy risks. To address these concerns, we introduce federated GRPO (FGRPO), a framework designed to decentralize the fine-tuning of reasoning models across heterogeneous data owners. To effectively mitigate the instability caused by divergent reward scales across heterogeneous tasks, FGRPO incorporates an adaptive aggregation mechanism based on relative performance gain. By characterizing each client's improvement relative to its personalized historical baseline, the framework dynamically prioritizes effective learning trajectories regardless of local task difficulty. FGRPO ensures robust convergence on non-IID data while preserving data privacy.

cs.LG↗

DECA: Decentralizing Block-Wise Adam for Efficient LLM Full-Parameter Fine-Tuning on Non-IID Data

Fine-tuning large language models (LLMs) in privacy-sensitive and resource-constrained environments remains challenging. Since training data are often distributed across multiple clients, decentralized fine-tuning offers a natural paradigm for collaborative adaptation without a central server. However, enabling full-parameter fine-tuning (FPFT) in this decentralized setting is difficult: FPFT provides strong adaptation capacity but incurs prohibitive resource consumption for billion-scale models. Existing decentralized LLM fine-tuning methods therefore mainly rely on parameter-efficient updates, which improve efficiency but may restrict downstream performance. Moreover, client data are typically non-IID, making decentralized optimization more vulnerable to client drift and unstable convergence. To address these challenges, we propose DECA, a resource-efficient decentralized FPFT framework for LLMs on non-IID data. DECA partitions model parameters into disjoint blocks and performs sequential block-wise Adam optimization, reducing resource consumption while preserving decentralized full-parameter adaptation. To stabilize training, DECA further introduces first- and second-order block-wise moment estimates with fresh local gradient statistics and consensus-derived discrepancy signals. We provide rigorous theoretical analysis and extensive experiments, showing that DECA achieves fast convergence, strong downstream performance, and significant resource efficiency.

cs.LG↗

LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation

Language-conditioned goal navigation (LGN) requires agents to locate user-specified targets without step-by-step guidance. However, existing benchmarks largely focus on category-level goals or rely on instance descriptions generated by vision-language models (VLMs), which often contain ambiguities and semantic errors, limiting systematic and reliable evaluation. We introduce HieraNav, an open-vocabulary LGN task with goals specified at four hierarchical semantic levels: scene, room, region, and instance. To this end, we present Language as a Map (LangMap), to our knowledge the first real-world 3D indoor navigation benchmark with human-verified semantic annotations to support tasks across all four goal levels. LangMap provides region labels and discriminative region and instance descriptions covering 414 object categories, produced through a rigorous contrastive annotation protocol comparing same-scene regions and instances, and contains over 18K tasks. Each target is paired with concise and detailed descriptions, enabling evaluation across instruction styles. Quantitative and qualitative analyses validate our annotation quality; notably, our instance descriptions outperform GOAT-Bench annotations by 23 percentage points in text-to-view matching. We further introduce PlaNaVid, a strong RGB-only baseline that combines Bounded Diverse Memory (BDM) with high-level planning to prime a reactive policy for multi-goal navigation. PlaNaVid achieves top-tier success rates without depth, 3D scene representations, or object masks. Further analysis shows that memory and richer context boost performance, while long-tailed categories, small objects, distant targets, and multi-goal completion remain open challenges. The benchmark is available at https://bo-miao.github.io/LangMap

cs.CV↗

What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness

Hallucination remains one of the key challenges undermining the reliability of Large Vision-Language Models (LVLMs). But what makes an LVLM hallucinate less? Many existing efforts focus on improving internal components of the model. We argue that hallucination fundamentally stems from how the model architecture is designed. To investigate this, we factor the architecture design into three dimensions: Linguistic Foundation (LF), Visual Representation (VR), and Semantic Alignment (SA), and categorize hallucinations into Co-occurrence, Similarity, and previously overlooked Uncertainty types. Building on this formulation, we propose CoSimUE, a benchmark that creates fine-grained hallucination scenarios through controlled textual perturbations and random perturbations, enabling mapping between design choices and hallucination behaviors. Experiments across 7 design aspects show that: 1) the widely emphasized scaling of model parameters has only limited impact on reducing all three types of hallucinations; 2) larger and better-trained language foundations can reduce co-occurrence hallucinations; 3) stronger visual encoders and higher resolutions mitigate similarity errors; 4) effective alignment strategies alleviate uncertainty hallucinations. 5) Furthermore, cross-dimensional analysis reveals that jointly enhancing visual fidelity and alignment quality yields the most comprehensive improvements. This study provides the first systematic exploration linking architecture-level design to hallucination robustness, offering practical guidance for developing reliable and efficient LVLMs.

cs.CV↗

Fundamental Physics and Cosmology with TianQin

The exploration of the surrounding world and the universe is an important theme in the legacy of humankind. The detection of gravitational waves is adding a new dimension to this grand effort. What are the fundamental physical laws governing the dynamics of the universe? What is the fundamental composition of the universe? How has the universe evolved in the past and how will it evolve in the future? These are the basic questions that press for answers. The space-based gravitational wave detector TianQin will tune in to gravitational waves in the millihertz frequency range ($10^{-4} \sim 1$ Hz, to be specific), opening a new gravitational wave spectrum window to explore many of the previously hidden sectors of the universe. TianQin will discover many astrophysical systems, populating the universe at different redshifts: some will be of new types that have never been detected before, some will have very high signal-to-noise ratios, and some will have very high parameter estimation precision. The plethora of information collected will bring us to new fronts on which to search for the breaking points of general relativity, the possible violation of established physical laws, the signature of possible new gravitational physics and new fundamental fields, and to improve our knowledge on the expansion history of the universe. In this white paper, we highlight the advances that TianQin can bring to fundamental physics and cosmology.

gr-qc↗

Design and Characterization of Racetrack 3D-Trench Silicon Sensor Based on 8-Inch Process with Excellent Time Resolution

In the extreme environments of high-luminosity colliders, traditional planar silicon sensors suffer severe radiation-induced performance degradation and fail to satisfy the stringent demands of high-precision tracking and high-speed timing in particle physics. 3D silicon sensors enhance radiation hardness by shortening charge collection distance, yet conventional designs with columnar or square-cell trench electrodes exhibit non-uniform electric fields, including saddle points and low-field regions, which degrade charge collection efficiency and timing resolution. This work presents a novel racetrack 3D-trench silicon sensor with continuous racetrack electrodes surrounding a long central collection electrode, aiming to eliminate electric field inhomogeneities. For the first time, a 23 $μ$m shallow-etched device was fabricated on an 8-inch platform, which provides a promising basis for its subsequent mass production and engineering applications. The device performance was systematically evaluated through theoretical analysis, 3D TCAD simulations, and characterization using semiconductor parameter analyzers and transient current technique (TCT) measurements. The sensor achieves leakage current below 0.2 nA, breakdown voltage above 110 V, full depletion voltage as low as a few volts, capacitance as low as 650 fF, collected charge of 4 fC, time response of about 640 ps, and time resolution of 50 ps. This large-scale manufacturable, shallow-etched racetrack 3D-trench silicon sensor provides a competitive device solution for portable radiation detection and next-generation 4D tracking under high-radiation and high-event-rate conditions.

physics.ins-det↗

Si/SiGe multi-channel superlattice structure epitaxial growth with segmented temperature control for Next-Generation Logic Devices

Stacking multiple SiSiGe channels in advanced logic devices faces severe thermal budget accumulation, which degrades interfaces via Ge-Si interdiffusion and strain relaxation.This strategy lowers the Ge diffusion coefficient to 5.6-7% of its value at 650C (Arrhenius estimate), suppressing interdiffusion and preserving pseudomorphic strain. The 4 + 4 channel stack exhibits clear XRD satellite peaks, fully coherent strain state (reciprocal space mapping), sharp interfaces (1.5-2.6 nm transition width) and low RMS roughness (0.08 nm). Quantitative analysis from bottom to top reveals that prolonged high-temperature exposure broadens bottom interfaces and dilutes Ge concentration (from 20% to 18.5%), while the top stack maintains design targets. This work provides a process-physics understanding of thermal budget effects in multi-channel superlattices and establishes a high-quality material foundation for advanced logic devices beyond 2 nm node.

cond-mat.mtrl-sci↗

FluxShard: Motion-Aware Feature Cache Reuse for Collaborative Video Analytics in Mobile Edge Computing

Caching and reusing intermediate features across consecutive frames is a common technique to reduce redundant computation and transmission for edge-cloud video analytics in mobile edge computation. Existing methods manage the cache in a fixed or globally shifted coordinate system, treating it as an indivisible whole. Under the non-uniform motion patterns of mobile scenes, this whole-scene granularity invalidates large portions of the cache even when most content has merely shifted spatially, wasting computation and bandwidth. The root cause is a granularity mismatch: the cache is managed per scene, yet motion varies per region. In this paper, we present FluxShard, a motion-aware edge-cloud video analytics system that uses codec-level block motion vectors (MVs) to manage feature cache reuse and recomputation at the granularity of individual motion regions. By re-indexing cached features along per-block MVs, FluxShard separates spatial displacement from content changes, recovering reusable content that whole-scene methods would otherwise discard. To ensure correct reuse under heterogeneous motion, the Receptive Field Alignment Principle (RFAP) identifies, from the input-level MV field alone, the positions that must be recomputed due to inconsistent spatial composition within receptive fields. To maintain cache coherence across frames, MV-guided cache remapping warps the entire feature cache to the current coordinate system each frame, sustaining a high reuse ratio over time. A profiling-driven dispatcher routes the remaining sparse workload between edge and cloud for lower latency. Evaluation across multiple vision tasks, dynamic video benchmarks, and network conditions shows that FluxShard reduces latency by 32.6-83.8% and energy by 14.9-64.0% over all baselines under the prescribed accuracy budget.

cs.NI↗

To Intervene or Not: Guiding Inference-time Alignment with Probabilistic Model Blending

The wide deployment of LLMs has made model alignment necessary to make newly trained models safely and effectively respond to user instructions. Among different methods, inference-time alignment is often cheaper as it intervenes (i.e., offers guidances) only during output generation. Existing proposals apply guidances extracted from certain aligned models without properly assessing their reliability. Nonetheless, our systematic evaluation reveals that guidance effectiveness varies drastically across models; since ineffective guidances lead to further confusion and thus further interventions, the resulting excessive interventions typically indicate poor performance. To make interventions more effective and thus more efficient, we introduce BlendIn, an inference-time alignment framework that shifts from binary decisions to creating hybrid distributions integrating both models' knowledge. BlendIn stabilizes inference-time alignment by performing quality-aware alignment and proportionally weighting each model's contribution based on reliability. Compared with existing works, it preserves beneficial guidance while downweighting unreliable suggestions. BlendIn provides both diagnostic signals and mitigation strategies for misaligned guidance, achieving consistent and up to 50% performance improvement on challenging model pairs. Our code is available at: https://github.com/DecayingSeart/BlendIn.

cs.LG↗

Bulk superconductivity up to 96 K in pressurized nickelate single crystals

Recently, the Ruddlesden-Popper bilayer nickelate $La_3Ni_2O_7$ has emerged as a superconductor with a transition temperature ($T_c$) of approximately 80 K above 14 GPa (Refs. 1-3). Achieving higher $T_c$ in nickelate superconductors, along with the synthesis of reproducible high-quality single crystals without relying on high-oxygen-pressure growth conditions, remains a significant challenge$^{[4-7]}$. Here we report superconductivity up to 96 K under high pressure in bilayer nickelate single crystals synthesized at ambient pressure. Energy-dispersive spectroscopy, single-crystal X-ray diffraction, nuclear quadrupole resonance and scanning transmission electron microscopy evidenced high crystal quality of the flux-grown $La_2SmNi_2O_{7-δ}$ single crystals. $La_2SmNi_2O_7$ exhibits clear bulk superconductivity, including zero resistivity ($T_{c,max}^{onset}$ = 92 K and $T_{c,max}^{zero}$ = 73 K at 21.6 GPa) and the Meissner effect ($T_c$= 60 K at 20.6 GPa). A low-temperature high-pressure structural study indicates that both monoclinic and tetragonal structures can support superconductivity in this bilayer nickelate. Furthermore, we established a correlation between higher $T_c$ under high pressures and larger in-plane lattice distortion under ambient conditions, corroborated by observing even higher $T_c^{onset}$ of 96 K in $La_{1.57}Sm_{1.43}Ni_2O_{7-δ}$. This study overcomes key limitations in growing nickelate superconductor crystals, resolves the crystal structure in the superconducting state and demonstrates an effective pathway towards achieving higher $T_c$.

cond-mat.supr-con↗

Physically-Induced Atmospheric Adversarial Perturbations: Enhancing Transferability and Robustness in Remote Sensing Image Classification

Adversarial attacks pose a severe threat to the reliability of deep learning models in remote sensing (RS) image classification. Most existing methods rely on direct pixel-wise perturbations, failing to exploit the inherent atmospheric characteristics of RS imagery or survive real-world image degradations. In this paper, we propose FogFool, a physically plausible adversarial framework that generates fog-based perturbations by iteratively optimizing atmospheric patterns based on Perlin noise. By modeling fog formations with natural, irregular structures, FogFool generates adversarial examples that are not only visually consistent with authentic RS scenes but also deceptive. By leveraging the spatial coherence and mid-to-low-frequency nature of atmospheric phenomena, FogFool embeds adversarial information into structural features shared across diverse architectures. Extensive experiments on two benchmark RS datasets demonstrate that FogFool achieves superior performance: not only does it exceed in white-box settings, but also exhibits exceptional black-box transferability (reaching 83.74% TASR) and robustness against common preprocessing-based defenses such as JPEG compression and filtering. Detailed analyses, including confusion matrices and Class Activation Map (CAM) visualizations, reveal that our atmospheric-driven perturbations induce a universal shift in model attention. These results indicate that FogFool represents a practical, stealthy, and highly persistent threat to RS classification systems, providing a robust benchmark for evaluating model reliability in complex environments.

cs.CV↗

Atoms of Compacta on Closed Surfaces

For any compact set $K$ lying on a closed surface $\mathcal{S}$ we introduce a closed equivalence relation $\sim$, called the {\em Schönflies equivalence} on $K$. We show that every class $[x]_\sim$ of $\sim$ is a continuum and that the resulting quotient space $K\!/\!\sim$ is a {\em Peano compactum}. By definition, all components of a Peano compactum are locally connected and for any $\varepsilon>0$ only finitely many of them have diameter greater than $\varepsilon$. The decomposition $\mathcal{D}_K=\{[x]_\sim: x\in K\}$ refines every other upper semicontinuous decomposition of $K$ into subcontinua that has a Peano compactum as its quotient space. In other words, $\mathcal{D}_K$ is the {\em core decomposition of $K$} with Peano quotient. The elements of $\mathcal{D}_K$ are called {\em atoms} of $K$. We also show that for any branched covering $f: \mathcal{S}^*\rightarrow \mathcal{S}$ from a closed surface $\mathcal{S}^*$ to $\mathcal{S}$, every atom of $f^{-1}(K)$ is sent into an atom of $K$. If $f$ is even a covering, it sends every atom of $f^{-1}(K)$ onto an atom of $K$. We illustrate our theory with examples and show that it cannot be generalized to $n$-manifolds with $n\ge 3$ by providing a detailed counterexample in~$\mathbb{R}^3$.

math.GN↗

Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents

Large Vision-Language Models (LVLMs) empower autonomous mobile agents, yet their security under realistic mobile deployment constraints remains underexplored. While agents are vulnerable to visual prompt injections, stealthily executing such attacks without requiring system-level privileges remains challenging, as existing methods rely on persistent visual manipulations that are noticeable to users. We uncover a consistent discrepancy between human and agent interactions: automated agents generate near-zero contact touch signals. Building on this insight, we propose a new attack paradigm, agent-only perceptual injection, where malicious content is exposed only during agent interactions, while remaining not readily perceived by human users. To accommodate mobile UI constraints and one-shot interaction settings, we introduce HG-IDA*, an efficient one-shot optimization method for constructing jailbreak prompts that evade LVLM safety filters. Experiments demonstrate that our approach induces unauthorized cross-app actions, achieving 82.5% planning and 75.0% execution hijack rates on GPT-4o. Our findings highlight a previously underexplored attack surface in mobile agent systems and underscore the need for defenses that incorporate interaction-level signals.

cs.CR↗

Dual-Envelope Constrained Nonlinear MPC for Distributed Drive Electric Vehicles Drifting Under Bounded Steering and Direct Yaw-Moment Control

Distributed drive electric vehicles offer superior yaw moment control for autonomous drifting in extreme maneuvers. Conventional drift analysis constructs stability boundaries from open loop equilibria points and assumes a fixed envelope structure. However, coupling among control inputs reshapes the phase plane and shifts saddle point location, which can invalidate open loop envelopes when used for closed loop drifting. To address this issue, a saddle point coordinate model is established in this paper by combining a nonlinear tire model with the handling diagram and explicitly accounting for road adhesion coefficient, longitudinal velocity, front wheel steering angle, and additional yaw moment. Based on saddle point properties, an extended dual envelope framework is constructed in the phase plane of slip angle and yaw rate. Using the convergence tendency of state points toward saddle points under bounded control inputs, the outer envelope defines a recoverable set under constraints on front wheel steering angle and additional yaw moment. The inner envelope characterizes the non-drifting stability region associated with unsaturated tire forces. Finally, a nonlinear model predictive control (NMPC) controller is developed using the extended dual envelope constraint. Hardware-in-the-loop experiments show that, compared with NMPC without envelope constraints, the proposed method enables smoother convergence toward the drift saddle point, reduces the steady-state tracking errors of vehicle speed, sideslip angle, and yaw rate by 33.07%, 71.18%, and 31.27%, respectively, and decreases the peak tracking error by 63.66% under road-friction mismatch.

eess.SY↗

Cascade of Spin Liquids in a Bilayer Triangular-lattice Antiferromagnet Rb_2Co_2(SeO_3)_3

In frustrated Ising magnets, classical spin liquids (CSLs) with macroscopic ground-state degeneracy can survive against conventional magnetic order, as exemplified by systems on triangular, kagome and pyrochlore lattices at zero field. Here we report the discovery of a high-field route toward spin liquids in a bilayer triangular lattice antiferromagnet, Rb$_2$Co$_2$(SeO$_3$)$_3$. We demonstrate that a cascade of CSLs -- characterized by doubly degenerate one-up-one-down local spin configurations and a residual entropy of 1/2(1-M/M_s)Rln2 per mole -- emerges through field-controlled dilution of Ising dimers. Owing to the interplay of intra- and inter-layer interactions, these CSLs are further stabilized by lattice symmetry breaking at fractional magnetization plateaus. Such field-induced spin liquids can be understood as a consequence of generalized ice rules, analogous to those governing in pyrochlore antiferromagnets. In particular, the 5/6-plateau state is a candidate quantum spin liquid. Our results thereby establish a new pathway for exploring diverse spin liquid states across both classical and quantum regimes.

cond-mat.str-el↗

LOPT: Learning Optimal Pigovian Tax in Sequential Social Dilemmas

In multi-agent reinforcement learning, each agent acts to maximize its individual accumulated rewards. Nevertheless, individual accumulated rewards could not fully reflect how others perceive them, resulting in selfish behaviors that undermine global performance. The externality theory, defined as ``the activities of one economic actor affect the activities of another in ways that are not reflected in market transactions,'' is applicable to analyze the social dilemmas in MARL. One of its most profound non-market solutions, ``Pigovian Tax'', which internalizes externalities by taxing those who create negative externalities and subsidizing those who create positive externalities, could aid in developing a mechanism to resolve MARL's social dilemmas. The purpose of this paper is to apply externality theory to analyze social dilemmas in MARL. To internalize the externalities in MARL, the \textbf{L}earning \textbf{O}ptimal \textbf{P}igovian \textbf{T}ax method (LOPT), is proposed, where an additional agent is introduced to learn the tax/allowance allocation policy so as to approximate the optimal ``Pigovian Tax'' which accurately reflects the externalities for all agents. Furthermore, a reward shaping mechanism based on the approximated optimal ``Pigovian Tax'' is applied to reduce the social cost of each agent and tries to alleviate the social dilemmas. Compared with existing state-of-the-art methods, the proposed LOPT leads to higher collective social welfare in both the Escape Room and the Cleanup environments, which shows the superiority of our method in solving social dilemmas.

cs.MA↗

HIPO: Instruction Hierarchy via Constrained Reinforcement Learning

Hierarchical Instruction Following (HIF) refers to the problem of prompting large language models with a priority-ordered stack of instructions. Standard methods like RLHF and DPO typically fail in this problem since they mainly optimize for a single objective, failing to explicitly enforce system prompt compliance. Meanwhile, supervised fine-tuning relies on mimicking filtered, compliant data, which fails to establish the priority asymmetry at the algorithmic level. In this paper, we introduce \textsc{HIPO}, a novel alignment framework that formulates HIF as a Constrained Markov Decision Process. \textsc{HIPO} elevates system prompts from mere input context to strict algorithmic boundaries. Using a primal-dual safe reinforcement learning approach, the algorithm dynamically enforces system prompt compliance as an explicit constraint, maximizing user utility strictly within this feasible region. Extensive evaluations across diverse model architectures (e.g., Qwen, Phi, Llama) demonstrate that \textsc{HIPO} significantly improves both system compliance and user utility. Furthermore, mechanistic analysis reveals that this constrained optimization autonomously drives the model to shift its attention toward long-range system tokens, providing a principled foundation for reliable LLM deployment in complex workflows.

cs.LG↗

Cheating Stereo Matching in Full-scale: Physical Adversarial Attack against Binocular Depth Estimation in Autonomous Driving

Though deep neural models adopted to realize the perception of autonomous driving have proven vulnerable to adversarial examples, known attacks often leverage 2D patches and target mostly monocular perception. Therefore, the effectiveness of Physical Adversarial Examples (PAEs) on stereo-based binocular depth estimation remains largely unexplored. To this end, we propose the first texture-enabled physical adversarial attack against stereo matching models in the context of autonomous driving. Our method employs a 3D PAE with global camouflage texture rather than a local 2D patch-based one, ensuring both visual consistency and attack effectiveness across different viewpoints of stereo cameras. To cope with the disparity effect of these cameras, we also propose a new 3D stereo matching rendering module that allows the PAE to be aligned with real-world positions and headings in binocular vision. We further propose a novel merging attack that seamlessly blends the target into the environment through fine-grained PAE optimization. It has significantly enhanced stealth and lethality upon existing hiding attacks that fail to get seamlessly merged into the background. Extensive evaluations show that our PAEs can successfully fool the stereo models into producing erroneous depth information.

cs.CV↗