Searcharxiv⌕ Search

arXiv subjects

Jun Luo

Publications and source records attributed to Jun Luo.

At least 55 records · Page 3Linked to original sources

HIPO: Instruction Hierarchy via Constrained Reinforcement Learning

Hierarchical Instruction Following (HIF) refers to the problem of prompting large language models with a priority-ordered stack of instructions. Standard methods like RLHF and DPO typically fail in this problem since they mainly optimize for a single objective, failing to explicitly enforce system prompt compliance. Meanwhile, supervised fine-tuning relies on mimicking filtered, compliant data, which fails to establish the priority asymmetry at the algorithmic level. In this paper, we introduce \textsc{HIPO}, a novel alignment framework that formulates HIF as a Constrained Markov Decision Process. \textsc{HIPO} elevates system prompts from mere input context to strict algorithmic boundaries. Using a primal-dual safe reinforcement learning approach, the algorithm dynamically enforces system prompt compliance as an explicit constraint, maximizing user utility strictly within this feasible region. Extensive evaluations across diverse model architectures (e.g., Qwen, Phi, Llama) demonstrate that \textsc{HIPO} significantly improves both system compliance and user utility. Furthermore, mechanistic analysis reveals that this constrained optimization autonomously drives the model to shift its attention toward long-range system tokens, providing a principled foundation for reliable LLM deployment in complex workflows.

cs.LG↗

Cheating Stereo Matching in Full-scale: Physical Adversarial Attack against Binocular Depth Estimation in Autonomous Driving

Though deep neural models adopted to realize the perception of autonomous driving have proven vulnerable to adversarial examples, known attacks often leverage 2D patches and target mostly monocular perception. Therefore, the effectiveness of Physical Adversarial Examples (PAEs) on stereo-based binocular depth estimation remains largely unexplored. To this end, we propose the first texture-enabled physical adversarial attack against stereo matching models in the context of autonomous driving. Our method employs a 3D PAE with global camouflage texture rather than a local 2D patch-based one, ensuring both visual consistency and attack effectiveness across different viewpoints of stereo cameras. To cope with the disparity effect of these cameras, we also propose a new 3D stereo matching rendering module that allows the PAE to be aligned with real-world positions and headings in binocular vision. We further propose a novel merging attack that seamlessly blends the target into the environment through fine-grained PAE optimization. It has significantly enhanced stealth and lethality upon existing hiding attacks that fail to get seamlessly merged into the background. Extensive evaluations show that our PAEs can successfully fool the stereo models into producing erroneous depth information.

cs.CV↗

Zoom to Essence: Trainless GUI Grounding by Inferring upon Interface Elements

Multimodal Large Language Model (MLLM)-based Graphical User Interface (GUI) agents develop rapidly, with visual grounding that maps natural language instructions to target UI elements serving as the core capability. Existing GUI agents typically fine-tune MLLM on massive datasets to handle challenges in understanding instructions and UI interfaces, which not only incurs high data annotation costs but also makes performance dependent on data quality and distribution. To avoid such cumbersome yet ineffective training, we notice that complex UI interfaces can be decomposed into basic visual elements directly understandable by common MLLMs. Consequently, we propose ZoomUI that leverages inference scaling to guide common MLLMs in progressively anchor instruction elements to increasingly detailed interface elements. Specifically, ZoomUI first optimizes the latent thinking to transform original instruction into element visual features description, and subsequently leverages internal attention to iteratively zoom in target element interface region. Evaluations on extensive benchmarks demonstrate that ZoomUI reaches or even surpasses SOTA baselines.

cs.LG↗

SceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene Generation

Text-to-3D scene generation from natural language is highly desirable for digital content creation. However, existing methods are largely domain-restricted or reliant on predefined spatial relationships, limiting their capacity for unconstrained, open-vocabulary 3D scene synthesis. In this paper, we introduce SceneAssistant, a visual-feedback-driven agent designed for open-vocabulary 3D scene generation. Our framework leverages modern 3D object generation model along with the spatial reasoning and planning capabilities of Vision-Language Models (VLMs). To enable open-vocabulary scene composition, we provide the VLMs with a comprehensive set of atomic operations (e.g., Scale, Rotate, FocusOn). At each interaction step, the VLM receives rendered visual feedback and takes actions accordingly, iteratively refining the scene to achieve more coherent spatial arrangements and better alignment with the input text. Experimental results demonstrate that our method can generate diverse, open-vocabulary, and high-quality 3D scenes. Both qualitative analysis and quantitative human evaluations demonstrate the superiority of our approach over existing methods. Furthermore, our method allows users to instruct the agent to edit existing scenes based on natural language commands. Our code is available at https://github.com/ROUJINN/SceneAssistant

cs.CV↗

EmoSURA: Towards Accurate Evaluation of Detailed and Long-Context Emotional Speech Captions

Recent advancements in speech captioning models have enabled the generation of rich, fine-grained captions for emotional speech. However, the evaluation of such captions remains a critical bottleneck: traditional N-gram metrics fail to capture semantic nuances, while LLM judges often suffer from reasoning inconsistency and context-collapse when processing long-form descriptions. In this work, we propose EmoSURA, a novel evaluation framework that shifts the paradigm from holistic scoring to atomic verification. EmoSURA decomposes complex captions into Atomic Perceptual Units, which are self-contained statements regarding vocal or emotional attributes, and employs an audio-grounded verification mechanism to validate each unit against the raw speech signal. Furthermore, we address the scarcity of standardized evaluation resources by introducing SURABench, a carefully balanced and stratified benchmark. Our experiments show that EmoSURA achieves a positive correlation with human judgments, offering a more reliable assessment for long-form captions compared to traditional metrics, which demonstrated negative correlations due to their sensitivity to caption length.

cs.SD↗

A Geometrically Convergent Solution to Spatial Hypercube Queueing Models

The hypercube queueing model was initially developed to address spatial queueing problems and has found wide applications in emergency services, such as ambulance and police systems. While the model was originally designed for homogeneous service rates, we extend it to handle heterogeneous service rates by devising an exact solution through a birth-death process and an equivalent reformulation. We demonstrate that our algorithm converges to the exact solution at a geometric rate. Additionally, we developed a parallel algorithm that leverages the convergence property and two structural features of the hypercube model, achieving more than 91% parallelization. Numerical experiments on emergency medical service systems show that our sequential algorithm is over 1,000 times faster than the sparse solver and more than 500 times faster than discrete-event simulation, while maintaining high accuracy. The parallel algorithm further improves efficiency, achieving an approximately eightfold speedup with 12 processing units, with additional gains possible when more computational resources are available. Overall, the proposed algorithms improve computational efficiency and enable the solution of large-scale problems that are otherwise intractable using traditional approaches.

math.OC↗

Learning Virtual Machine Scheduling in Cloud Computing through Language Agents

In cloud services, virtual machine (VM) scheduling is a typical Online Dynamic Multidimensional Bin Packing (ODMBP) problem, characterized by large-scale complexity and fluctuating demands. Traditional optimization methods struggle to adapt to real-time changes, domain-expert-designed heuristic approaches suffer from rigid strategies, and existing learning-based methods often lack generalizability and interpretability. To address these limitations, this paper proposes a hierarchical language agent framework named MiCo, which provides a large language model (LLM)-driven heuristic design paradigm for solving ODMBP. Specifically, ODMBP is formulated as a Semi-Markov Decision Process with Options (SMDP-Option), enabling dynamic scheduling through a two-stage architecture, i.e., Option Miner and Option Composer. Option Miner utilizes LLMs to discover diverse and useful non-context-aware strategies by interacting with constructed environments. Option Composer employs LLMs to discover a composing strategy that integrates the non-context-aware strategies with the contextual ones. Extensive experiments on real-world enterprise datasets demonstrate that MiCo achieves a 96.9\% competitive ratio in large-scale scenarios involving more than 10,000 virtual machines. It maintains high performance even under nonstationary request flows and diverse configurations, thus validating its effectiveness in complex and large-scale cloud environments.

cs.LG↗

Transformer-Based Multipath Congestion Control: A Decoupled Approach for Wireless Uplinks

The proliferation of artificial intelligence applications on edge devices necessitates efficient transport protocols that leverage multi-homed connectivity across heterogeneous networks. While Multipath TCP enables bandwidth aggregation, its in-kernel congestion control mechanisms lack the programmability and flexibility needed for achieving efficient transmission. Additionally, inherent measurement noise renders network state partially observable, challenging data-driven approaches like deep reinforcement learning (DRL). To address these challenges, we propose a Transformer-based Congestion Control Optimization (TCCO) framework for multipath transport. TCCO employs a decoupled architecture that offloads control decisions to an external decision engine via a lightweight in-kernel client and user-space proxy, enabling edge devices to leverage external computational resources while maintaining TCP/IP compatibility. The Transformer-based DRL agent in the external decision engine uses self-attention to capture temporal dependencies, filter noise, and coordinate control across subflows through a unified policy. Extensive evaluation on both simulated and real dual-band Wi-Fi testbeds demonstrates that TCCO achieves superior adaptability and performance than state-of-the-art baselines, validating the feasibility and effectiveness of TCCO for wireless networks.

cs.NI↗

NMR evidence of spin supersolid and Pomeranchuk effect behaviors in the triangular-lattice antiferromagnet Rb$_2$Ni$_2$(SeO$_3$)$_3$

We performed $^{85}$Rb nuclear magnetic resonance (NMR) measurements on the $S$ = 1 bilayer triangular-lattice antiferromagnet Rb$_2$Ni$_2$(SeO$_3$)$_3$ in magnetic fields up to 26 T. In the field range from 3 T to 26 T, the NMR spectral lines split and their respective spectral weight ratios reveal the existence of the magnetic up-up-down (UUD) phase, although the 1/3-plateau phase is only reached at fields above 16 T. Two distinct gapless regimes are further identified: one at low fields and low temperatures, and the other at high fields and high temperatures, consistent with the spin supersolid Y and V phases. Notably, the UUD-V phase boundary exhibits a negative slope in $dT/dH$, where the supersolid phase is located at temperatures above the solid phase due to strong low-energy spin fluctuations.

cond-mat.str-el↗

Reply to "Threefold error in the reported zero-field cooled magnetic moment of single crystal $La_2SmNi_2O_7$ (arXiv: 2602.23240)"

We respond to the critique by Aleksandr V. Korolev and Evgeny F. Talantsev on the superconducting phase fraction ($f$) calculations in Li et al. Nature 649, 871-878 (2026). First, the weak upturn in the low-temperature tail of our data has been confirmed to originate from the background, and the paramagnetic Meissner effect is absent in our case; thus, field-cooled (FC) data can be used for superconducting phase fraction calculations. Second, demagnetization effect must be calculated based on the actual measured moment as a function of $f$, which has been well-established and routinely employed in the superconductivity community. In contrast, Korolev and Talantsev treated the demagnetization field as a constant; thus, their calculation underestimates $f$ by a factor of $(1-Nχ_{meas})(1-N)$. This factor is close to 1/3, given $N$ = 0.849, $χ_{meas}$ = -1.313 in our study, which explains the origin of their deviated result (nearly three times smaller than our results). Third, our sample is a homogeneous high-quality bulk single crystal, evidenced by various techniques, making the existence of multiple discrete superconducting regions highly unlikely. We conclude that the superconducting phase fraction calculations reported in Li et al. Nature 649, 871-878 (2026) are not invalidated by the analyses presented in Korolev et al. arXiv: 2602.23240 (2026).

cond-mat.supr-con↗

Image Quality Assessment: Exploring Quality Awareness via Memory-driven Distortion Patterns Matching

Existing full-reference image quality assessment (FR-IQA) methods achieve high-precision evaluation by analysing feature differences between reference and distorted images. However, their performance is constrained by the quality of the reference image, which limits real-world applications where ideal reference sources are unavailable. Notably, the human visual system has the ability to accumulate visual memory, allowing image quality assessment on the basis of long-term memory storage. Inspired by this biological memory mechanism, we propose a memory-driven quality-aware framework (MQAF), which establishes a memory bank for storing distortion patterns and dynamically switches between dual-mode quality assessment strategies to reduce reliance on high-quality reference images. When reference images are available, MQAF obtains reference-guided quality scores by adaptively weighting reference information and comparing the distorted image with stored distortion patterns in the memory bank. When the reference image is absent, the framework relies on distortion patterns in the memory bank to infer image quality, enabling no-reference quality assessment (NR-IQA). The experimental results show that our method outperforms state-of-the-art approaches across multiple datasets while adapting to both no-reference and full-reference tasks.

cs.CV↗

Path to Diversity: A Primer on ISAC-izing Commodity Wi-Fi for Practical Deployments

Integrated Sensing and Communication (ISAC) has emerged as a key paradigm in next-generation wireless networks. While the ubiquity and low cost of commodity Wi-Fi make it an ideal platform for wide-scale sensing, it is the continuous evolution of Wi-Fi standards-towards higher frequency bands, wider bandwidths, and larger antenna arrays-that fundamentally unlocks the physical resources required for high-performance ISAC. To structure this rapidly expanding field, numerous surveys have appeared. However, prevailing literature predominantly adopts a top-down perspective, emphasizing upper-layer applications or deep learning models while treating the physical layer as an opaque abstraction. Consequently, these works often fail to touch the bottom layer of signal formation and lack technical guidance on overcoming the physical barriers that constrain sensing performance. To bridge this gap, this tutorial takes a bottom-up approach, systematically analyzing the sensing gains brought by Wi-Fi advancements through the lens of physical-layer diversity. We organize the framework around four orthogonal dimensions: i) Temporal Diversity addresses synchronization gaps to enable absolute ranging; ii) Frequency Diversity expands the effective bandwidth to sharpen range resolution; iii) Link Diversity leverages distributed topologies and digital feedback to achieve ubiquitous observability; and iv) Spatial Diversity utilizes multi-antenna arrays to combine passive angular discrimination with active directional control. Collectively, these orthogonal dimensions resolve fundamental ambiguities in time, range, and space, bridging physical capabilities with challenging sensing diversities. By synthesizing these dimensions, this tutorial provides a comprehensive guide for "ISAC-izing" commodity Wi-Fi, paving the way for future standardization and robust deployment.

cs.NI↗

Physics Guided Exponential Model Design of High Ge Content SiGe Selective Epitaxy for Gate All Around Source/Drain Applications

High germanium content silicon germanium (SiGe) epitaxy is critical for strain engineering in advanced gate all around (GAA) transistors. This paper demonstrates a physics guided exponential function model that quantitatively links selective epitaxial growth (SEG) parameters to Ge incorporation kinetics in nanoscale trenches. By coupling surface diffusion limited transport, gradient strain, and competitive adsorption dynamics, the model predicts optimal conditions for bottom-up filling with maximal Ge content. For trenches with widths of approximately 60 nm, the optimized process achieved a maximum Ge content of 57.93% and demonstrated 100% selectivity against silicon nitride (SiN) and silicon dioxide (SiO). Cross sectional TEM and EDS analyses reveal a graded Ge profile that minimizes interfacial defects and strain energy. Our results show that the established process physics correlation will significantly facilitate the development of GAA devices with 5nm CMOS technology nodes and beyond.

physics.app-ph↗

Quasi-Medial Distance Field (Q-MDF): A Robust Method for Approximating and Discretizing Neural Medial Axes

The medial axis, a lower-dimensional descriptor that captures the extrinsic structure of a shape, plays an important role in digital geometry processing. Despite its importance, computing the medial axis transform robustly from diverse inputs, especially point clouds with defects, remains a challenging problem. In this paper, we propose a new implicit method that deviates from traditional explicit medial axis computation. Our key technical insight is that the difference between the signed distance field (SDF) and the medial field (MF) of a solid shape relates to the unsigned distance field (UDF) of the shape's medial axis. This observation allows us to formulate medial axis extraction as an implicit reconstruction problem. By employing a modified double covering strategy, we recover the medial axis as the zero level-set of the UDF. Extensive experiments demonstrate that our method achieves higher accuracy and robustness in learning compact medial axis transforms from challenging meshes and point clouds, outperforming existing approaches.

cs.CV↗

SIDeR: Semantic Identity Decoupling for Unrestricted Face Privacy

With the deep integration of facial recognition into online banking, identity verification, and other networked services, achieving effective decoupling of identity information from visual representations during image storage and transmission has become a critical challenge for privacy protection. To address this issue, we propose SIDeR, a Semantic decoupling-driven framework for unrestricted face privacy protection. SIDeR decomposes a facial image into a machine-recognizable identity feature vector and a visually perceptible semantic appearance component. By leveraging semantic-guided recomposition in the latent space of a diffusion model, it generates visually anonymous adversarial faces while maintaining machine-level identity consistency. The framework incorporates momentum-driven unrestricted perturbation optimization and a semantic-visual balancing factor to synthesize multiple visually diverse, highly natural adversarial samples. Furthermore, for authorized access, the protected image can be restored to its original form when the correct password is provided. Extensive experiments on the CelebA-HQ and FFHQ datasets demonstrate that SIDeR achieves a 99% attack success rate in black-box scenarios and outperforms baseline methods by 41.28% in PSNR-based restoration quality.

cs.CV↗

Intertwined Charge and Spin Density Waves in Trilayer Nickelate La$_4$Ni$_3$O$_{10}$ Revealed by $^{139}$La NQR

The discovery of superconducting transitions in pressurized La$_3$Ni$_2$O$_{7}$ and La$_4$Ni$_3$O$_{10}$ has highlighted the pivotal role of density wave (DW) orders in nickelate superconductors. To gain a comprehensive understanding of the superconducting state, it is essential to elucidate the nature of the DW order. In this study, we utilized $^{139}$La nuclear quadrupole resonance (NQR) to investigate the charge density wave (CDW) and spin density wave (SDW) orders in both single-crystal and polycrystalline La$_4$Ni$_3$O$_{10}$. Near $T_{\rm{DW}} \approx 133$ K, an abrupt change in both the linewidth and frequency of the La(2) site in the single-crystal sample provides compelling evidence for a first-order-like phase transition. The pronounced broadening of the NQR lines indicates the incommensurate nature of the DW order. Furthermore, the spin-lattice relaxation rate divided by temperature 1/$T_1$$T$ exhibits a strong enhancement at $T_{\rm{DW}}$, indicating the strong spin fluctuations above the first-order DW transition. These observations suggest an intricate interplay between incommensurate CDW and SDW orders. Our findings offer critical insights into the microscopic mechanisms of the DW state in La$_4$Ni$_3$O$_{10}$ and establish an essential framework for exploring the interplay between DW and superconducting phases in nickelate superconductors.

cond-mat.supr-con↗

Spin-density-wave transition in monolayer-trilayer La3Ni2O7 single crystals

The recent discovery of high-temperature superconductivity in pressurized Ruddlesden-Popper nickelates stimulated intense research into their correlated electron physics. Establishing the diversity of ground states across different Ruddlesden-Popper phases is crucial for elucidating the superconducting mechanisms in these nickelates. Motivated by the recent report of superconductivity in hybrid 1212-type La5Ni3O11, we synthesized and investigated the long-range-ordered hybrid 1313-type La3Ni2O7. In contrast to its bilayer counterpart, the 1313-type La3Ni2O7 exhibits characteristic semiconducting behavior at ambient pressure, displaying a distinct anomaly at 170 K. This behavior is consistently evidenced by measurements of both magnetic susceptibility and specific heat. Nuclear magnetic resonance spectroscopy unambiguously indicates a spin-density-wave transition occurring at 170 K. High-pressure electrical transport measurements demonstrate the induction of metallization under pressure, yet reveal no discernible traces of superconductivity up to 65 GPa. Our findings establish hybrid 1313-type La3Ni2O7 as a new member of the Ruddlesden-Popper nickelate family exhibiting a distinct spin-density-wave transition, and offers a new platform for investigating the interplay among crystal structure, electronic orders, and superconductivity in hybrid nickelates.

cond-mat.supr-con↗

TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning

The demand of customized large language models (LLMs) has led to commercial LLMs offering black-box fine-tuning APIs, yet this convenience introduces a critical security loophole: attackers could jailbreak the LLMs by fine-tuning them with malicious data. Though this security issue has recently been exposed, the feasibility of such attacks is questionable as malicious training dataset is believed to be detectable by moderation models such as Llama-Guard-3. In this paper, we propose TrojanPraise, a novel finetuning-based attack exploiting benign and thus filter-approved data. Basically, TrojanPraise fine-tunes the model to associate a crafted word (e.g., "bruaf") with harmless connotations, then uses this word to praise harmful concepts, subtly shifting the LLM from refusal to compliance. To explain the attack, we decouple the LLM's internal representation of a query into two dimensions of knowledge and attitude. We demonstrate that successful jailbreak requires shifting the attitude while avoiding knowledge shift, a distortion in the model's understanding of the concept. To validate this attack, we conduct experiments on five opensource LLMs and two commercial LLMs under strict black-box settings. Results show that TrojanPraise achieves a maximum attack success rate of 95.88% while evading moderation.

cs.CR↗