SearcharxivSearch

arXiv subjects

Ji Zhou

Publications and source records attributed to Ji Zhou.

At least 19 recordsLinked to original sources

A Large-Language-Model Supported Personalized Driving Framework for Lane Change in Highway Scenarios

Personalized driving can improve the user acceptance of automated driving systems. However, existing methods still provide limited support for translating natural-language driving preferences, especially when such preferences are expressed implicitly, into executable and distinguishable driving behaviors. This paper proposes a large language model (LLM)-supported personalized driving framework for highway lane-change scenarios. The framework maps natural-language driving commands to executable planning parameters in the open-source Apollo automated driving stack according to three driving styles: aggressive, normal, and conservative. To establish this mapping, candidate planning parameters are evaluated based on the resulting lane-change behaviors, and style-specific parameter sets are constructed through clustering and style-intensity ranking. For command interpretation, a retrieval dataset is constructed to support retrieval-augmented generation (RAG), enabling LLM-based interpretation of implicit user commands. Experimental results show that the derived parameter sets generate distinguishable personalized lane-change behaviors, while RAG consistently improves preference interpretation, particularly for implicit commands. These results indicate the potential of integrating LLM-based natural-language interaction with Apollo to support personalized lane-change behavior generation. The source code and the relevant datasets are available at: https://github.com/ftgTUGraz/LLM-Personalized-Driving.

cs.RO

State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition

Conversational multimodal emotion recognition (MER) requires reliable prediction when language, acoustic, or visual observations are missing or unreliable. Many missing-modality methods reconstruct absent inputs, yet such recovery can be non-unique in dialogue context, and nonverbal cues may conflict with the target utterance. To this end, we propose CoRe-KD (Complete-view Reference-guided Knowledge Distillation), a state-anchored, conflict-regularized complete-view distillation framework for robust conversational MER. A complete-view teacher provides structured references, including prediction-level references, fused states, and modality-specific states. Complete-view State Anchoring (CSA) aligns incomplete-view student predictions and states with these references, while Nonverbal Conflict Exposure (NCE) trains on target-preserving nonverbal conflict views to reduce donor-label bias. Experiments on IEMOCAP and MELD, with CMU-MOSEI as a supplementary utterance-level check, show consistent gains under fixed- and random-missing protocols. Comprehensive ablation studies and further analyses support the role of CSA and the complementary effect of NCE.

cs.MM

Beyond Isolated Utterances: Cue-Guided Interaction for Context-Dependent Conversational Multimodal Understanding

Conversational multimodal understanding aims to infer the meaning or label of the current utterance from its preceding dialogue context together with textual, acoustic, and visual signals. Existing methods mainly strengthen contextual modeling through enhanced encoding, fusion, or propagation, but rarely abstract the context-utterance dependency into an explicit cue and incorporate it into later multimodal reasoning. To address this issue, we propose CUCI-Net for conversational multimodal understanding. CUCI-Net fully preserves the structural distinction between context and utterance during encoding, effectively abstracts their dependency into an interpretation cue by combining local modality evidence with global contextual evidence, and seamlessly integrates the resulting cue into the final multimodal interaction stage for context-conditioned prediction. Extensive experiments on mainstream benchmark datasets fully demonstrate the effectiveness of the proposed method.

cs.MM

Reconfigurable Resonant Multimode Nonlinear Coupling for UV-to-infrared Frequency Generation

On-chip coherent visible and near-infrared (NIR) light generation has broad applications in metrology, bio-sensing, and quantum information. High-Q microresonators are ideal candidates for generating light across such broad wavelength ranges via efficient second- ($\chi^{(2)}$) and third-order ($\chi^{(3)}$) nonlinear optical processes. However, harnessing these diverse nonlinearities simultaneously in a single microresonator remains elusive yet highly attractive both fundamentally and technologically. Here, we demonstrate coherent light generation from the ultraviolet to NIR in a silicon nitride microresonator pumped by a single continuous-wave telecom laser. This broad frequency generation arises from the interplay of $\chi^{(2)}$ and $\chi^{(3)}$ nonlinear processes. A cascade of nonlinear processes, including harmonic generation and optical parametric oscillation (OPO), is initiated by the photoinduced second harmonic generation enabled by all-optical poling. The dynamic reconfigurability of this $\chi^{(2)}$ nonlinearity enables access to different transverse spatial modes at the second harmonic, enabling highly tunable OPO processes triggered by hybrid modal phase matching conditions and yielding milliwatt-level NIR light. This work sheds new insights into the fundamental physics of cooperative nonlinear multimode interactions in resonant systems and provides a versatile approach for reconfigurable OPOs, highlighting their potential to generate light at wavelengths beyond the reach of photonic integrated lasers.

physics.optics

From stacking to function: emergent states and quantum devices in 2D superconductor heterostructures

Two-dimensional (2D) superconductors provide a powerful building block for engineering emergent quantum states shaped by reduced dimensionality, enhanced quantum fluctuations, and interfacial symmetry breaking. In van der Waals heterostructures, atomically sharp and lattice-mismatch-free interfaces enable superconductivity to be deliberately coupled with magnetism, spin orbit interaction, and band topology, allowing collective electronic orders to be combined and reconfigured in ways unattainable in bulk materials. This Review summarizes recent advances in vdW heterostructures of 2D superconductors, focusing on superconductor/magnet, superconductor/topological material, and superconductor/superconductor junctions. We discuss the microscopic mechanisms underlying proximity effects and highlight how interfacial exchange fields, spin orbit coupling, and twist-controlled tunneling give rise to unconventional pairing, long-range spin-triplet supercurrents, nonreciprocal Josephson transport, and topological superconductivity potentially hosting Majorana bound states. Beyond their fundamental significance, the ability to controllably generate topological and nonreciprocal superconducting states positions 2D superconductor heterostructures as promising building blocks for emerging quantum technologies, including ultra-sensitive quantum sensing, programmable superconducting logic, and energy-efficient quantum and neuromorphic computing architectures. Looking forward, advances in materials synthesis, interface engineering, and device integration are expected to further expand the scope and functionality of 2D superconductor heterostructures, reinforcing their role as a central platform for exploring and controlling emergent quantum phases.

cond-mat.mes-hall

Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025

Human identification at a distance (HID) is challenging because traditional biometric modalities such as face and fingerprints are often difficult to acquire in real-world scenarios. Gait recognition provides a practical alternative, as it can be captured reliably at a distance. To promote progress in gait recognition and provide a fair evaluation platform, the International Competition on Human Identification at a Distance (HID) has been organized annually since 2020. Since 2023, the competition has adopted the challenging SUSTech-Competition dataset, which features substantial variations in clothing, carried objects, and view angles. No dedicated training data are provided, requiring participants to train their models using external datasets. Each year, the competition applies a different random seed to generate distinct evaluation splits, which reduces the risk of overfitting and supports a fair assessment of cross-domain generalization. While HID 2023 and HID 2024 already used this dataset, HID 2025 explicitly examined whether algorithmic advances could surpass the accuracy limits observed previously. Despite the heightened difficulty, participants achieved further improvements, and the best-performing method reached 94.2% accuracy, setting a new benchmark on this dataset. We also analyze key technical trends and outline potential directions for future research in gait recognition.

cs.CV

Topological Metamaterial for Magnetic Resonance Imaging

Magnetic Resonance Imaging (MRI) is crucial in global healthcare, but the traditional receive coils, as a core component of MRI, SNR enhancement is limited due to the optimization of channel number and magnetic field strength faces high cost and complexity challenges. Here, we demonstrate the use of a topological material to enhance MRI signal reception. Designed with a stack of weak couplings, this material forms quasi-two-dimensional dual topological boundary states. High properties are achieved through low-loss signal transmission via these topological states, as well as only enhanced local magnetic fields and increased number of channels. Initial tests demonstrate superior performance and accessibility compared to commercial coils, suggesting significant potential. This concept introduces a transformative paradigm for all MRI coil designs.

physics.app-ph

A Comparative Evaluation of Large Vision-Language Models for 2D Object Detection under SOTIF Conditions

Reliable environmental perception remains one of the main obstacles for safe operation of automated vehicles. Safety of the Intended Functionality (SOTIF) concerns safety risks from perception insufficiencies, particularly under adverse conditions where conventional detectors often falter. While Large Vision-Language Models (LVLMs) demonstrate promising semantic reasoning, their quantitative effectiveness for safety-critical 2D object detection is underexplored. This paper presents a systematic evaluation of ten representative LVLMs using the PeSOTIF dataset, a benchmark specifically curated for long-tail traffic scenarios and environmental degradations. Performance is quantitatively compared against two specialized detectors: the anchor-based YOLOv5 and the transformer-based RT-DETRv4. Experimental results reveal a critical trade-off: top-performing LVLMs (e.g., Gemini 3) surpass the YOLOv5 in recall by over 25% and closely match RT-DETRv4 under natural visual degradation, while specialized detectors retain an advantage in geometric precision for handcrafted perturbations. These findings highlight the complementary strengths of semantic reasoning versus geometric regression, supporting the use of LVLMs as high-level safety validators in SOTIF-oriented automated driving systems.

cs.CV

Low-Latency and Low-Complexity MLSE for Short-Reach Optical Interconnects

To meet the high-speed, low-latency, and low-complexity demand for optical interconnects, simplified maximum likelihood sequence estimation (MLSE) is proposed in this paper. Simplified MLSE combines computational simplification and reduced state in MLSE. MLSE with a parallel sliding block architecture reduces latency from linear order to logarithmic order. Computational simplification reduces the number of multipliers from exponential order to linear order. Incorporating the reduced state with computational simplification further decreases the number of adders and comparators. The simplified MLSE is evaluated in a 112-Gbit/s PAM4 transmission over 2-km standard single-mode fiber. Experimental results show that the simplified MLSE significantly outperforms the FFE-only case in bit error ratio (BER) performance. Compared with simplified 1-step MLSE, the latency of simplified MLSE is reduced from 34 delay units in linear order to 7 delay units in logarithmic order. The simplified scheme in MLSE reduces the number of variable multipliers from 512 in exponential order to 33 in linear order without BER performance deterioration, while reducing the number of adders and comparators to 37.2% and 8.4%, respectively, with nearly identical BER performance.

cs.IT

A Survey on the Applications of Generative Artificial Intelligence in Automated Driving Systems Test Scenario Generation Methods

Ensuring the safety and reliability of Automated Driving Systems (ADS) remains a critical challenge, as traditional verification methods such as large-scale on-road testing are prohibitively costly and time-consuming.To address this,scenario-based testing has emerged as a scalable and efficient alternative,yet existing surveys provide only partial coverage of recent methodological and technological advances.This review systematically analyzes 31 primary studies,and 10 surveys identified through a comprehensive search spanning 2015~2025;however,the in-depth methodological synthesis and comparative evaluation focus primarily on recent frameworks(2023~2025),reflecting the surge of Artificial Intelligent(AI)-assisted and multimodal approaches in this period.Traditional approaches rely on expert knowledge,ontologies,and naturalistic driving or accident data,while recent developments leverage generative models,including large language models,generative adversarial networks,diffusion models,and reinforcement learning frameworks,to synthesize diverse and safety-critical scenarios.Our synthesis identifies three persistent gaps:the absence of standardized evaluation metrics,limited integration of ethical and human factors,and insufficient coverage of multimodal and Operational Design Domain (ODD)-specific scenarios.To address these challenges,this review contributes a refined taxonomy that incorporates multimodal extensions,an ethical and safety checklist for responsible scenario design,and an ODD coverage map with a scenario-difficulty schema to enable transparent benchmarking.Collectively,these contributions provide methodological clarity for researchers and practical guidance for industry,supporting reproducible evaluation and accelerating the safe deployment of higher-level ADS.

cs.SE

Training for Identity, Inference for Controllability: A Unified Approach to Tuning-Free Face Personalization

Tuning-free face personalization methods have developed along two distinct paradigms: text embedding approaches that map facial features into the text embedding space, and adapter-based methods that inject features through auxiliary cross-attention layers. While both paradigms have shown promise, existing methods struggle to simultaneously achieve high identity fidelity and flexible text controllability. We introduce UniID, a unified tuning-free framework that synergistically integrates both paradigms. Our key insight is that when merging these approaches, they should mutually reinforce only identity-relevant information while preserving the original diffusion prior for non-identity attributes. We realize this through a principled training-inference strategy: during training, we employ an identity-focused learning scheme that guides both branches to capture identity features exclusively; at inference, we introduce a normalized rescaling mechanism that recovers the text controllability of the base diffusion model while enabling complementary identity signals to enhance each other. This principled design enables UniID to achieve high-fidelity face personalization with flexible text controllability. Extensive experiments against six state-of-the-art methods demonstrate that UniID achieves superior performance in both identity preservation and text controllability. Code will be available at https://github.com/lyuPang/UniID

cs.CV

Photoinduced giant modulation of terahertz nonlinearity from metasurfaces

Active control of optical nonlinearity is essential for advancing next-generation electronics and photonics, including high-speed wireless communications, optical information processing, and nonlinear signal manipulation. However, achieving tunable nonlinearity at terahertz (THz) frequencies faces significant challenges due to the lack of materials that combine high nonlinear responses with strong sensitivity to external stimuli in this spectral regime. Here, we show giant modulation of THz nonlinearity by optically tailoring the valley degree of freedom in semiconductor-based metasurfaces. Mediated by the resonant behaviors of metasurfaces, photoexcited electrons transition into different valleys in the conduction band in response to the driving THz field, with the transition rate controlled by light intensity. Since THz nonlinearities vary significantly with electron dynamics in different valleys, various nonlinear effects-such as nonlinear transmission and generation-can be efficiently enhanced and modulated within a single metasurface using weak optical pumping. With optical energy as low as a few picojoules, we achieve on-off switching of THz third harmonic generation with a modulation depth exceeding 20000%, along with effective tunability of its nonperturbative behaviors. Our approach breaks new ground in active THz devices fully compatible with semiconductor industry standards, indicating a promising building block for ultrafast THz signal processing, all-optical computing, and nonlinear optical elements.

physics.optics

All-Electrical Self-Switching of van der Waals Chiral Antiferromagnet

Antiferromagnets have garnered significant attention due to their negligible stray field and ultrafast magnetic dynamics, which are promising for high-density and ultrafast spintronic applications. Their dual functionality as both spin sources and information carriers could enable all-electrical self-induced switching of antiferromagnetic order, offering great potential for ultra-compact spintronic devices. However, related progress is still elusive. Here, we report the deterministic switching of chiral antiferromagnetic orders induced by charge current at zero external magnetic field in the van der Waals (vdW) magnetically intercalated transition metal dichalcogenide CoTa3S6. This system exhibits strong interactions between cobalt atom magnetic moment lattice and itinerant electrons within the metallic layers, as demonstrated by temperature-dependent angle-resolved photoemission, scanning tunneling spectroscopy, and topological Nernst effect measurements. Notably, the itinerant-localization interactions lead to current-induced chiral spin orbit torques as well as Ruderman-Kittel-Kasuya-Yosida (RKKY) exchange torques that interact with the localized magnetic moments, facilitating all-electrical switching of the chiral magnetic order in the CoTa3S6 flake. Our work opens a promising avenue for manipulating antiferromagnetic orders by delicately engineering the synergistic interactions between magnetic moments and itinerant electrons.

cond-mat.mes-hall

Broadband spectral mapping of photo-induced second-harmonic generation in silicon nitride microresonators

By employing a pump-probe technique for enhanced spectral mapping of the dynamics in nonlinear frequency conversion, we demonstrate that photo-induced second-harmonic generation (SHG) in silicon nitride (Si3N4) microresonators can persist when transitioning from the preferred doubly resonant condition--where the resonances of the optical harmonics are required to be matched--to a highly detuned state where the generated second harmonic is significantly shifted away from its corresponding resonance. This results in an unconventionally broad conversion bandwidth. Other intriguing phenomena, such as detuning-dependent all-optical poling and nonlinear multi-mode interaction, are also presented for the first time with direct experimental evidence. Our findings provide new insights into the physics of photo-induced second-order (\chi^{(2)}) nonlinearity, highlighting its potential applications for nonlinear \chi^{(2)} photonics in integrated Si3N4 platform

physics.optics

An Efficient Deep Template Matching and In-Plane Pose Estimation Method via Template-Aware Dynamic Convolution

In industrial inspection and component alignment tasks, template matching requires efficient estimation of a target's position and geometric state (rotation and scaling) under complex backgrounds to support precise downstream operations. Traditional methods rely on exhaustive enumeration of angles and scales, leading to low efficiency under compound transformations. Meanwhile, most deep learning-based approaches only estimate similarity scores without explicitly modeling geometric pose, making them inadequate for real-world deployment. To overcome these limitations, we propose a lightweight end-to-end framework that reformulates template matching as joint localization and geometric regression, outputting the center coordinates, rotation angle, and independent horizontal and vertical scales. A Template-Aware Dynamic Convolution Module (TDCM) dynamically injects template features at inference to guide generalizable matching. The compact network integrates depthwise separable convolutions and pixel shuffle for efficient matching. To enable geometric-annotation-free training, we introduce a rotation-shear-based augmentation strategy with structure-aware pseudo labels. A lightweight refinement module further improves angle and scale precision via local optimization. Experiments show our 3.07M model achieves high precision and 14ms inference under compound transformations. It also demonstrates strong robustness in small-template and multi-object scenarios, making it highly suitable for deployment in real-time industrial applications. The code is available at:https://github.com/ZhouJ6610/PoseMatch-TDCM.

cs.CV

Quantum Optics Nature of the Elementary Excitations in Few-Layer WSe2 Semiconductors

A fully quantized description of a two-level system resonantly coupled with an electromagnetic field (light) is among the central topics of quantum electrodynamics, which is theorized by the quantum Rabi model. It is also a fundamental issue of light-matter interactions. The rapid development of two-dimensional (2D) transition-metal dichalcogenide (TMDC) atomically-thin semiconductors brings excellent great chance to test and demonstrate some basic predictions by quantum optics theory in textbook, i.e., by the quantum Rabi model. Such test and demonstration are of both scientific and technological significance because of the quick emergence of the second quantum revolution. In this Letter, we show a quantum optics demonstration of the variable-temperature optical responses of the elementary excitations in few-layer WSe2 flakes. It is unraveled that the variable-temperature reflectance and fluorescence patterns of the elementary excitations (i.e., the band-edge excitons) of monolayer, bilayer and hBN capped WSe2 match well with the predictions by the quantum Rabi model under the rotating wave approximation. Decoherence times, Rabi frequencies, and transition matrix elements of the elementary excitations in these few-layer WSe2 flakes are found to be all negatively correlated with temperature, and to show dependence on layer number and capping layer. These findings may provide a novel perspective for comprehending the fundamental quantum physical properties of two-dimensional materials.

physics.optics

Comparison of FTN-NOFDM and PCS-OFDM for Long-Haul Coherent Optical Communications

Single-wavelength 400G coherent optical communications have become a critical solution to meet the explosive traffic demands. However, the single-carrier modulation using low-order modulation formats requires a broader wavelength division multiplexing grid and expands the occupied optical bandwidth. In this paper, we propose the faster-than-Nyquist non-orthogonal frequency division multiplexing (FTN-NOFDM) to improve the spectral efficiency for long-haul coherent optical communications. The subcarrier number is set to eight to enable low-complexity FTN-NOFDM signal generation using a pruned inverse fast Fourier transform and inter-carrier interference (ICI) cancellation. To deal with the conventional timing recovery (TR) failure, a frequency tone-based TR is proposed for FTN-NOFDM. A time-domain multiple-input multiple-output equalizer is designed to update the tap coefficients based on outputs of conventional iterative detection (ID). To further mitigate ICI, a low-density parity check-assisted ID is integrated into the conventional ID module. FTN-NOFDM, probabilistic constellation shaping (PCS)-OFDM, and quadrature phase shift keying-OFDM are experimentally compared in a 400G coherent optical communication system over 11 cascaded 125-GHz wavelength-selective switches (WSSs) and 2000 km transmission. Results show that the FTN-NOFDM exhibits comparable WSS filtering tolerance to PCS-OFDM and superior nonlinearity tolerance, while PCS-OFDM achieves the best bit error ratio performance.

cs.NI