SearcharxivSearch

arXiv subjects

Yu Han

Publications and source records attributed to Yu Han.

At least 37 records · Page 2Linked to original sources

A Key Exchange Construction using Mihailova Subgroups in Braid groups

In this paper, we propose a modified Anshel-Anshel-Goldfeld (AAG) key exchange construction.The algebraic motivation underlying this construction comes from the membership problem for Mihailova subgroups of the braid group, a problem that is algorithmically unsolvable. We show that this perspective leads naturally to a quotient-group formulation involving Mihailova subgroups modulo the center of Bn. We also explain, however,that these algebraic facts do not by themselves provide a complete security proof for the protocol,because recovering a functionally equivalent conjugator modulo the center may already suffice for an adversary. Thus, the construction should be regarded as an algebraically motivated candidate whose full cryptographic security requires further study.

math.GR

Continuous-Variable Source-Independent Quantum Random Number Generation with General POVMs

Continuous-variable source-independent quantum random number generators offer the highest generation rates among semi-device-independent protocols. In reality, the protocol design is limited due to permissible measurement configurations. In this work, we propose a rigorous security proof framework that accommodates general, infinite-dimensional positive-operator-valued measures. Building upon the numerical security proof framework, we evaluate the randomness lower bound by maximizing the eavesdropper's guessing probability. Specifically, we transform the inherently infinite-dimensional semidefinite program in Fock space into a tractable finite-dimensional one, rigorously proving that latter provides a strict upper bound to the guessing probability of the original infinite-dimensional problem. Our framework showcases its capability by certifying secure randomness using unbalanced homodyne detection with only a single quadrature measurement, thereby bypassing the traditional requirement of measuring two conjugate quadratures such as $X$ and $P$. We experimentally validate our protocol on an optical platform using vacuum and weak coherent states, achieving a maximum secure randomness extraction of 1.11 bits per sample and an ultra-high generation rate of 1.776 Gbps. This work provides a flexible design for practical, high-speed quantum random number generators.

quant-ph

High-order tensor neural network for iteration-free structure relaxation

Structure relaxation is important for the discovery of new materials, yet conventional ab initio optimization remains a major bottleneck in high-throughput screening workflows. Machine learning potentials have accelerated relaxation by orders of magnitude, but they still rely on iterative optimization and high-quality DFT force labels. Here, we present HotRelax, a high-order tensor message-passing neural network for one-shot, end-to-end prediction of relaxed structures. Trained directly on paired unrelaxed and relaxed structures, HotRelax requires no DFT force labels and predicts relaxed structures in a single forward pass, without iterative inference or post-processing. Across five diverse datasets spanning 3D bulk crystals, 2D layered materials and catalysts, HotRelax shows strong performance relative to state-of-the-art end-to-end relaxation models, achieving lower prediction errors on several benchmarks while maintaining a compact model size and efficient inference. Extensive DFT calculations further show that the predicted structures are close in energy to their DFT-relaxed counterparts. When integrated into catalytic workflows, HotRelax also improves the accuracy and generalization of relaxed-state energy prediction models. Together, these results support HotRelax as an efficient and widely applicable framework for end-to-end structure relaxation, with strong potential to accelerate high-throughput materials discovery.

physics.comp-ph

Scaling Learning-based AEB with Massive Unlabeled Data

This paper studies how to scale learning-based automatic emergency braking (AEB) with massive unlabeled fleet data under production constraints. Our approach is based on meta-feedback semi-supervised learning (MF-SSL), where a teacher generates pseudo labels for unlabeled driving data and is updated using a small labeled anchor set as safety-critical feedback. In production, anchor ambiguity and labeled-unlabeled mismatch can amplify systematic pseudo-label errors, leading to spurious triggers. We propose a stabilized MF-SSL framework with (i) Noise-Aware Decoupling, which removes ambiguity-prone anchors from the teacher's supervised update path, and (ii) kinematics-gated pseudo-labeling with a teacher conflict penalty to suppress mismatch-induced risk hallucinations on unlabeled data while maintaining broad coverage. Extensive experiments show consistent gains as unlabeled data scale from 1M to 1B windows, improving safety while keeping comfort stable. The 1B-trained student model is deployed to hundreds of thousands of vehicles and validated over \$10^9$ km of driving, achieving a positive-to-false activation ratio exceeding 100:1 and a 35% improvement in accident-free driving mileage over a production rule-only baseline.

cs.LG

Digital Twin-Based Channel Generation Toolchain and Foundation Model for Low-Altitude XL-MIMO

The rapid development of the low-altitude economy (LAE) has created growing demand for reliable aerial communication systems. Extremely large-scale multiple-input multiple-output (XL-MIMO) is a promising enabler for such systems due to its high spatial resolution and robust connectivity. However, three-dimensional (3D) mobility together with near-field propagation makes it difficult to obtain dedicated high-fidelity wireless datasets, hindering systematic algorithm development and evaluation. To address this issue, we develop LAETwin-XL, a digital twin (DT)-based toolchain and dataset for XL-MIMO research in LAE scenarios. Built on the Sionna ray-tracing (RT) module, the proposed toolchain simulates near-field and far-field channels with diverse wireless labels for practical environments. Building on this dataset, we further develop a conditional denoising diffusion implicit model (CDDIM)-based generative foundation model that is pretrained to learn transferable XL-MIMO channel representations from incomplete channel observations. Unlike conventional task-specific or foundation models that rely on relatively complete channel inputs, the proposed model can generatively infer informative channel representations from partially observed channels. Experimental results demonstrate that the proposed framework achieves effective zero-shot channel extrapolation performance. Furthermore, using lightweight task heads and limited training data, it enables parameter-efficient transfer to various downstream tasks (e.g., channel estimation, classification, and localization), delivering high accuracy and robustness even under sparse antenna observations. The codes and dataset are available at https://github.com/Lmyxxn/LAETwin-XL.

eess.SP

XL-ChannelDiff: An Efficient Diffusion-Based Multi-Domain Near-Field Channel Extrapolation Framework for XL-MIMO Systems

Accurate channel state information (CSI) acquisition is essential for unleashing the performance gains of extremely large-scale multiple-input multiple-output (XL-MIMO) systems. However, in near-field regions, CSI acquisition is much more challenging than in the far field due to the high-dimensional channel representation and spherical wavefront propagation. To address this, in this paper, we propose an efficient multi-domain near-field channel extrapolation framework for XL-MIMO systems. Leveraging the conditional denoising diffusion implicit model (CDDIM), our approach enables accurate channel extrapolation across the antenna, frequency, and spatial domains. Specifically, we design a physics-aware CDDIM backbone that incorporates position-embedded patch tokenization and a mask-guided multi-head attention mechanism, enabling the model to exploit position-dependent channel correlations induced by near-field spherical-wave propagation. To ensure high-fidelity extrapolation, we incorporate a Wasserstein GAN (WGAN) discriminator that provides adversarial supervision to the CDDIM during both the training and reverse sampling phases. Additionally, a RePaint-style refinement scheme is introduced to optimize the sampling trajectory, further boosting extrapolation accuracy. Extensive experiments demonstrate the superiority of the proposed framework, achieving superior extrapolation accuracy and robust generalization across diverse domains, varied configurations, and severe masking conditions.

eess.SP

Vision-Based Efficient Joint Trajectory and Channel Tracking in Near-Field XL-MIMO Systems

Accurate joint tracking of mobile users, surrounding scatterers, and dynamic channels is a critical task for sixth-generation (6G) wireless systems, essential for both ensuring high-quality communications and empowering advanced selsing applications such as autonomous driving and immersive extended reality. While extremely large-scale multiple-input multiple-output (XL-MIMO) inherently offers strong support for this task through its high spatial resolution and spectral efficiency, its massive scale of antenna arrays, coupled with near-field propagation characteristics, makes joint trajectory and channel tracking time-consuming and hardware-intensive. To address these challenges, we rethink the problem from a vision-based signal perspective. Specifically, we design a subarray-based partially connected hybrid beamforming (PC-HBF) architecture with a tailored time-multiplexed (TM) mechanism. This effectively compensates for the aperture loss caused by limited radio frequency (RF) chains, generating high-fidelity Cartesian-domain signal images that inherently capture near-field spatial features. Based on this visual representation, we propose an improved CenterNet to perform accurate one-shot path localization, circumventing the path-iterative search required by conventional compressed-sensing-based methods. Building upon this to further improve the accuracy and exploit temporal correlation, a local small-scale orthogonal matching pursuit (OMP) refiner and a lightweight cascaded OMP tracker are developed. Finally, a Hungarian-based trajectory association module is incorporated to maintain track continuity and provide trajectory-level information for environment monitoring. Simulation results show that the proposed framework consistently outperforms representative baselines in position and channel tracking accuracy, especially under low-SNR and limited-hardware conditions.

eess.SP

3SPO: State-Score-Supervised Policy Optimization for LLM Agents

Training large language models (LLMs) as autonomous agents via reinforcement learning (RL) has enabled frontier models to achieve superhuman performance in long-horizon tasks. However, existing RL algorithms operate at the trajectory level, performing policy optimization only after collecting complete episode rollouts. This coarse-grained approach faces fundamental challenges in multi-turn agent settings where rewards are sparse, delayed, and credit assignment across individual steps is critical. In this work, we propose \textbf{State-Score-Supervised Policy Optimization (3SPO)}, a novel RL algorithm that performs post-step policy optimization with dynamic state score supervision. At each step, 3SPO computes the state score based on historical success rates, supervising step-wise credit assignment, adaptive rollout and post-step policy optimization without requiring value function estimation or additional auxiliary models. Theoretically, under a per-state bandit abstraction, we show that the proposed score-supervised allocation mechanism achieves logarithmic allocation regret and provide sample-complexity guarantees for action identification, score distinguishability, and filtering stability. Experiments on ALFWorld and WebShop with Qwen2.5-1.5B/7B-Instruct show that 3SPO consistently outperforms GRPO by $+22.6\%$ on ALFWorld and $+15.6$ points on WebShop, while using comparable resources to achieve $2.4\times$ more state exploration and $1.8\times$ faster convergence. Code is available at https://github.com/genalyu/3SPO.

cs.LG

Hybrid-plasticity Photonic Synapses Enabling Hardware-Level Neural Reuse

Biological intelligence is distinguished by neural reuse, the capacity to preserve established learning memory while repurposing it for new tasks and dynamic environments. Bringing this capability to photonic hardware requires hybrid plasticity, namely the coexistence of long-term synaptic plasticity for persistent weight storage and short-term synaptic plasticity for rapid, reversible adaptation within a single synaptic element; however, current photonic architectures lack such a unified mechanism. Here, we demonstrate a hybrid-plasticity photonic synapse on thin-film lead zirconate titanate (PZT) that couples non-volatile and volatile modes to enable hardware-level neural reuse. Crucially, high-speed refresh operations can be superimposed without perturbing the stored weight. Such a neural-reuse framework yields a convergence speedup of over 20-fold and reduces the weight updates by approximately 30-fold compared with random initialization. These results establish hybrid-plasticity photonic synapses as a pathway toward on-chip learning systems that are both memory-preserving and rapidly adaptable.

physics.optics

NF-TrackLLM: Joint Prediction of UAV Trajectory and Near-Field Beam for LAE XL-MIMO Systems

User localization and beam management are tightly linked in extremely large-scale multiple-input multiple-output (XL-MIMO) systems, especially in dense low-altitude economy (LAE) scenarios. However, the near-field propagation in XL-MIMO introduces strong distance sensitivity and complex spatial coupling, which makes joint trajectory and beam prediction challenging. Meanwhile, large language models (LLMs) have attracted attention in physical-layer transmission for modeling long-range dependencies. In this paper, we propose NF-TrackLLM, a multi-modal semantic-aware framework for near-field unmanned aerial vehicles (UAVs) positioning and beam prediction in XL-MIMO systems. By incorporating visual and LiDAR sensing into a Sionna-based channel generation pipeline, environmental semantics and GPS are utilized to guide trajectory and beam prediction. Built upon the aligned multi-modal representation, a GPT-2-based spatiotemporal reasoning backbone, and a cascaded prediction strategy are employed, where future trajectories are first inferred and then used to guide beam prediction as geometric priors. Simulation results demonstrate that NF-TrackLLM achieves accurate beam prediction and reliable UAV trajectory tracking in dense urban low-altitude scenarios.

eess.SP

SGAgent: Suggestion-Guided LLM-Based Multi-Agent Framework for Repository-Level Software Repair

Large Language Models (LLMs) have enabled intelligent agents that autonomously interact with environments and invoke external tools. Recently, agent-based software repair has drawn wide attention, as repair agents can localize bugs, generate patches, and achieve state-of-the-art performance on repository-level benchmarks (e.g., SWE-Bench). However, existing approaches usually adopt a localize-then-fix paradigm, jumping directly from "where the bug is" to "how to fix it", leaving a fundamental reasoning gap. To this end, we propose SGAgent, a Suggestion-Guided multi-Agent framework for repository-level software repair, which follows a localize-suggest-fix paradigm. SGAgent introduces a suggestion phase to strengthen the transition from localization to repair: the suggester starts from the buggy locations, incrementally retrieves relevant context until it fully understands the bug, and provides actionable repair suggestions. We further construct a Knowledge Graph (KG) from the target repository and develop a KG-based toolkit to strengthen SGAgent's global contextual awareness and repository-level reasoning. Three specialized sub-agents (i.e., localizer, suggester, and fixer) collaborate to achieve automated end-to-end software repair. We evaluate SGAgent on SWE-Bench-Lite. SGAgent with Claude-3.5 achieves 51.3% repair accuracy, 81.2% file-level, and 52.4% function-level localization accuracy at an average cost of $1.48 per instance, outperforming all baselines using the same base model. SGAgent also generalizes well across base LLMs, reaching a 60.7% resolution rate with Claude-4. When extended to vulnerability repair, it achieves 48.0% on VUL4J and VJBench, demonstrating strong generalization across tasks and programming languages.

cs.SE

Multimodal-NF: A Wireless Dataset for Near-Field Low-Altitude Sensing and Communications

Environment-aware 6G wireless networks demand the deep integration of multimodal and wireless data. However, most existing datasets are confined to 2D terrestrial far-field scenarios, lacking the 3D spatial context and near-field characteristics crucial for low-altitude extremely large-scale multiple-input multiple-output (XL-MIMO) systems. To bridge this gap, this letter introduces Multimodal-NF, a large-scale dataset and specialized generation framework. Operating in the upper midband, it synchronizes high-fidelity near-field channel state information (CSI) and precise wireless labels (e.g., Top-5 beam indices, LoS/NLoS) with comprehensive sensory modalities (RGB images, LiDAR point clouds, and GPS). Crucially, these multimodal priors provide spatial semantics that help reduce the near-field search space and thereby lower the overhead of wireless sensing and communication tasks. Finally, we validate the dataset through representative case studies, demonstrating its utility and effectiveness. The open-source generator and dataset are available at https://lmyxxn.github.io/6GXLMIMODatasets/.

eess.SP

Distributed Near-Field Channel Estimation for U6G XL-MIMO Systems under Beam Squint

Since the beam squint and near-field effects both inherently exist in upper-6 GHz (U6G) extremely large-scale multiple-input multiple-output (XL-MIMO) systems, wideband near-field channel estimation faces severe challenges, such as higher computational complexity, and higher pilot overhead particularly at hybrid architectures with fewer radio frequency (RF) chains. To precisely reduce the complexity and number of pilots, the parametric symmetry of wideband near-field channels is explored, such that the channel parameters, including angle, distance, and range, can be decoupled based on the delay variations observed by different antennas. Based on this, a distributed parametric symmetry-based (DPS) algorithm, applicable to U6G XL-MIMO, is proposed. The delays observed by different subarrays are estimated and extrapolated across the local processing units (LPUs) firstly, and then, the channel parameters are decoupled and estimated at the central processing unit (CPU), by only linearly combining the delays from different LPUs. The path gains are calculated at different LPUs, respectively, to reconstruct the channel with low complexity. Since the proposed algorithm does not rely on scanning the polar-domain dictionary, only a single pilot is required even with hybrid architectures. Furthermore, the computational complexity, multiple-path resolution, Cramer-Rao lower bound (CRLB) and lower bound (LB) of the estimates in hybrid architectures and the DPS algorithm, respectively, are analyzed, to evaluate the realizable potential of the proposed algorithm. The simulation results prove that the proposed algorithm has a higher estimation accuracy, while requiring less complexity and pilots.

eess.SP

Analytical Framework of Airy Beams in Near-Field XL-MIMO: From Ideal Optics to Wireless Reality

The synthesis of Airy-profiled wavefronts has emerged as a pivotal paradigm for advanced electromagnetic engineering, attributed to their intrinsic non-diffractive propagation, transverse self-acceleration, and structural self-healing properties. While the advent of extremely large-scale multiple-input multiple-output (XL-MIMO) and the elevation in frequency bands for sixth generation wireless systems provide the physical foundation for generating such structured beams, their wireless realization is fundamentally governed by hybrid precoding architectures, finite array apertures, and discrete antenna topologies. These constraints induce significant deviations from ideal optical Airy beam models, necessitating a rigorous re-characterization of Airy beams in practical wireless contexts. Consequently, this paper establishes an analytical theoretical framework to explicitly characterize Airy beam propagation in near-field XL-MIMO and derives the constraints on array aperture and antenna spacing to sustain distortion-free main lobe trajectories. Furthermore, quantitative metrics are developed to rigorously evaluate the performance trade-offs between Airy beams and Gaussian focusing beams, thereby providing systematic guidelines for their deployment in scenario-dependent wireless applications. Numerical results corroborate the proposed analytical theoretical framework of Airy beams in near-field XL-MIMO, and demonstrate the potential to achieve robust communication and spectral efficiency (SE) improvement in certain scenarios.

eess.SP

Wideband Precoding for U6G XL-MIMO Systems: Beam Squint Boundaries and Channel Slicing

The unconventionally large aperture of extremely large-scale multiple-input multiple-output (XL-MIMO) arrays, in conjunction with the wider bandwidths in the upper-6 GHz (U6G) frequency bands, will very likely lead to non-negligible beam squint effects. In the context of a limited number of radio frequency (RF) chains, and by adopting hybrid precoding, the beams at different subcarriers may point to different positions and compromise the spectral efficiency (SE). Moreover, the existence of \textit{multiple paths} in U6G XL-MIMO channels also entails practical challenges for wideband precoding. It is therefore essential to ascertain whether the beam squint effect is pronounced for U6G XL-MIMO systems and design efficient wideband precoding schemes. To address these challenges, precise antenna-domain and frequency-domain wideband boundaries are derived from the near-field and far-field perspectives, respectively. These boundaries can inform the design of wideband precoding in future system settings. Subsequently, a channel slicing scheme is proposed for wideband precoding. The process involves the segmentation of U6G XL-MIMO channels into multiple blocks, with the objective of mitigating the beam squint effect for each path. The antenna-domain and frequency-domain slicing methods are developed for multipath and multiuser scenarios, respectively. The simulation results prove that the beam squint effect remains a significant issue for U6G XL-MIMO systems, while the near-field effect invariably precedes the beam squint effect as the array size and the bandwidth increase. In addition, the proposed scheme can greatly improve the SE.

eess.SP

GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference

Sparse Mixture of Experts (SMoE) enables scalable parameter growth in large language models (LLMs) by selectively activating a subset of experts, and its large parameter count necessitates distributed deployment for inference. However, distributed inference faces a critical dilemma: although communication overhead constitutes the primary bottleneck, reducing it often exacerbates computational load imbalance, leading to resource waste. In this paper, we present GRACE-MoE, which stands for Grouping and Replication with Locality-Aware Routing for SMoE inference. GRACE-MoE is a lossless co-optimization framework that integrates expert grouping to reduce communication and dynamic replication to correct load skew, together with locality-aware routing to resolve replica selection. To underpin this coordinated optimization in multi-node settings, GRACE-MoE adopts a hierarchical sparse communication design that reduces cross-node traffic while implicitly aligning execution across nodes, thereby mitigating synchronization overhead. Experiments on diverse models and multi-node, multi-GPU environments demonstrate that GRACE-MoE efficiently reduces end-to-end inference latency, achieving up to 4.66x speedup over existing systems, and the code will be released upon acceptance.

cs.DC

Capacity drop accounting for microscopic vehicle interaction effects: analytical model and validation with high-resolution trajectories

Capacity drop is a traffic phenomenon in which the discharge flow from a queue is lower than the theoretical infrastructure capacity. This paper proposes a generic analytical method to estimate the queue discharge flow of freeway traffic. Capacity drop is primarily attributed to hesitant vehicles, defined as vehicles that stochastically and temporarily enter an acceleration delay state and generate voids (i.e., extra gaps) in front of them. The proposed method estimates the expected total void length generated by all hesitant vehicles, based on the distributions of their spatial and temporal locations as well as the associated delays. It also accounts for interactions between the waves triggered by downstream hesitant vehicles and the voids generated by upstream ones. Our analysis reveals that this interaction is the key mechanism behind the differing extents of capacity drop observed between standing queues and jam waves in previous studies. The accuracy of the model is validated through both numerical simulations and real-world trajectories. Overall, the proposed method offers a deeper understanding of capacity drop, which can be leveraged in traffic flow modeling and control.

physics.soc-ph

The Swarm Intelligence Freeway-Urban Trajectories (SWIFTraj) Dataset -- Part II: A Graph-Based Approach for Trajectory Connection

In Part I of this companion paper series, we introduced SWIFTraj, a new open-source vehicle trajectory dataset collected using a unmanned aerial vehicle (UAV) swarm. The dataset has two distinctive features. First, by connecting trajectories across consecutive UAV videos, it provides long-distance continuous trajectories, with the longest exceeding 4.5 km. Second, it covers an integrated traffic network consisting of both freeways and their connected urban roads. Obtaining such long-distance continuous trajectories from a UAV swarm is challenging, due to the need for accurate time alignment across multiple videos and the irregular spatial distribution of UAVs. To address these challenges, this paper proposes a novel graph-based approach for connecting vehicle trajectories captured by a UAV swarm. An undirected graph is constructed to represent flexible UAV layouts, and an automatic time alignment method based on trajectory matching cost minimization is developed to estimate optimal time offsets across videos. To associate trajectories of the same vehicle observed in different videos, a vehicle matching table is established using the Hungarian algorithm. The proposed approach is evaluated using both simulated and real-world data. Results from real-world experiments show that the time alignment error is within three video frames, corresponding to approximately 0.1 s, and that the vehicle matching achieves an F1-score of about 0.99. These results demonstrate the effectiveness of the proposed method in addressing key challenges in UAV-based trajectory connection and highlight its potential for large-scale vehicle trajectory collection.

physics.soc-ph