SearcharxivSearch

arXiv subjects

Yueheng Li

Publications and source records attributed to Yueheng Li.

At least 19 recordsLinked to original sources

Guided Policy Optimization under Partial Observability

Reinforcement Learning (RL) in partially observable environments poses significant challenges due to the complexity of learning under uncertainty. While additional information, such as that available in simulations, can enhance training, effectively leveraging it remains an open problem. To address this, we introduce Guided Policy Optimization (GPO), a framework that co-trains a guider and a learner. The guider takes advantage of privileged information while ensuring alignment with the learner's policy that is primarily trained via imitation learning. We theoretically demonstrate that this learning scheme achieves optimality comparable to direct RL, thereby overcoming key limitations inherent in existing approaches. Empirical evaluations show strong performance of GPO across various tasks, including continuous control with partial observability and noise, and memory-based challenges, significantly outperforming existing methods.

cs.LG

Multi-Agent Guided Policy Optimization

Due to practical constraints such as partial observability and limited communication, Centralized Training with Decentralized Execution (CTDE) has become the dominant paradigm in cooperative Multi-Agent Reinforcement Learning (MARL). However, existing CTDE methods often underutilize centralized training or lack theoretical guarantees. We propose Multi-Agent Guided Policy Optimization (MAGPO), a novel framework that better leverages centralized training by integrating centralized guidance with decentralized execution. MAGPO uses an autoregressive joint policy for scalable, coordinated exploration and explicitly aligns it with decentralized policies to ensure deployability under partial observability. We provide theoretical guarantees of monotonic policy improvement and empirically evaluate MAGPO on 43 tasks across 6 diverse environments. Results show that MAGPO consistently outperforms strong CTDE baselines and matches or surpasses fully centralized approaches, offering a principled and practical solution for decentralized multi-agent learning. Our code and experimental data can be found in https://github.com/liyheng/MAGPO.

cs.AI

Analysis of Sensing in OFDM-based ISAC under the Influence of Sampling Jitter

To enable integrated sensing and communication (ISAC) in cellular networks, a wide range of additional requirements and challenges are either imposed or become more critical. One such impairment is sampling jitter (SJ), which arises due to imperfections in the sampling instants of the clocks of digital-to-analog converters (DACs) and analog-to-digital converters (ADCs). While SJ is already well studied for communication systems based on orthogonal frequency-division multiplexing (OFDM), which is expected to be the waveform of choice for most sixth-generation (6G) scenarios where ISAC could be possible, the implications of SJ on the OFDM-based radar sensing must still be thoroughly analyzed. Considering that phase-locked loop (PLL)-based oscillators are used to derive sampling clocks, which leads to colored SJ, i.e., SJ with non-flat power spectral density, this article analyzes the resulting distortion of the adopted digital constellation modulation and sensing performance in OFDM-based ISAC for both baseband (BB) and bandpass (BP) sampling strategies and different oversampling factors. For BB sampling, it is seen that SJ induces intercarrier interference (ICI), while for BP sampling, it causes carrier phase error and more severe ICI due to a phase noise-like effect at the digital intermediate frequency. Obtained results for a single-input single-output OFDM-based ISAC system with various OFDM signal parameterizations demonstrate that SJ-induced degradation becomes non-negligible for both BB and BP sampling only for root mean square (RMS) SJ values above 10^-11 s at both DAC and ADC, which corresponds to 0.5*10^-2 times the considered critical sampling period without oversampling. Based on the achieved results, it can be concluded that state-of-the-art hardware enables sufficient communication and sensing robustness against SJ, as RMS SJ values in the femtosecond range can be achieved.

eess.SP

Joint BS Deployment and Power Optimization for Minimum EMF Exposure with RL in Real-World Based Urban Scenario

Base station (BS) deployment remains a critical task with successive wireless communication generations and increasing data rates demands, while the electromagnetic field (EMF) exposure is often underrated, yielding potential health implications. Therefore, this paper proposes a workflow that adjusts BS deployment and radiated power in a 3D urban scenario to jointly consider EMF exposure and coverage. To achieve this ambition, firstly, a novel least-time shoot-and-bounce ray (SBR) ray-launching (RL) tool is developed to improve computational efficiency, and simultaneously enhance diffraction modeling} for accurate EMF exposure calculation, validated with real-world measurements. To efficiently extend the computation across the target urban area, the adaptive grid refinement (AGR) algorithm is designed based on the spatial stability of the effective channel while accounting for BS beamforming, enabling global estimation of EMF exposure and signal coverage. Subsequently, to better represent real-world communication network behaviors, the actual maximum transmit power, intercell interference, and channel state information imperfection are incorporated on the BS side, while mobility over the EMF exposure averaging interval is captured on the user equipment side. Upon the aforementioned aspects, the coverage-guaranteed EMF exposure minimization problem is formulated in a realistic and accurate manner, and solved by a geometry-aware algorithm adapted to deterministic channel models, yielding the optimal BS deployment and power configuration. In comparison to a baseline that relies on an empirical channel model, the proposed workflow delivers more reliable estimation of EMF exposure and provides practical guidance for BS construction and operations.

eess.SP

Facet Specific Electron Conduction in Pentavalent (W5+) WO3 Drives Superior Photocatalytic CO 2 Reduction in (002) Plane

This article reports a concept of heat-induced topological modifications of non-layered WO 3 followed by successful synthesis of oxygen-vacant more-porous nanosheets with exposed active (002) facet. Experimental measurements and Density Functional Theory (DFT) calculations have revealed that the photoexcited electrons are found to accumulate preferentially on (002) facet to yield enhanced electron conduction, and consequently, strengthen the reduction potential as active catalytic sites for photocatalytic CO2 reduction. Owing to these beneficial properties, the more-porous nanosheets of WO 3 with (002) facet have exhibited superior performance than that of less-porous nanosheets of WO3 with (220) facet and bulk WO3 with (205) facet. This study therefore provides a new understanding of regulating physical, optical, and electronic properties through intricate atomic structure modulation of WO3, and may find widespread application in optoelectronics, sensors, and energy conversion.

cond-mat.mtrl-sci

Distributed Specialization: Rare-Token Neurons in Large Language Models

Large language models (LLMs) struggle with representing and generating rare tokens despite their importance in specialized domains. We investigate whether LLMs develop internal specialization mechanisms through discrete modular architectures or distributed parameter-level differentiation. Through systematic analysis of final-layer MLP neurons across multiple model families, we discover that rare-token processing emerges via \textit{distributed specialization}: functionally coordinated but spatially distributed subnetworks that exhibit three distinct organizational principles. First, we identify a reproducible three-regime influence hierarchy comprising highly influential plateau neurons(also termed as rare-token neurons), power-law decay neurons, and minimally contributing neurons, which is absent in common-token processing. Second, plateau neurons demonstrate coordinated activation patterns (reduced effective dimensionality) while remaining spatially distributed rather than forming discrete clusters. Third, these specialized mechanisms are universally accessible through standard attention pathways without requiring dedicated routing circuits. Training dynamics reveal that functional specialization emerges gradually through parameter differentiation, with specialized neurons developing increasingly heavy-tailed weight correlation spectra consistent with Heavy-Tailed Self-Regularization signatures. Our findings establish that LLMs process rare-tokens through distributed coordination within shared architectures rather than mixture-of-experts-style modularity. These results provide insights for interpretable model editing, computational efficiency optimization, and understanding emergent functional organization in transformer networks.

cs.AI

Multi-RIS Deployment Optimization for mmWave ISAC Systems in Real-World Environments

Reconfigurable intelligent surface-assisted integrated sensing and communication (RIS-ISAC) presents a promising system architecture to leverage the wide bandwidth available at millimeter-wave (mmWave) frequencies, while mitigating severe signal propagation losses and reducing infrastructure costs. To enhance ISAC functionalities in the future air-ground integrated network applications, RIS deployment must be carefully designed and evaluated, which forms the core motivation of this paper. To ensure practical relevance, a multi-RIS-ISAC system is established, with its signal model at mmWave frequencies demonstrated using ray-launching calibrated to real-world environments. On this basis, an energy-efficiency-driven optimization problem is formulated to minimize the multi-RIS size-to-coverage sum ratio, comprehensively considering real-world RIS deployment constraints, positions, orientations, as well as ISAC beamforming strategies at both the base station and the RISs. To solve the resulting non-convex mixed-integer problem, a simplified reformulation based on equivalent gain scaling method is introduced. A two-step iterative algorithm is then proposed, in which the deployment parameters are determined under fixed RIS positions in the first step, and the RIS position set is updated in the second step to progressively approach the optimum solution. Simulation results based on realistic parameter benchmarks present that the optimized RISs deployment significantly enhances communication coverage and sensing accuracy with the minimum RIS sizes, outperforming existing approaches.

eess.SP

Emergent Specialization: Rare Token Neurons in Language Models

Large language models struggle with representing and generating rare tokens despite their importance in specialized domains. In this study, we identify neuron structures with exceptionally strong influence on language model's prediction of rare tokens, termed as rare token neurons, and investigate the mechanism for their emergence and behavior. These neurons exhibit a characteristic three-phase organization (plateau, power-law, and rapid decay) that emerges dynamically during training, evolving from a homogeneous initial state to a functionally differentiated architecture. In the activation space, rare token neurons form a coordinated subnetwork that selectively co-activates while avoiding co-activation with other neurons. This functional specialization potentially correlates with the development of heavy-tailed weight distributions, suggesting a statistical mechanical basis for emergent specialization.

cs.AI

System Concept and Demonstration of Bistatic MIMO-OFDM-based ISAC

In future sixth-generation (6G) mobile networks, radar sensing is expected to be offered as an additional service to its original purpose of communication. Merging these two functions results in integrated sensing and communication (ISAC) systems. In this context, bistatic ISAC appears as a possibility to exploit the distributed nature of cellular networks while avoiding highly demanding hardware requirements such as full-duplex operation. Recent studies have introduced strategies to perform required synchronization and data exchange between nodes for bistatic ISAC operation, based on orthogonal frequency-division multiplexing (OFDM), however, only for single-input single-output architectures. In this article, a system concept for a bistatic multiple-input multiple-output (MIMO)-OFDM-based ISAC system with beamforming at both transmitter and receiver is proposed, and a distribution synchronization concept to ensure coherence among the different receive channels for direction-of-arrival estimation is presented. After a discussion on the ISAC processing chain, including relevant aspects for practical deployments such as transmitter digital pre-distortion and receiver calibration, a 4x8 MIMO measurement setup at 27.5 GHz and results are presented to validate the proposed system and distribution synchronization concepts.

eess.SP

On the Sensing Performance of OFDM-based ISAC under the Influence of Oscillator Phase Noise

Integrated sensing and communication (ISAC) is a novel capability expected for sixth generation (6G) cellular networks. To that end, several challenges must be addressed to enable both mono- and bistatic sensing in existing deployments. A common impairment in both architectures is oscillator phase noise (PN), which not only degrades communication performance, but also severely impairs radar sensing. To enable a broader understanding of orthogonal-frequency division multiplexing (OFDM)-based sensing impaired by PN, this article presents an analysis of sensing peformance in OFDM-based ISAC for different waveform parameter choices and settings in both mono- and bistatic architectures. In this context, the distortion of the adopted digital constellation modulation is analyzed and the resulting PN-induced effects in range-Doppler radar images are investigated both without and with PN compensation. These effects include peak power loss of target reflections and higher sidelobe levels, especially in the Doppler shift direction. In the conducted analysis, these effects are measured by the peak power loss ratio, peak-to-sidelobe level ratio, and integrated sidelobe level ratio parameters, the two latter being evaluated in both range and Doppler shift directions. In addition, the signal-to-interference ratio is analyzed to allow not only quantifying the distortion of a target reflection, but also measuring the interference floor level in a radar image. The achieved results allow to quantify not only the PN-induced impairments to a single target, but also how the induced degradation may impair the sensing performance of OFDM-based ISAC systems in multi-target scenarios.

eess.SP

Pilot-Based SFO Estimation for Bistatic Integrated Sensing and Communication

Enabling bistatic radar sensing within the context of integrated sensing and communication (ISAC) for future sixth generation mobile networks demands strict synchronization accuracy, which is particularly challenging to be achieved with over-the-air synchronization. Existing algorithms handle time and frequency offsets adequately, but provide insufficiently accurate sampling frequency offset (SFO) estimates that result in degradation of obtained radar images in the form of signal-to-noise ratio loss and migration of range and Doppler shift. This article introduces an SFO estimation algorithm named tilt inference of time offset (TITO) for orthogonal frequency-division multiplexing (OFDM)-based ISAC. Using available pilot subcarriers, TITO obtains channel impulse response estimates and extracts information on the SFO-induced delay migration to a dominant reference path with constant range, Doppler shift, and angle between transmit and receive ISAC nodes. TITO then adaptively selects the delay estimates that are only negligibly impaired by SFO-induced intersymbol interference, ultimately employing them to estimate the SFO. Assuming a scenario without a direct line-of-sight (LoS) between the aforementioned transmitting and receiving ISAC nodes, a system concept with a relay reflective intelligent surface (RIS) is used to create the aforementioned reference path is proposed. Besides a mathematical derivation of accuracy bounds, simulation and measurements at 26.2 GHz are presented to demonstrate TITO's superiority over existing methods in terms of SFO estimation accuracy and robustness.

eess.SP

Bistatic OFDM-based ISAC with Over-the-Air Synchronization: System Concept and Performance Analysis

Integrated sensing and communication (ISAC) has been defined as one goal for 6G mobile communication systems. In this context, this article introduces a bistatic ISAC system based on orthogonal frequency-division multiplexing (OFDM). While the bistatic architecture brings advantages such as not demanding full duplex operation with respect to the monostatic one, the need for synchronizing transmitter and receiver is imposed. In this context, this article introuces a bistatic ISAC signal processing framework where an incoming OFDM-based ISAC signal undergoes over-the-air synchronization based on preamble symbols and pilots. Afterwards, bistatic radar processing is performed using either only pilot subcarriers or the full OFDM frame. The latter approach requires estimation of the originally transmitted frame based on communication processing and therefore error-free communication, which can be achieved via appropriate channel coding. The performance and limitations of the introduced system based on both aforementioned approaches are assessed via an analysis of the impact of residual synchronization mismatches and data decoding failures on both communication and radar performances. Finally, the performed analyses are validated by proof-of-concept measurement results.

eess.SP

Breast Cancer Immunohistochemical Image Generation: a Benchmark Dataset and Challenge Review

For invasive breast cancer, immunohistochemical (IHC) techniques are often used to detect the expression level of human epidermal growth factor receptor-2 (HER2) in breast tissue to formulate a precise treatment plan. From the perspective of saving manpower, material and time costs, directly generating IHC-stained images from Hematoxylin and Eosin (H&E) stained images is a valuable research direction. Therefore, we held the breast cancer immunohistochemical image generation challenge, aiming to explore novel ideas of deep learning technology in pathological image generation and promote research in this field. The challenge provided registered H&E and IHC-stained image pairs, and participants were required to use these images to train a model that can directly generate IHC-stained images from corresponding H&E-stained images. We selected and reviewed the five highest-ranking methods based on their PSNR and SSIM metrics, while also providing overviews of the corresponding pipelines and implementations. In this paper, we further analyze the current limitations in the field of breast cancer immunohistochemical image generation and forecast the future development of this field. We hope that the released dataset and the challenge will inspire more scholars to jointly study higher-quality IHC-stained image generation.

eess.IV

Improving Adaptive Real-Time Video Communication Via Cross-layer Optimization

Effective Adaptive BitRate (ABR) algorithm or policy is of paramount importance for Real-Time Video Communication (RTVC) amid this pandemic to pursue uncompromised quality of experience (QoE). Existing ABR methods mainly separate the network bandwidth estimation and video encoder control, and fine-tune video bitrate towards estimated bandwidth, assuming the maximization of bandwidth utilization yields the optimal QoE. However, the QoE of a RTVC system is jointly determined by the quality of compressed video, fluency of video playback, and interaction delay. Solely maximizing the bandwidth utilization without comprehensively considering compound impacts incurred by both network and video application layers, does not assure the satisfactory QoE. And the decoupling of network and video layer further exacerbates the user experience due to network-codec incoordination. This work therefore proposes the Palette, a reinforcement learning based ABR scheme that unifies the processing of network and video application layers to directly maximize the QoE formulated as the weighted function of video quality, stalling rate and delay. To this aim, a cross-layer optimization is proposed to derive fine-grained compression factor of upcoming frame(s) using cross-layer observations like network conditions, video encoding parameters, and video content complexity. As a result, Palette manages to resolve the network-codec incoordination and to best catch up with the network fluctuation. Compared with state-of-the-art schemes in real-world tests, Palette not only reduces 3.1%-46.3% of the stalling rate, 20.2%-50.8% of the delay, but also improves 0.2%-7.2% of the video quality with comparable bandwidth consumption, under a variety of application scenarios.

cs.MM

Mamba: Bringing Multi-Dimensional ABR to WebRTC

Contemporary real-time video communication systems, such as WebRTC, use an adaptive bitrate (ABR) algorithm to assure high-quality and low-delay services, e.g., promptly adjusting video bitrate according to the instantaneous network bandwidth. However, target bitrate decisions in the network and bitrate control in the codec are typically incoordinated and simply ignoring the effect of inappropriate resolution and frame rate settings also leads to compromised results in bitrate control, thus devastatingly deteriorating the quality of experience (QoE). To tackle these challenges, Mamba, an end-to-end multi-dimensional ABR algorithm is proposed, which utilizes multi-agent reinforcement learning (MARL) to maximize the user's QoE by adaptively and collaboratively adjusting encoding factors including the quantization parameters (QP), resolution, and frame rate based on observed states such as network conditions and video complexity information in a video conferencing system. We also introduce curriculum learning to improve the training efficiency of MARL. Both the in-lab and real-world evaluation results demonstrate the remarkable efficacy of Mamba.

cs.MM

Enabling Joint Radar-Communication Operation in Shift Register-Based PMCW Radars

This article introduces adaptations to the conventional frame structure in binary phase-modulated continuous wave (PMCW) radars with sequence generation via linear-feedbck shift registers and additional processing steps to enable joint radar-communication (RadCom) operation. In this context, a preamble structure based on pseudorandom binary sequences (PRBSs) that is compatible with existing synchronization algorithms is outlined, and the allocation of pilot PRBS blocks is discussed. Finally, results from proof-of-concept measurements are presented to illustrate the effects of the choice of system and signal parameters and validate the investigated PMCW-based RadCom system and synchronization strategy.

eess.SP

Improving ABR Performance for Short Video Streaming Using Multi-Agent Reinforcement Learning with Expert Guidance

In the realm of short video streaming, popular adaptive bitrate (ABR) algorithms developed for classical long video applications suffer from catastrophic failures because they are tuned to solely adapt bitrates. Instead, short video adaptive bitrate (SABR) algorithms have to properly determine which video at which bitrate level together for content prefetching, without sacrificing the users' quality of experience (QoE) and yielding noticeable bandwidth wastage jointly. Unfortunately, existing SABR methods are inevitably entangled with slow convergence and poor generalization. Thus, in this paper, we propose Incendio, a novel SABR framework that applies Multi-Agent Reinforcement Learning (MARL) with Expert Guidance to separate the decision of video ID and video bitrate in respective buffer management and bitrate adaptation agents to maximize the system-level utilized score modeled as a compound function of QoE and bandwidth wastage metrics. To train Incendio, it is first initialized by imitating the hand-crafted expert rules and then fine-tuned through the use of MARL. Results from extensive experiments indicate that Incendio outperforms the current state-of-the-art SABR algorithm with a 53.2% improvement measured by the utility score while maintaining low training complexity and inference time.

cs.MM

Discrete-Fresnel Domain Channel Estimation in OCDM-based Radar Systems

In recent years, orthogonal chirp-division multiplexing (OCDM) has been increasingly considered as an alternative multicarrier scheme, e.g., to orthogonal frequency-division multiplexing, in digital communication applications. Among reasons for thar are its demonstrated superior performance resulting from its robustness to impairments such as frequency selectivity of channels and intersymbol interference. Furthermore, the so-called unbiased channel estimation in the discrete-Fresnel domain has also been investigated for both communication and sensing systems, however without considering the effects of frequency shifts. This article investigates the suitability of the aforementioned discrete-Fresnel domain channel estimation in OCDM-based radar systems as an alternative to the correlation-based processing previously adopted, e.g., in the radar-communication (RadCom) literature, which yields high sidelobe level depending on the symbols modulated onto the orthogonal subchirps. In this context, a mathematical formulation for the aforementioned channel estimation approach is introduced. Additionally, extensions to multi-user/multiple-input multiple-output and RadCom operations are proposed. Finally, the performance of the proposed schemes is analyzed, and the presented discussion is supported by simulation and measurement results. In summary, all proposed OCDM-based schemes yield comparable radar sensing performance to their orthogonal frequency-division multiplexing counterpart, while achieving improved peak-to-average power ratio and, in the RadCom case, communication performance.

eess.SP