SearcharxivSearch

arXiv subjects

Dapeng Li

Publications and source records attributed to Dapeng Li.

At least 19 recordsLinked to original sources

Knowledge Distillation Driven Semantic NOMA with GAN Refinement for 6G Robotic Vehicle Networks

To achieve sustainable intelligent mobility, 6G-empowered robotic vehicles (RVs) require high-fidelity visual perception under stringent bandwidth and energy constraints. Semantic communication offers a spectral-efficient solution but suffers from severe interference in uplink non-orthogonal multiple access (NOMA) RV networks. To address this, we propose a knowledge distillation-driven and generative models-enhanced NOMA framework for robust and green RV communications, named KDG-SemNOMA. First, we develop a ConvNeXt-based deep joint source-channel coding (DeepJSCC) architecture with an enhanced attention feature (AF) module for dynamic channel adaptation. Second, to mitigate interference without inference overhead, an orthogonal transmission teacher model guides the NOMA student model via a two-stage knowledge distillation strategy. Finally, to address the over-smoothing artifacts of pixel-wise optimization, we introduce a channel-conditional GAN (cGAN). By explicitly taking the Stage-I initial reconstruction and channel states as conditional inputs, this module refines coarse outputs into high-fidelity images with realistic textures. Experiments on FFHQ-256 demonstrate that KDG-SemNOMA significantly outperforms state-of-the-art methods in both pixel-level accuracy and perceptual fidelity.

cs.IT

Multi-Domain Iterative Detection for Massive Connectivity in LEO Satellite Networks

Grant-Free (GF) random access is promising for low Earth orbit satellite Internet due to its reduced access latency. However, existing schemes suffer from poor performance in massive connectivity scenarios. To address this challenge, we firstly propose an iterative residual feedback multi-measurement vector approximate message passing algorithm. This algorithm leverages multi-domain synergistic sparsity in the spatial-frequency and angular-delay domains to alternately perform active user terminal detection (AUD) and channel estimation (CE). Additionally, a residual feedback mechanism is incorporated to suppress error accumulation, thereby enhancing AUD performance. Furthermore, conventional data detection (DD) methods significantly degrade when active user terminals are spatially close or outnumber the satellite's receive antennas, making the demodulation problem rank-deficient or underdetermined. To mitigate this, we design a data modulation scheme via joint spatial-frequency multi-domain spreading, which utilizes observations from both spatial and frequency domains to facilitate multi-domain DD. Simulation results demonstrate that the proposed scheme significantly outperforms existing GF methods in terms of AUD accuracy, CE precision, and bit error rate, especially under conditions of low effective pilot length and practical signal-to-noise ratios.

eess.SP

DJSCC-Enabled Multi-User Semantic CSI Feedback for Hybrid Beamforming in Dual-Polarized cmWave Massive MIMO

Driven by the ultra-high throughput requirements of 6G, wireless communications are migrating to centimeter wave (cmWave) bands to overcome the limitations of current spectral resources. Massive multiple-input multiple-output (MIMO) and orthogonal frequency division multiplexing (OFDM) systems aim to achieve high spectral efficiency in cmWave regimes but are often constrained by the heavy overhead of downlink channel state information (CSI) feedback. This paper proposes a deep learning scheme based on the multi-axis multi-layer perceptron for image processing (MAXIM) architecture for joint semantic CSI feedback and hybrid beamforming in multi-user cmWave MIMO-OFDM systems, which maximizes the downlink sum rate by end-to-end optimization. Specifically, distributed encoders at multiple user equipments (UEs) perform limited CSI feedback, while the decoder at the base station (BS) jointly designs the hybrid beamforming matrices without explicit CSI reconstruction. The uplink transmission is implemented via deep joint source-channel coding (DJSCC) to enhance CSI compression efficiency and noise robustness. Furthermore, considering the high correlation between vertical and horizontal polarization channels in dual-polarized massive MIMO systems, a cross-polarization interaction module is introduced at the UEs to exploit polarization correlations for joint CSI compression. Simulation results demonstrate that the proposed method improves the downlink sum rate under various signal-to-noise ratio (SNR) conditions with a limited number of feedback symbols, validating its robustness and superiority in multi-user dual-polarized cmWave MIMO-OFDM systems.

eess.SP

Adaptive Channel Estimation and Hybrid Beamforming for RIS aided Vehicular Communication

Reconfigurable intelligent surface (RIS) constitutes a disruptive technology for enhancing vehicular communication performance through reconfigurable propagation environments. In this paper, we propose an adaptive channel estimation framework and hybrid beamforming optimization strategy for RIS-aided vehicular multiple-input multiple-output (MIMO) systems operating in high-mobility scenarios. To address severe Doppler effects and rapid channel variations, we design a velocity-aware pilot scheme that progressively estimates cascaded channels across two timescales, leveraging tensor decomposition and adaptive grouping of passive elements. This framework dynamically balances channel estimation accuracy and spectral efficiency, significantly reducing training overhead. Furthermore, we develop a low-complexity hybrid beamforming algorithm for both narrowband single vehicle user equipment (VUE) and broadband multi-VUE systems. For single-VUE scenarios, we derive closed-form active beamforming solutions and optimize passive beamforming via alternating optimization. For multi-VUE broadband systems, we jointly optimize subcarrier allocation, power distribution, and beamforming to maximize system throughput while mitigating inter-carrier interference (ICI) caused by Doppler spread, subject to quality-of-service (QoS) constraints and RIS hardware limitations. Our simulation results demonstrate that the proposed methods achieve substantial performance gains in channel estimation efficiency, beamforming robustness, and system throughput compared to conventional schemes, particularly under high mobility conditions.

eess.SY

QSIM: Mitigating Overestimation in Multi-Agent Reinforcement Learning via Action Similarity Weighted Q-Learning

Value decomposition (VD) methods have achieved remarkable success in cooperative multi-agent reinforcement learning (MARL). However, their reliance on the max operator for temporal-difference (TD) target calculation leads to systematic Q-value overestimation. This issue is particularly severe in MARL due to the combinatorial explosion of the joint action space, which often results in unstable learning and suboptimal policies. To address this problem, we propose QSIM, a similarity weighted Q-learning framework that reconstructs the TD target using action similarity. Instead of using the greedy joint action directly, QSIM forms a similarity weighted expectation over a structured near-greedy joint action space. This formulation allows the target to integrate Q-values from diverse yet behaviorally related actions while assigning greater influence to those that are more similar to the greedy choice. By smoothing the target with structurally relevant alternatives, QSIM effectively mitigates overestimation and improves learning stability. Extensive experiments demonstrate that QSIM can be seamlessly integrated with various VD methods, consistently yielding superior performance and stability compared to the original algorithms. Furthermore, empirical analysis confirms that QSIM significantly mitigates the systematic value overestimation in MARL. Code is available at https://github.com/MaoMaoLYJ/pymarl-qsim.

cs.MA

How to Build Robust, Scalable Models for GSV-Based Indicators in Neighborhood Research

A substantial body of health research demonstrates a strong link between neighborhood environments and health outcomes. Recently, there has been increasing interest in leveraging advances in computer vision to enable large-scale, systematic characterization of neighborhood built environments. However, the generalizability of vision models across fundamentally different domains remains uncertain, for example, transferring knowledge from ImageNet to the distinct visual characteristics of Google Street View (GSV) imagery. In applied fields such as social health research, several critical questions arise: which models are most appropriate, whether to adopt unsupervised training strategies, what training scale is feasible under computational constraints, and how much such strategies benefit downstream performance. These decisions are often costly and require specialized expertise. In this paper, we answer these questions through empirical analysis and provide practical insights into how to select and adapt foundation models for datasets with limited size and labels, while leveraging larger, unlabeled datasets through unsupervised training. Our study includes comprehensive quantitative and visual analyses comparing model performance before and after unsupervised adaptation.

cs.CV

Balancing Rewards in Text Summarization: Multi-Objective Reinforcement Learning via HyperVolume Optimization

Text summarization is a crucial task that requires the simultaneous optimization of multiple objectives, including consistency, coherence, relevance, and fluency, which presents considerable challenges. Although large language models (LLMs) have demonstrated remarkable performance, enhanced by reinforcement learning (RL), few studies have focused on optimizing the multi-objective problem of summarization through RL based on LLMs. In this paper, we introduce hypervolume optimization (HVO), a novel optimization strategy that dynamically adjusts the scores between groups during the reward process in RL by using the hypervolume method. This method guides the model's optimization to progressively approximate the pareto front, thereby generating balanced summaries across multiple objectives. Experimental results on several representative summarization datasets demonstrate that our method outperforms group relative policy optimization (GRPO) in overall scores and shows more balanced performance across different dimensions. Moreover, a 7B foundation model enhanced by HVO performs comparably to GPT-4 in the summarization task, while maintaining a shorter generation length. Our code is publicly available at https://github.com/ai4business-LiAuto/HVO.git

cs.CL

Latency Minimization for Hybrid-Frequency UHD Upload in Double-IRS-Aided HSR Networks

Real-time mechanical fault diagnosis in high-speed railway (HSR) networks requires ultra-reliable and low-latency upload of ultra-high-definition (UHD) video streams. However, energy constraints of trackside cameras and severe transmission latency pose critical challenges. This paper proposes a novel 6G infrastructure-to-vehicle (I2V) architecture employing double intelligent reflecting surfaces (IRSs) to enhance wireless powered communication network (WPCN) and hybrid-frequency data transmission. Crucially, to guarantee the quality of experience (QoE) for in-cabin passengers using Mobile Multimedia Broadcasting Services (MBMS), a strict zero-forcing spatial interference isolation constraint is imposed via the window-mounted IRS. We formulate a weighted latency minimization problem and develop a block coordinate descent (BCD) algorithm. Downlink energy beamforming and uplink information transmission are alternately optimized utilizing difference of convex (DCA) and semi-definite relaxation (SDR) techniques. Additionally, a low-complexity heuristic algorithm is proposed to mitigate the severe Doppler spread induced by train mobility. Simulation results demonstrate that the proposed scheme significantly reduces upload latency to meet stringent URLLC thresholds while ensuring interference isolation within the carriage.

eess.SP

Calibration-free Rydberg Atomic Receiver for Sub-MHz Wireless Communications and Sensing

The exploitation of sub-MHz (\textless 1 MHz) can be beneficial for a plethora of applications like underwater vehicular communication, subsurface exploration, low-frequency navigation etc. The traditional electrical receivers in this band are either hundreds of meters long or, when miniaturized, inefficient and bandwidth-limited, making them inapplicable for practical underwater implementations. Such obstacles can be circumvented by the emerging Rydberg atomic receiving technology, which is capable of detecting fields from DC up to the terahertz regime with compact structure. Against this background, we propose a method to detect sub-MHz electric fields without further calibration. Specifically, a physics-based model of the combined DC and AC-Stark response is established. Based on the model, we modulate the DC-Stark spectrum with the received signal and extract its amplitude by fitting the cycle-averaged, symmetric Stark-split peaks. Then we map this swing directly to the intrinsic atomic polarizability. By such operations, the proposed method can remove the dependence on electrode spacing or field-amplitude references. For performance evaluation, six-level Lindblad simulations and experiments are conducted at a low-frequency field of 30 kHz demonstrate a minimum detectable field of 5.3 \text{mV}/\text{cm}, with stable readout across practical optical-power variations. The approach manages to expand operating range of Rydberg atomic receivers below 1 MHz, and enables compact, calibration-free quantum front ends for underwater and subsurface receivers.

physics.ins-det

Beyond Local Views: Global State Inference with Diffusion Models for Cooperative Multi-Agent Reinforcement Learning

In partially observable multi-agent systems, agents typically only have access to local observations. This severely hinders their ability to make precise decisions, particularly during decentralized execution. To alleviate this problem and inspired by image outpainting, we propose State Inference with Diffusion Models (SIDIFF), which uses diffusion models to reconstruct the original global state based solely on local observations. SIDIFF consists of a state generator and a state extractor, which allow agents to choose suitable actions by considering both the reconstructed global state and local observations. In addition, SIDIFF can be effortlessly incorporated into current multi-agent reinforcement learning algorithms to improve their performance. Finally, we evaluated SIDIFF on different experimental platforms, including Multi-Agent Battle City (MABC), a novel and flexible multi-agent reinforcement learning environment we developed. SIDIFF achieved desirable results and outperformed other popular algorithms.

cs.MA

Verco: Learning Coordinated Verbal Communication for Multi-agent Reinforcement Learning

In recent years, multi-agent reinforcement learning algorithms have made significant advancements in diverse gaming environments, leading to increased interest in the broader application of such techniques. To address the prevalent challenge of partial observability, communication-based algorithms have improved cooperative performance through the sharing of numerical embedding between agents. However, the understanding of the formation of collaborative mechanisms is still very limited, making designing a human-understandable communication mechanism a valuable problem to address. In this paper, we propose a novel multi-agent reinforcement learning algorithm that embeds large language models into agents, endowing them with the ability to generate human-understandable verbal communication. The entire framework has a message module and an action module. The message module is responsible for generating and sending verbal messages to other agents, effectively enhancing information sharing among agents. To further enhance the message module, we employ a teacher model to generate message labels from the global view and update the student model through Supervised Fine-Tuning (SFT). The action module receives messages from other agents and selects actions based on current local observations and received messages. Experiments conducted on the Overcooked game demonstrate our method significantly enhances the learning efficiency and performance of existing methods, while also providing an interpretable tool for humans to understand the process of multi-agent cooperation.

cs.MA

KnowledgeNavigator: Leveraging Large Language Models for Enhanced Reasoning over Knowledge Graph

Large language model (LLM) has achieved outstanding performance on various downstream tasks with its powerful natural language understanding and zero-shot capability, but LLM still suffers from knowledge limitation. Especially in scenarios that require long logical chains or complex reasoning, the hallucination and knowledge limitation of LLM limit its performance in question answering (QA). In this paper, we propose a novel framework KnowledgeNavigator to address these challenges by efficiently and accurately retrieving external knowledge from knowledge graph and using it as a key factor to enhance LLM reasoning. Specifically, KnowledgeNavigator first mines and enhances the potential constraints of the given question to guide the reasoning. Then it retrieves and filters external knowledge that supports answering through iterative reasoning on knowledge graph with the guidance of LLM and the question. Finally, KnowledgeNavigator constructs the structured knowledge into effective prompts that are friendly to LLM to help its reasoning. We evaluate KnowledgeNavigator on multiple public KGQA benchmarks, the experiments show the framework has great effectiveness and generalization, outperforming previous knowledge graph enhanced LLM methods and is comparable to the fully supervised models.

cs.CL

Adaptive parameter sharing for multi-agent reinforcement learning

Parameter sharing, as an important technique in multi-agent systems, can effectively solve the scalability issue in large-scale agent problems. However, the effectiveness of parameter sharing largely depends on the environment setting. When agents have different identities or tasks, naive parameter sharing makes it difficult to generate sufficiently differentiated strategies for agents. Inspired by research pertaining to the brain in biology, we propose a novel parameter sharing method. It maps each type of agent to different regions within a shared network based on their identity, resulting in distinct subnetworks. Therefore, our method can increase the diversity of strategies among different agents without introducing additional training parameters. Through experiments conducted in multiple environments, our method has shown better performance than other parameter sharing methods.

cs.AI

Controlling Large Language Model-based Agents for Large-Scale Decision-Making: An Actor-Critic Approach

The remarkable progress in Large Language Models (LLMs) opens up new avenues for addressing planning and decision-making problems in Multi-Agent Systems (MAS). However, as the number of agents increases, the issues of hallucination in LLMs and coordination in MAS have become increasingly prominent. Additionally, the efficient utilization of tokens emerges as a critical consideration when employing LLMs to facilitate the interactions among a substantial number of agents. In this paper, we develop a modular framework called LLaMAC to mitigate these challenges. LLaMAC implements a value distribution encoding similar to that found in the human brain, utilizing internal and external feedback mechanisms to facilitate collaboration and iterative reasoning among its modules. Through evaluations involving system resource allocation and robot grid transportation, we demonstrate the considerable advantages afforded by our proposed approach.

cs.AI

Stackelberg Decision Transformer for Asynchronous Action Coordination in Multi-Agent Systems

Asynchronous action coordination presents a pervasive challenge in Multi-Agent Systems (MAS), which can be represented as a Stackelberg game (SG). However, the scalability of existing Multi-Agent Reinforcement Learning (MARL) methods based on SG is severely constrained by network structures or environmental limitations. To address this issue, we propose the Stackelberg Decision Transformer (STEER), a heuristic approach that resolves the difficulties of hierarchical coordination among agents. STEER efficiently manages decision-making processes in both spatial and temporal contexts by incorporating the hierarchical decision structure of SG, the modeling capability of autoregressive sequence models, and the exploratory learning methodology of MARL. Our research contributes to the development of an effective and adaptable asynchronous action coordination method that can be widely applied to various task types and environmental configurations in MAS. Experimental results demonstrate that our method can converge to Stackelberg equilibrium solutions and outperforms other existing methods in complex scenarios.

cs.MA

From Explicit Communication to Tacit Cooperation:A Novel Paradigm for Cooperative MARL

Centralized training with decentralized execution (CTDE) is a widely-used learning paradigm that has achieved significant success in complex tasks. However, partial observability issues and the absence of effectively shared signals between agents often limit its effectiveness in fostering cooperation. While communication can address this challenge, it simultaneously reduces the algorithm's practicality. Drawing inspiration from human team cooperative learning, we propose a novel paradigm that facilitates a gradual shift from explicit communication to tacit cooperation. In the initial training stage, we promote cooperation by sharing relevant information among agents and concurrently reconstructing this information using each agent's local trajectory. We then combine the explicitly communicated information with the reconstructed information to obtain mixed information. Throughout the training process, we progressively reduce the proportion of explicitly communicated information, facilitating a seamless transition to fully decentralized execution without communication. Experimental results in various scenarios demonstrate that the performance of our method without communication can approaches or even surpasses that of QMIX and communication-based methods.

cs.MA

SEA: A Spatially Explicit Architecture for Multi-Agent Reinforcement Learning

Spatial information is essential in various fields. How to explicitly model according to the spatial location of agents is also very important for the multi-agent problem, especially when the number of agents is changing and the scale is enormous. Inspired by the point cloud task in computer vision, we propose a spatial information extraction structure for multi-agent reinforcement learning in this paper. Agents can effectively share the neighborhood and global information through a spatially encoder-decoder structure. Our method follows the centralized training with decentralized execution (CTDE) paradigm. In addition, our structure can be applied to various existing mainstream reinforcement learning algorithms with minor modifications and can deal with the problem with a variable number of agents. The experiments in several multi-agent scenarios show that the existing methods can get convincing results by adding our spatially explicit architecture.

cs.MA

Inducing Stackelberg Equilibrium through Spatio-Temporal Sequential Decision-Making in Multi-Agent Reinforcement Learning

In multi-agent reinforcement learning (MARL), self-interested agents attempt to establish equilibrium and achieve coordination depending on game structure. However, existing MARL approaches are mostly bound by the simultaneous actions of all agents in the Markov game (MG) framework, and few works consider the formation of equilibrium strategies via asynchronous action coordination. In view of the advantages of Stackelberg equilibrium (SE) over Nash equilibrium, we construct a spatio-temporal sequential decision-making structure derived from the MG and propose an N-level policy model based on a conditional hypernetwork shared by all agents. This approach allows for asymmetric training with symmetric execution, with each agent responding optimally conditioned on the decisions made by superior agents. Agents can learn heterogeneous SE policies while still maintaining parameter sharing, which leads to reduced cost for learning and storage and enhanced scalability as the number of agents increases. Experiments demonstrate that our method effectively converges to the SE policies in repeated matrix game scenarios, and performs admirably in immensely complex settings including cooperative tasks and mixed tasks.

cs.MA