SearcharxivSearch

arXiv subjects

Dongming Wang

Publications and source records attributed to Dongming Wang.

At least 19 recordsLinked to original sources

Geodesic strong convexity does not imply forward invariance under gradient flow on SO(3): a certified counterexample

Let $mathcal{C}=\overline{\mathcal{B}}_{\rho}(R_c)$ be a geodesic ball of radius $\rho<\pi/2$ in SO(3) with the bi-invariant metric, and let $f$ be geodesically strongly convex on $\mathcal{C}$ with an interior minimizer. It is tempting to expect the gradient flow $\dot R=R(-\nabla f)^\wedge$ to keep $\mathcal{C}$ forward invariant: the flow is attracted to an interior point, and strong convexity appears to leave no room for outward motion. We show this expectation is false by an explicit, fully certified construction with $\rho=0.3$: a cost, quadratic in the principal logarithmic chart with off-diagonal coupling $0.7$, whose geodesic Hessian satisfies $\Hess f\succeq\mu I_3$ on all of $\mathcal{C}$ with a machine-certified modulus $\mu\geq0.172$, rigorous ball arithmetic over exact rational inputs, yet whose descent velocity at a boundary point has the exact rational outward radial component $21/500$. A continuity corollary of the exact rate certifies that the flow exits the ball; numerical integration puts the peak excursion near $0.3143$ before convergence to the minimizer. The mechanism is elementary: strong convexity constrains the projection of the gradient onto the minimizer direction, not onto the inward radial direction. Code reproducing every certified constant and figure accompanies the note.

eess.SY

Securing Cooperative Sensing in UAV Swarms Against Conformity-Driven Byzantine Attacks

In integrated sensing and communication (ISAC) enabled 6G unmanned aerial vehicle (UAV) swarm networks, the widely adopted imitation-based conformity cooperation mechanism can be exploited by Byzantine attackers to fabricate false consensus, causing the effective error probability of normal UAVs to evolve dynamically and far exceed their inherent sensing errors, which invalidates conventional fusion methods built on the independence assumption. This paper proposes a conformity-aware Byzantine-resilient fusion framework that couples evolutionary game theory with maximum a posteriori (MAP) estimation. First, the strategy updates of normal UAVs are characterized by bounded-rational opinion dynamics, and the evolution dynamics of the misinformation ratio together with its evolutionarily stable state (ESS) are derived under death birth updating. Three theoretical results are then established: under heterogeneous per-node sensing errors, the zeroth-order ESS depends on the error distribution only through its mean; a closed-form first-order weak-selection correction to the ESS is obtained, together with an exact mean-field fixed point valid for arbitrary selection intensity; and it is revealed that swarm level misinformation can overwhelm the majority if and only if the attack probability exceeds one half, with this threshold independent of both the sensing error and the malicious ratio. Embedding the predicted error dynamics into a per-node MAP rule, the resulting fusion mechanism achieves nearly 100% situation-inference accuracy under different network topologies, attack intensities, network scales, and sensing-error distributions, and maintains accuracy above 99% under +-20% parameter mismatch. In contrast, majority voting, reputation weighting, and independent fusion collapse completely once the majority-flip threshold is crossed.

eess.SY

Reliability-Constrained Hybrid Beamforming for Multistatic ISAC in Vehicular Networks

This letter investigates reliability constrained hybrid beamforming for transceiver separated multistatic integrated sensing and communication in vehicular networks. A target position Cramer Rao bound minimization problem is formulated under outage probability, transmit-power, and analog constant modulus constraints. To handle the constrained non convex problem, we develop a proportional-integral Lagrangian proximal policy optimization algorithm. Simulation results show that the proposed algorithm keeps the average outage probability at or below the reliability threshold, around 8%-10%, improves constraint satisfaction, and achieves stable sensing performance.

eess.SY

Human-in-the-Loop Distributed Control of Grid-Interactive Buildings for Demand Response Participation

This paper proposes a human-in-the-loop distributed consensus control approach for demand-side management across multiple buildings. Specifically, a novel framework is introduced in which a human acts as the non-autonomous leader in consensus control of cooperative buildings participating in demand response programs. In this system, the facility manager in the facility building serves as the leader, determining the participation level of cooperative buildings in demand response events while simultaneously considering occupants' comfort. Cooperative buildings align their responses with the facility management building, despite lacking direct access to the facility manager's decisions, which presents a challenge for observer design. To address this, a nonlinear unknown input sliding-mode observer is proposed, tailored for leader-follower multi-agent systems (MASs). Furthermore, a human-in-the-loop leader-follower consensus protocol is introduced, enabling a framework to flexibly manage energy use and balance thermal comfort during demand response events. Simulation results validate the effectiveness of the proposed approach, demonstrating its ability to achieve consensus, maintain system performance, and enhance the adaptability of power grid operations under various demand response scenarios.

eess.SY

CHMAS: A Coupled Hierarchical Framework for Multi-Agent Reinforcement Learning

Multi-agent reinforcement learning (MARL) systems face fundamental challenges in balancing global coordination with local execution across different temporal scales. This paper introduces the Coupled Hierarchical Multi-Agent System (CHMAS), a novel framework that decomposes multi-agent decision-making into centralized strategic planning and distributed tactical execution with bidirectional information flow. The strategic layer integrates all agents' states with an exclusive global environmental state to generate guidance actions every $T$ timesteps, while tactical agents execute distributed policies augmented by strategic guidance and local neighborhood observations. Unlike existing hierarchical approaches with unidirectional control, CHMAS establishes a feedback mechanism where accumulated tactical rewards influence strategic objectives through a coupling coefficient $\lambda$, ensuring strategic plans remain grounded in tactical feasibility. To address the non-stationarity inherent in hierarchical learning, we propose an asynchronous update protocol where strategic parameters update every $N_f$ tactical episodes, allowing tactical policies to converge to quasi-stationary points between strategic changes. We present both a general bi-level formulation capturing full system dynamics and a tractable additive approximation enabling rigorous analysis. Theoretical analysis proves that this asynchronous scheme achieves $\mathcal{O}(\log K/\sqrt{K})$ convergence for the strategic layer after $K$ strategic updates under standard assumptions. Experimental validation in a multi-agent foraging domain demonstrates successful learning of spatially partitioned exploration strategies, with both layers converging stably despite hierarchical coupling.

cs.MA

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces

We develop the Continuous Distributed Coupled Policy Gradient (CDCPG) algorithm for cooperative reinforcement learning in networked Markov decision processes with continuous state and action spaces. Each agent maintains a local actor over a bounded graph neighborhood, and a localized least-squares temporal-difference critic evaluates a truncated action-value function through a spectral random-feature representation of the local transition kernel. The analysis makes four contributions. First, the truncated action-value function is constructed as a conditional expectation over the neighborhood, yielding a well-posed localized Bellman theory that removes the continuation-kernel mismatch of naive truncation arguments. Second, we expose a dimensional obstruction to temporal-difference stability for normalized random features and prove an unconditional excitation bound that reduces stability to a symmetric persistence-of-excitation condition, monitorable through an online matrix-concentration certificate. Third, under exponential spatial decay of agent interactions, the excitation condition, and smoothness of the objective, CDCPG drives an averaged per-agent stationarity measure to within any excess $\epsilon$ of an explicitly characterized approximation floor using $\widetilde{\mathcal{O}}(\epsilon^{-2})$ shared-oracle samples, and the excess dependence matches the smooth nonconvex first-order rate; per-agent computation and communication are governed by the neighborhood size rather than the network size. Fourth, an adaptive-locality rule selects the radius that balances truncation and graph-decay residuals against the target accuracy. Experiments on a networked linear-quadratic benchmark corroborate the locality and feature-dimension predictions.

cs.MA

Asymptotically Optimal Local Receiver in Uplink CF-mMIMO: A Functional-Variational Analysis

In cell-free massive multiple-input multiple-output (CF-mMIMO) systems, the canonical uplink local receiver is the local minimum mean square error (LMMSE) receiver with large-scale fading decoding (LSFD) at the central processing unit (CPU). The LSFD coefficients are derived under the use-and-then-forget (UatF) lower bound of the ergodic rate, and computing these coefficients introduces additional fronthaul overhead and computational complexity at the CPU. This paper investigates local receiver design directly from the true ergodic-rate objective under perfect local channel state information (CSI). By introducing an expectation-based constraint and leveraging large-system random matrix theory, we develop a functional-variational approach that yields the asymptotically optimal quasi-LMMSE (Q-LMMSE) receiver in closed form. A key insight is that the Q-LMMSE receiver shares the same direction as the conventional LMMSE receiver, differing only by an instantaneous CSI-dependent scalar, and thus incurs the same per-access point (AP) complexity. More importantly, this scalar varies across APs and implicitly provides adaptive weighting for the direct summation at the CPU, thereby completely eliminating the need for statistical LSFD coefficients and the associated CPU-side computational overhead. Numerical results demonstrate that the proposed Q-LMMSE receiver consistently outperforms the LMMSE-LSFD benchmark in terms of the ergodic rate, achieving approximately a {5\%} gain when the number of antennas per AP is low, while operating with strictly lower system-level complexity.

eess.SP

Adaptive Joint Beamforming and Fluid Antenna System Design for 6G ISAC

Fixed-Position Antennas (FPAs) are constrained by static physical topologies and struggle to adapt to rapidly varying wireless environments. By dynamically reconfiguring the antenna positions, Fluid Antenna Systems (FASs) introduce additional spatial Degrees of Freedom (DoF) for wireless optimization. This paper investigates the joint optimization of Fluid Antenna System (FAS) topology reconfiguration and active beamforming for mobile Integrated Sensing and Communication (ISAC) systems. To enable real-time decision making, an end-to-end optimization framework based on the Soft Actor-Critic (SAC) algorithm is proposed. Simulation results show that the proposed scheme achieves an online inference latency of only 4 ms. Compared to the widely used alternating optimization, it improves communication performance by 42%. Moreover, it achieves performance comparable to the SCA-SDR benchmark while requiring 57% fewer antennas, demonstrating superior hardware efficiency.

eess.SY

User-Centric Clustering for uRLLC in Cell-Free RAN via Extreme Value Theory

Ultra-reliable low-latency communication (uRLLC) is a pivotal enabler for B5G/6G networks, yet it faces severe challenges from rare but critical extreme events, which are characterized by heavy tails in the delay distribution. While the cell-free radio access network (CF-RAN) architecture offers essential spatial diversity to combat these uncertainties, conventional user-centric clustering designs typically focus on average metrics, thereby inadequately addressing such tail behaviors. We propose a novel, tail-risk-aware, user-centric clustering framework operating within the finite blocklength (FBL) regime. Our approach employs extreme value theory (EVT), specifically the peaks-over-threshold (POT) model, to accurately quantify the probability of queue latency violations. This framework is applied to formulate an energy efficiency (EE) maximization problem under strict tail latency constraints. The problem is solved via an efficient online algorithm that integrates Lyapunov optimization with successive convex approximation (SCA). Simulation results demonstrate that the proposed scheme, through its dynamic adaptation of cluster formation to mitigate tail risks, achieves a superior reliability-efficiency trade-off and leads to a significant suppression of extreme latency events.

cs.IT

Distributed Zeroth-Order Policy Gradient for Networked Multi-agent Reinforcement Learning from Human Feedback

We study a networked multi-agent reinforcement learning (NMARL) problem with human feedback in an infinite-horizon setting, where agents interact over an underlying network with localized state dependencies and aim to collaboratively maximize the average discounted return. Existing approaches with preference feedback are primarily developed for single-agent settings and rely on centralized training, which limits their scalability and applicability to large-scale networked multi-agent systems. To address this, we introduce a novel human feedback mechanism based on spatiotemporally truncated trajectories, defined as $H$-horizon trajectory pairs aggregated over each agent's $κ$-hop neighborhood. Building on this, we develop a distributed zeroth-order policy gradient algorithm, where each agent estimates its local policy gradient using human preference feedback generated from both the current joint policy and a perturbed joint policy drawn from zero-mean Gaussian distribution. Specifically, the algorithm is fully distributed, as the feedback received by each agent depends solely on the state-action information within its $κ$-hop neighborhood and does not require explicit reward signals or centralized control. We further rigorously establish that the proposed algorithm converges to an $ε$-stationary point with polynomial sample complexity. Finally, simulation results in a stochastic GridWorld environment and a predator-prey environment further demonstrate that the effectiveness and scalability of the proposed algorithm in achieving collaborative optimization based solely on human preference feedback.

cs.MA

CRS-LLM: Cooperative Beam Prediction with a GPT-Style Backbone and Switch-Gated Fusion

Millimeter-wave (mmWave) communication depends on highly directional beamforming, while fast mobility, blockage, and rapid geometry changes in vehicle-to-everything (V2X) scenarios make beam tracking challenging. In cooperative multi-base-station (BS) systems, conventional hierarchical methods usually separate BS selection and beam selection, which may cause error propagation when beam states change abruptly. To address this issue, this paper proposes Cooperative Radio Sensing with Large Language Models (CRS-LLM), a cooperative beam prediction framework for next-step joint BS-beam prediction. CRS-LLM formulates beam tracking as a single classification problem over the joint BS-beam space, avoiding cascaded decision errors. To adapt channel state information (CSI) to large language models, a dual-view CSI tokenizer extracts frequency-domain and delay-domain channel features through a lightweight CNN front-end and temporal tokenization module. A truncated GPT-style backbone is then used for temporal modeling with parameter-efficient adaptation. In addition, a transition-aware switch-gated predictor combines a stable branch, a residual flip branch, and a low-rank transition prior to capture both smooth evolution and abrupt changes. Simulation results show that CRS-LLM outperforms CSI-Transformer, Hierarchical BS-Beam, and representative CNN- and recurrent-neural-network baselines in Top-1 accuracy and normalized beam gain under different SNR conditions, while also showing strong few-shot performance and promising zero-shot transferability.

eess.SP

A Framework for Uplink ISAC Receiver Designs: Performance Analysis and Algorithm Development

Uplink integrated sensing and communication (ISAC) systems have recently emerged as a promising research direction, enabling simultaneous uplink signal detection and target sensing. {In this paper, we propose the flexible projection (FP)-type receiver that unifies the projection-type receiver and the successive interference cancellation (SIC)-type receiver by using a flexible tradeoff factor to adapt to dynamically changing uplink ISAC scenarios.} The FP-type receiver addresses the joint signal detection and target response estimation problem through two coordinated phases: 1) Communication signal detection using a reconstructed signal whose composition is controlled by the tradeoff factor, followed by 2) Target response estimation performed through subtraction of the detected communication signal from the received signal. With adjustable tradeoff factors, the FP-type receiver can balance the enhancement of the signal-to-interference-plus-noise ratio (SINR) with the reduction of correlation in the reconstructed signal for communication signal detection. The pairwise error probability (PEP) expressions are analyzed for both the maximum likelihood (ML) and the zero-forcing (ZF) detectors, revealing that the optimal tradeoff factor should be determined based on the adopted detection algorithm and the relative power of the sensing and communication (S\&C) signals. A homotopy optimization framework is first applied for the FP-type receiver with a fixed tradeoff factor. This framework is then extended to develop the dynamic flexible projection (DFP)-type receiver, which iteratively adjusts the tradeoff factor for improved algorithm performance and environmental adaptability. Finally, we show that the length of the jointly processed signal should scale with the antenna size to fully unleash the potential of the uplink ISAC receiver.

cs.IT

Experimental Performance of Bidirectional Phase Coherent Transmission and Sensing for mmWave Cell-free Massive MIMO Systems with Reciprocity Calibration

Phase synchronization among distributed transmission reception points (TRPs) is a prerequisite for enabling coherent joint transmission and high-precision sensing in millimeter wave (mmWave) cell-free massive multiple-input and multiple-output (MIMO) systems. This paper proposes a bidirectional calibration scheme and a calibration coefficient estimation method for phase synchronization, and presents a calibration coefficient phase tracking method using unilateral uplink/downlink channel state information (CSI). Furthermore, this paper introduces the use of reciprocity calibration to eliminate non-ideal factors in sensing and leverages sensing results to achieve calibration coefficient phase tracking in dynamic scenarios, thus enabling bidirectional empowerment of both communication and sensing. Simulation results demonstrate that the proposed method can effectively implement reciprocal calibration with lower overhead, enabling coherent collaborative transmission, and resolving non-ideal factors to acquire lower sensing error in sensing applications. Experimental results show that, in the mmWave band, over-the-air (OTA) bidirectional calibration enables coherent collaborative transmission for both collaborative TRPs and collaborative user equipments (UEs), achieving beamforming gain and long-time coherent sensing capabilities.

eess.SP

Base Station Sleeping Strategy Based on Load Sharing in Ultra-Dense Networks

To address the issues of high operational costs and low energy efficiency (EE) caused by the dense deployment of small base stations (s-BSs) in 5G ultra-dense networks (UDNs), this paper first constructs a multi-objective mathematical optimization model targeting maximizing EE and minimizing the number of active BSs. The model incorporates key constraints including BS operational state, user equipment (UE)-BS connection relationship, and load threshold, laying a theoretical foundation for the coordinated optimization of energy conservation and quality of service. Based on this model, an integrated solution combining UE-BS initial connection optimization and load-sharing based BS sleeping is proposed. In the initial connection phase, with communication quality and BS load as dual constraints, efficient matching between UEs and optimal BSs is achieved through three sequential steps: communication feasibility screening, redundant connection removal, and overload load redistribution. This resolves the problems of load imbalance and difficult identification of redundant BSs in UDNs arising from unordered initial connections. In the BS sleeping phase, a BS sleeping index, comprehensively considering UE transferability and backup BS resources, is innovatively introduced to quantify BS dormancy priority. Through a closed-loop process involving low-load BS screening, adjacent BS load evaluation, and load sharing by two takeover BSs based on their capacity, accurate dormancy of redundant BSs and collaborative load migration are realized. Simulation results in a typical UDNs scenario demonstrate that, compared with the traditional baseline scheme, the proposed solution exhibits significant advantages in convergence speed, optimization of the number of active BSs, and EE improvement.

eess.SY

Flexible-Duplex Cell-Free Architecture for Secure Uplink Communications in Low-Altitude Wireless Networks

Low-altitude wireless networks (LAWNs) are expected to play a central role in future 6G infrastructures, yet uplink transmissions of uncrewed aerial vehicles (UAVs) remain vulnerable to eavesdropping due to their limited transmit power, constrained antenna resources, and highly exposed air-ground propagation conditions. To address this fundamental bottleneck, we propose a flexible-duplex cell-free (CF) architecture in which each distributed access point (AP) can dynamically operate either as a receive AP for UAV uplink collection or as a transmit AP that generates cooperative artificial noise (AN) for secrecy enhancement. Such AP-level duplex flexibility introduces an additional spatial degree of freedom that enables distributed and adaptive protection against wiretapping in LAWNs. Building upon this architecture, we formulate a max-min secrecy-rate problem that jointly optimizes AP mode selection, receive combining, and AN covariance design. This tightly coupled and nonconvex optimization is tackled by first deriving the optimal receive combiners in closed form, followed by developing a penalty dual decomposition (PDD) algorithm with guaranteed convergence to a stationary solution. To further reduce computational burden, we propose a low-complexity sequential scheme that determines AP modes via a heuristic metric and then updates the AN covariance matrices through closed-form iterations embedded in the PDD framework. Simulation results show that the proposed flexible-duplex architecture yields substantial secrecy-rate gains over CF systems with fixed AP roles. The joint optimization method attains the highest secrecy performance, while the low-complexity approach achieves over 90% of the optimal performance with an order-of-magnitude lower computational complexity, offering a practical solution for secure uplink communications in LAWNs.

cs.IT

NeRF-VIO: Map-Based Visual-Inertial Odometry with Initialization Leveraging Neural Radiance Fields

A prior map serves as a foundational reference for localization in context-aware applications such as augmented reality (AR). Providing valuable contextual information about the environment, the prior map is a vital tool for mitigating drift. In this paper, we propose a map-based visual-inertial localization algorithm (NeRF-VIO) with initialization using neural radiance fields (NeRF). Our algorithm utilizes a multilayer perceptron model and redefines the loss function as the geodesic distance on \(SE(3)\), ensuring the invariance of the initialization model under a frame change within \(\mathfrak{se}(3)\). The evaluation demonstrates that our model outperforms existing NeRF-based initialization solution in both accuracy and efficiency. By integrating a two-stage update mechanism within a multi-state constraint Kalman filter (MSCKF) framework, the state of NeRF-VIO is constrained by both captured images from an onboard camera and rendered images from a pre-trained NeRF model. The proposed algorithm is validated using a real-world AR dataset, the results indicate that our two-stage update pipeline outperforms MSCKF across all data sequences.

cs.CV

Distributed scalable coupled policy algorithm for networked multi-agent reinforcement learning

This paper studies networked multi-agent reinforcement learning (NMARL) with interdependent rewards and coupled policies. In this setting, each agent's reward depends on its own state-action pair as well as those of its direct neighbors, and each agent's policy is parameterized by its local parameters together with those of its $κ_{p}$-hop neighbors, with $κ_{p}\geq 1$ denoting the coupled radius. The objective of the agents is to collaboratively optimize their policies to maximize the discounted average cumulative reward. To address the challenge of interdependent policies in collaborative optimization, we introduce a novel concept termed the neighbors' averaged $Q$-function and derive a new expression for the coupled policy gradient. Based on these theoretical foundations, we develop a distributed scalable coupled policy (DSCP) algorithm, where each agent relies only on the state-action pairs of its $κ_{p}$-hop neighbors and the rewards of its $(κ_{p}+1)$-hop neighbors. Specially, in the DSCP algorithm, we employ a geometric 2-horizon sampling method that does not require storing a full $Q$-table to obtain an unbiased estimate of the coupled policy gradient. Moreover, each agent interacts exclusively with its direct neighbors to obtain accurate policy parameters, while maintaining local estimates of other agents' parameters to execute its local policy and collect samples for optimization. These estimates and policy parameters are updated via a push-sum protocol, enabling distributed coordination of policy updates across the network. We prove that the joint policy produced by the proposed algorithm converges to a first-order stationary point of the objective function. Finally, the effectiveness of DSCP algorithm is demonstrated through simulations in a robot path planning environment, showing clear improvement over state-of-the-art methods.

cs.MA

Bayesian Probability Fusion for Multi-AP Collaborative Sensing in Mobile Networks

Integrated sensing and communication is widely acknowledged as a foundational technology for next-generation mobile networks. Compared with monostatic sensing, multi-access point (AP) collaborative sensing endows mobile networks with broader, more accurate, and resilient sensing capabilities, which are critical for diverse location-based sectors. This paper focuses on collaborative sensing in multi-AP networks and proposes a Bayesian probability fusion framework for target parameter estimation using orthogonal frequency-division multiplexing waveform. The framework models multi-AP received signals as probability distributions to capture stochastic observations from channel noise and scattering coefficients. Prior information is then incorporated into the joint probability density function to cast the problem as a constrained maximum a posteriori estimation. To address the high-dimensional optimization, we develop a prior-constrained gradient ascent (PCGA) algorithm that decouples correlated parameters and performs efficient gradient updates guided by the target prior. Theoretical analysis covers optimal fusion weights for global signal-to-noise ratio maximization, PCGA convergence, and the Cramer-Rao lower bound of the estimator, with insights applicable to broader fusion schemes. Extensive numerical simulations and real-world experiments with commercial devices show the framework reduces transmission overhead by 90% versus signal fusion and lowers estimation error by 41% relative to parameter fusion. Notably, field tests achieve submeter accuracy with 50% probability in typical coverage of mmWave APs. These improvements highlight a favorable balance between communication efficiency and estimation accuracy for practical multi-AP sensing deployment. The dataset is released for research purposes and is publicly available at: http://pmldatanet.com.cn/dataapp/multimodal

eess.SP