Searcharxiv⌕ Search

arXiv subjects

Zhaoyang Zhang

Publications and source records attributed to Zhaoyang Zhang.

At least 19 recordsLinked to original sources

MIMO-OFDM AI Receiver Based on Incrementally Conditioned Diffusion with Soft Decision

Conventional iterative receiver usually begins with channel estimation using sparse pilot observations and follows with data detection based on the channel estimates, and then updates channel estimation using decision feedback, and so on. In Artificial Intelligence (AI)-based receiver design, it is also crucial to make use of such progressively enriched data observations to enhance the generative channel estimation. However, the statistical characteristics and reliability of the data decisions always evolve with the channel estimation processes, which brings great challenges to the design of the overall learning framework and algorithms. In this paper, we propose \textit{Diff-Rx}, an incrementally conditioned diffusion-based receiver with co-designed data detection and channel estimation, for multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) systems. Specifically, we develop a condition-adaptive post-training method which enables the generative channel estimator to adapt to the conditioning inputs that are progressively enriched by soft data decisions, and also to implicitly align the pilot- and data-induced channel feature spaces to mitigate potential estimation errors. The threshold-free soft decisions provide smooth condition updates without condition-specific reliability tuning. We further develop a conditional diffusion transformer that is capable of performing robust channel estimation under noisy observations and various pilot patterns while reducing the conventional multi-step diffusion generation to one single step. Simulations on both statistical and site-specific ray-tracing channels show that, Diff-Rx exhibits consistent gains across different noise levels, pilot densities and modulation schemes, and works well at a pilot density as low as 1/32 while achieving significant improvement in channel estimation and data detection performances.

eess.SP↗

TelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks

Large language models (LLMs) offer great potential to automate a broad range of telecom engineering tasks by reasoning over standards, network configurations, mathematical models, source code, and operational logs. However, existing telecom LLMs struggle to reliably reason across these diverse tasks and data types. General-purpose LLMs often lack reliable grounding in telecom-specific knowledge, while telecom-specialized models are typically developed for narrower task families and exhibit limited multi-task performance. To fill this gap, we introduce TelecomGPT-R1, a family of open source unified telecom reasoning models structured around four complementary axes: protocol, knowledge, modeling, and fault. We first develop an axis-aware data generation framework that refines coarse public telecom artifacts into verified question-answer pairs and high quality chain-of-thought (CoT) reasoning trajectories, yielding a training corpus containing 104,880 examples. Building on this corpus, supervised fine-tuning (SFT) instills telecom knowledge and evidence-grounded reasoning patterns to overcome the cold start barrier for reinforcement learning (RL). We then apply dynamic sampling policy optimization (DAPO) with task-routed rubric rewards to keep RL updates informative and stable across heterogeneous telecom reasoning tasks. These rewards decompose axis-specific CoT traces into verifiable reasoning units and combine grounded dense process credit with outcome correctness, allowing RL to learn generalizable problem solving behaviors from verifiable telecom evidence. We release the TelecomGPT-R1 models and a reproducible training recipe to support further community development. Evaluations on seven benchmarks of the GSMA Open Telco Leaderboard show that the open-source TelecomGPT-R1-27B achieves an 89.64% mean score, outperforming leading proprietary models, including GPT-5, Claude, and Gemini.

cs.CL↗

Prediction--Loss Alignment for Sampler--Robust Flow Matching Training

Recent work has popularized a practical recipe in diffusion and flow matching: predict the clean signal $x$, convert it to a velocity, and train through a velocity-space loss. The conversion contains a singular endpoint amplification and therefore appears prone to unstable optimization, yet recent systems obtain strong empirical results with this recipe. We investigate this tension through the integrability of the pre-optimizer stochastic-gradient second moment. Under stated initialization conditions, the moment diverges under Uniform sampling; boundary-suppressing sampling can restore integrability under an additional upper-growth condition. We then show that prediction--loss alignment eliminates this conversion-induced source of non-integrability. Under a uniform moment bound, alignment yields a finite second moment for every timestep density, including Uniform sampling. Controlled experiments across continuous and binary settings reproduce the predicted sampler-dependent instability and show that aligned objectives remain trainable across the tested samplers. These results reconcile pointwise amplification with sampler-dependent empirical success and support alignment as a principled route to more robust flow-matching training.

cs.LG↗

RIDE: Relocalization-Informed Depth Estimation with 3D Gaussian Splatting

Render--match--PnP relocalization establishes correspondences between query image pixels and 3D map points for camera pose recovery, but their potential to support dense depth estimation is often overlooked. To exploit this geometric information, we present RIDE, which estimates dense metric depth from a robot's RGB stream. Given a metrically scaled 3D Gaussian Splatting (3DGS) model, RIDE combines sparse metric depth observations derived from PnP-RANSAC inlier correspondences with the geometric prior of a pretrained video-depth model. To handle uneven and intermittent observations, it integrates global and local depth correction with temporal memory, supporting depth estimation through short observation gaps after metric scale initialization. Trained on public RGB-D videos, RIDE is evaluated on robot sequences without fine tuning. Experiments show improved depth accuracy and temporal consistency over scale-only calibration, demonstrating how localization geometry can support both pose recovery and dense robot perception.

cs.RO↗

Floquet-sideband-enhanced shortwave electrometry with Rydberg atoms

Rydberg atomic electric-field sensors, under the framework of optical excitation and readout, can overcome the size-to-wavelength constraint imposed by the Chu limit. However, their sensitivity for decametric-wavelength shortwave electric fields is substantially lower than that for microwave ones, stemming from the off-resonant nature of low-frequency signals with Rydberg transitions. Here, we demonstrate a high-sensitivity heterodyne shortwave sensor based on microwave-dressed Rydberg atoms, leveraging precisely modulated Floquet sidebands. Around local- and microwave-field-engineered Floquet sidebands, the steep response gradient, arising from the enhanced atom-shortwave interaction through additionally created coherent channels, induces pronounced amplification of the heterodyne intermediate-frequency signal. As a result, compared to the same atomic heterodyne setup without microwave modulation, such Floquet-sideband-enhanced shortwave measurement boosts the sensitivity by four orders of magnitude, yielding a sensitivity of -122.7 dBm/Hz for shortwave at 30 MHz. This work offers a potential route to high-sensitive portable shortwave receivers in radio astronomy, radar and long-distance communications.

physics.atom-ph↗

DBRepro: Automated Database Synthesis via a Hybrid Constraint-Solving Approach for Reproducing Slow Queries

Slow queries frequently cause severe performance bottlenecks in database management systems. Diagnosing their root causes online risks exacerbating resource contention, while data privacy regulations often prohibit copying production data to test environments. Synthesizing a proxy database from non-intrusive metadata that induces the query optimizer to generate the same physical execution plans is therefore critical for offline diagnosis. High-fidelity reproduction requires preserving global statistical distributions while enforcing exact local cardinalities. Existing data-driven and workload-aware approaches cannot satisfy both requirements simultaneously. We present DBRepro, an automated end-to-end framework that formulates database generation as a constrained distribution synthesis problem. DBRepro initializes a global distribution from lightweight column statistics, extracts execution constraints from target queries, and progressively adjusts the distribution to satisfy these constraints while preserving the global distribution. Experiments on TPC-H and SSB show that DBRepro reduces cardinality error by up to 20.3% over a data-driven baseline while maintaining identical plan consistency. Compared with a workload-aware baseline, it reproduces 15% more consistent execution plans and reduces latency proportion error by 21.5%. We further validate DBRepro on a nearly 1 TB real-world dataset managed by KingbaseES, where it reproduces the execution performance of complex slow queries with high fidelity.

cs.DB↗

Recursive Flow: A Generative Framework for MIMO Channel Estimation

Channel estimation is a fundamental challenge in massive multiple-input multiple-output systems, where estimation accuracy governs the spectral efficiency and link reliability. In this work, we introduce Recursive Flow (RC-Flow), a novel solver that leverages pre-trained flow matching priors to robustly recover channel state information from noisy, under-determined measurements. Different from conventional open-loop generative models, our approach establishes a closed-loop refinement framework via a serial restart mechanism and anchored trajectory rectification. By synergizing flow-consistent prior directions with data-fidelity proximal projections, the proposed RC-Flow achieves robust channel reconstruction and delivers state-of-the-art performance across diverse noise levels, particularly in noise-dominated scenarios. The framework is further augmented by an adaptive dual-scheduling strategy, offering flexible management of the trade-off between convergence speed and reconstruction accuracy. Theoretically, we analyze the Jacobian spectral radius of the recursive operator to prove its global asymptotic stability. Numerical results demonstrate that RC-Flow reduces inference latency by two orders of magnitude while achieving a 2.7 dB performance gain in low signal-to-noise ratio regimes compared to the score-based baseline.

cs.IT↗

Cross-Domain Joint DDoS Detection in Multi-Controller SDN via Confidence-Based Entropy Fusion

In multi-controller Software-Defined Networking (SDN), Distributed Denial-of-Service (DDoS) attacks exhibit a "dispersed source, concentrated target" pattern across domains, i.e., attack traffic originates from multiple edge-controller domains but converges on a victim in a single aggregation controller domain. While entropy-based DDoS detectors are effective in single-controller settings, their direct application in multi-controller SDN reveals a previously overlooked anomaly. Through systematic experiments, we identify an aggregation bias: during the post-attack transition phase, the aggregation controller continues to generate excessive false positives, while edge controllers have already returned to normal. We attribute this phenomenon to the coupled effects of OpenFlow statistics lag and unconstrained dynamic-threshold drift. To address this issue, we propose a cross-domain confidence-fusion framework that leverages lightweight edge-side messages to calibrate aggregation-controller decisions without sharing raw traffic data. The framework is non-intrusive, communication-efficient, and incrementally deployable. Experiments on a three-controller linear Mininet testbed with 24 hosts over 10 runs show that the method preserves edge-controller performance while reducing the aggregation false positive rate from 8.87% to 1.96% and increasing the F1 score from 89.04% to 96.89%.

cs.CR↗

Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching

Model Inversion Attacks (MIAs) aim to reconstruct representative training samples of target identities from face recognition models, exposing critical security vulnerabilities. Existing methods typically rely on indirect guidance or highly stochastic guidance, making it difficult to stably optimize generation trajectories toward target facial images. In this paper, we propose Steering Flow Model Inversion (SFMI), a novel two-stage white-box model inversion method that reformulates inversion as a trajectory-steering task. Specifically, Step I, Learning a Generic Flow Matching Prior, pre-trains a generic unconditional Flow Matching model to encode the manifold of human faces as a robust prior. Step II, Attacking with Progressive Guidance Scheduler (PGS), injects time-dependent target-specific gradients during sampling. By backpropagating through the target model to obtain gradients from intermediate generated states, PGS progressively injects adaptive guidance signals into the vector field. This process effectively steers the current generative flow from random noise toward the high-density regions of the target class. Under an identity-disjoint cross-evaluation setting using the CelebA dataset, SFMI achieves an ACC of 0.9248, an FID of 22.61, and an LPIPS of 0.3874 on the ArcFace target. Extensive experiments on multiple target models demonstrate that SFMI achieves competitive state-of-the-art performance in attack success and visual fidelity under the evaluated white-box protocol.

cs.CV↗

When Time Meets Space: Entropy Integration and Dynamic Threshold for Adaptive DDoS Detection in SDN

Entropy-based Distributed Denial of Service (DDoS) detection in Software-Defined Networking (SDN) commonly relies on spatial traffic distributions and static or loosely adaptive thresholds, making it vulnerable to legitimate traffic fluctuations in Internet of Things (IoT) environments. This paper proposes a lightweight spatiotemporal entropy-based detector for DDoS attacks. Spatial entropy is computed from dynamically selected traffic attribute pairs, while temporal entropy captures the randomness of packet inter-arrival times. The two normalized entropy measures are fused into a unified indicator and evaluated using a constrained second-order Exponentially Weighted Moving Average threshold that jointly tracks entropy trend and volatility. To prevent attack-contaminated observations from biasing threshold adaptation, threshold updates are performed only for windows classified as normal. Testbed results show 99.26% recall, a 0.9737 F1-score, and a 3.2% false positive rate (41.74% below that of spatial entropy alone). On CICDDoS2019, the method achieves an FPR of 0 and remains competitive with machine-learning-based methods. It requires 3.95 ms of core processing per window and 11.65% system-wide CPU utilization, supporting resource-constrained edge and IoT deployment.

cs.CR↗

Federated Unlearning Over Wireless Networks

To comply with stringent data privacy regulations, federated unlearning (FU) has emerged as a critical paradigm. However, its implementation over wireless networks introduces severe communication latency and reliability challenges due to iterative calibration requirements and physical-layer channel uncertainties. In this paper, we investigate the problem of delay minimization for federated unlearning networks (FUN). Specifically, we establish a comprehensive system model that jointly incorporates the convergence behavior of the FUN algorithm, local device computation dynamics, and a worst-case robust transmission model operating under bounded channel state information (CSI) error. To solve the resulting non-convex joint resource allocation problem, we propose an efficient iterative algorithm. By exploiting the monotonicity and convexity properties of the system constraints, the problem is decomposed via a uniform scan over the local accuracy parameter, within which the optimal delay, bandwidth, power, and computation frequency are determined utilizing nested bisection and golden-section searches. Both theoretical analysis and extensive numerical results demonstrate that the proposed algorithm achieves polynomial complexity and significantly reduces the overall unlearning completion time compared to conventional baseline schemes.

cs.IT↗

Energy Efficiency Maximization for FAS-Assisted Downlink Communication in Mobile Embodied AI Networks (MEAN) over Interference Channels

In this paper, we investigate a fluid antenna system (FAS)-assisted downlink mobile embodied AI network (MEAN) over interference channels, where multiple base station (BS)-agent pairs reuse the same spectrum. The BSs employ FASs to improve the communication quality, while the mobile embodied artificial intelligence (AI) agents can adjust their positions according to environment-aware channel information, such as a channel-to-interference-plus-noise map (CINM). Considering both co-channel interference and the energy consumption caused by communication and agent movement, we formulate an energy efficiency (EE) maximization problem by jointly optimizing the agent positions, FAS port selections, and transmit powers. To solve this mixed-integer non-convex problem, we first derive the optimal transmit power in closed form for given agent positions and FAS ports. We then develop an iterative algorithm with adaptive FAS-port optimization and sequential agent-position optimization, together with a low-complexity power-update method. Simulation results demonstrate that the proposed design outperforms the considered benchmark schemes and provides improved feasibility under severe noise conditions.

eess.SP↗

Analogical Learning for Cross-Scenario Generalization: Framework and Application to Intelligent Localization

Modern learning systems often struggle with joint learning across diverse scenarios and immediate adaptation to new ones, because they rely heavily on the scenario-dependent absolute data-label representations. Here, we propose analogical learning (AL), a learning framework that explores the inherent invariance of the underlying physical processes across scenarios, to improve the cross-scenario generalization. Specifically, we introduce the physical concepts of reference frames and relativity into the neural modeling. The resultant framework explicitly employs intra-scenario data-label pairs as reference anchors and enforces the network to mediate its data-to-label transformation through data-domain relative metrics that factor out the scenario-dependent variations. We instantiate AL with Mateformer, a bipartite Transformer-based neural architecture. Each layer of the auxiliary Transformer extracts certain feature space of the current data, while the corresponding layer of the primary Transformer computes attention among the data feature space and then use it as a relativity metric to weight the current label feature to synthesize the next label feature and, ultimately, the final prediction. We apply AL to intelligent wireless localization, a representative multi-scenario learning task. Across synthetic, real-world, and city-scale datasets, AL enables robust cross-scenario transfer and multi-scenario joint learning, achieving wavelength-scale localization accuracy that matches or surpasses state-of-the-art methods. This physics-inspired learning framework provides a promising alternative for other cross-scenario learning tasks and applications.

cs.LG↗

P-Flow: Proxy-gradient Flows for Linear Inverse Problems

Generative models based on flow matching have emerged as a powerful paradigm for inverse problems, offering straighter trajectories and faster sampling compared to diffusion models. However, existing approaches often necessitate differentiating through unrolled paths, leading to numerical instability and prohibitive computational overhead. To address this, we propose P-Flow, a framework that stabilizes the reconstruction process by leveraging a proxy gradient to update the source point. This approach effectively circumvents the numerical instability and memory overhead of long-chain differentiation. To ensure consistency with the prior distribution, we employ a Gaussian spherical projection motivated by the concentration of measure phenomenon in high-dimensional spaces. We further provide a theoretical analysis for P-Flow based on Bayesian theory and Lipschitz continuity. Experiments across diverse restoration tasks demonstrate that P-Flow delivers competitive performance, especially under extreme degradations such as severely ill-posed conditions and high measurement noise.

cs.LG↗

Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models

End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LLM--TTS pipelines. Although these tokens support expressive and low-latency spoken interaction, they may also preserve sensitive speaker characteristics. We investigate whether exposed speech tokens leak voiceprints and formulate this risk as a speaker inversion attack. We introduce Audio BERT (AuB), a trainable model that constructs token embeddings from discrete codebooks and aggregates them into speaker-sensitive representations, and propose SpInv, a two-stage inversion method built on AuB to recover embeddings in the space of an attacker-specified speaker encoder. We evaluate Moshi, Higgs3, Kimi-Audio, and Qwen3-Omni using speaker-disjoint protocols on the VoxCeleb dataset. Extensive experiments show that, with only three seconds of frontend output, SpInv achieves cosine similarities above 0.70 in the attacker-specified speaker-encoder space.

cs.SD↗

Agentic AI-RAN Empowering Synergetic Sensing, Communication, Computing, and Control

Future sixth-generation (6G) networks are expected to support low-altitude wireless networks (LAWNs), where unmanned aerial vehicles (UAVs) and aerial robots operate in highly dynamic three-dimensional environments under stringent latency, reliability, and autonomy requirements. In such scenarios, autonomous task execution at the network edge demands holistic coordination among sensing, communication, computing, and control (SC3) processes. Agentic Artificially Intelligent Radio Access Networks (Agentic AI-RAN) offer a promising paradigm by enabling the edge network to function as an autonomous decision-making entity for low-altitude agents with limited onboard resources. In this article, we propose a task-oriented Agentic AI-RAN architecture that enables SC3 task execution within a single edge node. The proposed architecture addresses the challenge of coordinating heterogeneous workloads in resource-constrained edge environments. To validate this framework, we prototype a representative low-altitude UAV system on a general-purpose Graphics Processing Unit (GPU) platform and evaluate it through an autonomous drone-navigation case study. The current prototype instantiates the platform-agnostic design through Multi-Instance GPU (MIG) partitioning and containerized deployment, providing physical resource isolation and coordinated execution between real-time communication and multimodal inference. Experimental results demonstrate low closed-loop latency, robust bidirectional communication, and stable performance under dynamic runtime conditions, highlighting the feasibility of the proposed framework for mission-critical low-altitude wireless networks in 6G.

eess.SY↗

Digital Twin Synchronization Over Mobile Embodied AI Network With Agentic Intelligence

Efficient digital twin (DT) synchronization relies on maintaining high-fidelity virtual representations with minimal age of information (AoI). However, the synergistic potential of cooperative sensing and autonomous mobility of the sensing agent remains underexplored in existing DT synchronization frameworks. In this paper, we propose an agentic AI-empowered mobile embodied AI network (MEAN) framework for DT synchronization. In the proposed hybrid architecture, the base station (BS) conducts global orchestration, while the agents autonomously execute a five-stage closed-loop workflow: move-to-sense, cooperative sensing, onboard semantic processing, channel-aware mobility, and uplink transmission. To optimize synchronization performance, we formulate a joint topology dispatching and multidimensional resource allocation problem aimed at minimizing the maximum twin deviation across regions, subject to heterogeneous sensing fidelity and energy budget constraints. To tackle this, we develop a hierarchical two-layer optimization algorithm, where the outer-layer refines multi-agent assignment via a dynamic matching game, and the inner-layer iteratively optimizes the continuous resources. Extensive simulation results verify the convergence of the proposed algorithm and demonstrate its substantial superiority over multiple baseline schemes in reducing synchronization deviation. Furthermore, the results reveal that semantic compression serves as a vital substitute for channel resources in latency reduction under constrained bandwidth, while autonomous velocity adaptation provides an essential degree of freedom for the system to navigate the fundamental energy-time trade-off.

cs.IT↗

AC$^2$P$^2$SL: Adaptive Communication-Computation Pipeline Parallel Split Learning over Edge Networks

In wireless edge networks, split learning (SL) enables base station (BS) to utilize the distributed data and computing power across user equipments (UEs) to achieve collaborative model training while protecting local data privacy. However, the inherent sequential execution of computation and communication processes in conventional SL usually leads to long training times. To overcome this limitation, this paper proposes an adaptive communication-computation pipeline parallel split learning (AC$^2$P$^2$SL) framework. By conceptualizing the communication and computation processes of UEs and the BS as a unified pipeline, AC$^2$P$^2$SL achieves fine-grained pipeline parallelism across multiple micro-batches. Through this approach, effective overlapping of communication and computation is achieved which results in significant reduction of the overall training latency. Moreover, by considering the system constraints in the communication, computation, and storage dimensions as well as the heterogeneity of UEs, we formulate a joint optimization problem to minimize the training time and propose a corresponding split and pre-allocation algorithm to further enhance the pipeline efficiency. Additionally, accounting for the practical dynamic environments for the UEs, we design an adaptive re-allocation strategy to enhance the system resilience. Extensive experimental results demonstrate the effectiveness and robustness of AC$^2$P$^2$SL in reducing training time while ensuring data privacy preservation.

cs.DC↗