SearcharxivSearch

arXiv subjects

Chau Yuen

Publications and source records attributed to Chau Yuen.

At least 19 recordsLinked to original sources

A Fully Wave-Domain Wideband MU MIMO OFDM Transmitter via Stacked Intelligent Metasurfaces

This paper proposes an advanced realization principle for wideband multiuser multiple-input multiple-output orthogonal frequency-division multiplexing (MU-MIMO OFDM) transmitters, where the conventional transmitter-side baseband chain is physically synthesized in the wave domain. For design and optimization purposes, this fully wave-domain wideband MU-MIMO OFDM transmitter implemented by a cascaded SIM structure is functionally partitioned into two cascaded SIM blocks. The first block, denoted as SIM_1, integrates symbol loading and channel-adaptive MU-MIMO precoding updated at the channel-coherence timescale, mapping the user streams to a virtual port-subcarrier representation. The second block, denoted as SIM_2, acts as an offline-configured sampling-rate modulator that materializes the inverse discrete Fourier transform (IDFT) and cyclic prefix (CP) insertion directly in the wave domain. This baseband-free architecture establishes a virtual-to-physical transition from information bits to radiated CP-extended OFDM waveforms. To account for practical nonidealities, SIM_2 is optimized to fit the ideal multi-port CP-OFDM operator, and its residual response is mapped into an effective coupling matrix. Then, SIM_1 is optimized in a communication-oriented manner by jointly adapting discrete phase shifts and stream-subcarrier power loading to maximize the sum spectral efficiency. Results demonstrate the convergence, architecture trade-off, wave-domain OFDM materialization accuracy, and competitive performance of the proposed baseband-free transmitter.

eess.SP

Multi-Stream Spatiotemporal Channel Coding for MIMO Systems: Transmission Scheme Design and Achievable Rate Optimization

Spatiotemporal channel coding (STCC) can improve the achievable rate over traditional temporal channel coding (TCC) by leveraging spatial degrees of freedom to extend the codeword length. Although several information-theoretic foundations on STCC have been established, the investigation of transmission schemes from a communication-theoretic perspective remains in its early stages. This paper proposes a multi-stream over multi-subchannel STCC (STCC-MSC) under full channel state information assumption and optimizes its achievable rate in the finite blocklength regime. We first formulate the transmission architecture of STCC-MSC in a point-to-point MIMO system, which introduces a stream-subchannel matching mechanism. We then maximize the achievable rate of STCC-MSC by jointly optimizing the subchannel assignment and power allocation strategies, which is formulated as a mixed-integer-nonlinear-programming problem. Next, a penalized alternating convex approximation (PACA) algorithm is proposed to solve this problem. Subsequently, we extend the point-to-point STCC-MSC designs to the more general multi-user MIMO systems, including both uplink and downlink scenarios. Finally, simulation results indicate that the PACA algorithm achieves a 9.85% rate improvement over the benchmark algorithm within the STCC-MSC scheme. Furthermore, the joint STCC-MSC-PACA scheme improves the achievable rate by 28.68% over TCC scheme.

eess.SP

Multi-Hop RIS ISAC for Target Positioning: A Tensor Decomposition-based Approach

Reconfigurable intelligent surface (RIS) has demon- strated remarkable potential to enhance the performance of integrated sensing and communication (ISAC), particularly when the line-of-sight (LoS) paths are obstructed. By controlling the reconfigurable elements on the surface, RIS can establish virtual LoS paths and provide considerable passive beamforming gains, thereby significantly improving the received signal quality. In this paper, we design a novel multi-hop RIS ISAC system for target positioning, where multiple RISs are deployed to assist the communication from a transmitter to associated users while simultaneously enhancing receiver sensing performance in target positioning. Specifically, we formulate an optimization problem to minimize the root mean square error (RMSE) of the target detection while guaranteeing the communication requirements of the users. To solve this problem, we first unfold the cascaded sens- ing channel through parallel factor decomposition, and develop a low-rank CANDECOMP/PARAFAC decomposition (CPD)-based scheme to extract the location parameters (i.e., angle of arrival, angle of departure and delay) of the sensing targets. Then, we develop a scheme for jointly selecting the transmit beamforming and RIS phase shift configurations to maximize the sensing energy at the receiver, which in turn leads to improved accuracy in target positioning. We also provide a uniqueness analysis, complexity analysis, and Cram\'er-Rao lower bound (CRLB) of the parameters estimated by our methodology. Simulation results validate the improvement in target positioning obtained by our design relative to baselines.

eess.SP

Baseband-Free Wideband MU-MIMO OFDM via Stacked Intelligent Metasurfaces

This paper proposes a novel baseband-free wideband MU-MIMO OFDM transmitter architecture enabled by two cascaded stacked intelligent metasurfaces (SIMs). Unlike conventional wireless transmitters, the proposed design shifts symbol loading, MU-MIMO precoding, and OFDM modulation from digital baseband to the wave domain. Specifically, in the first SIM (SIM$_1$), a programmable symbol-loading interface first loads one stacked virtual-subcarrier coefficient vector onto a common monochromatic carrier, while the remaining layers perform channel-adaptive MU-MIMO precoding from the data streams to the port-subcarrier domain. Then, the second SIM (SIM$_2$) realizes offline-configured wave-domain OFDM modulation by implementing the inverse discrete Fourier transform (IDFT) and cyclic-prefix (CP) insertion operator directly in the wave domain. Based on this architecture, we develop a wave-domain system model that explicitly characterizes the virtual-to-physical transition from block-level coefficients to radiated OFDM subcarriers. The designs of SIM$_1$ and SIM$_2$ are formulated as two operator-fitting problems, respectively targeting a block-diagonal wideband precoder and the ideal OFDM modulation operator. Under practical discrete phase constraints, we develop a quantization-aware training-based gradient descent (QAT-GD) framework for both SIM stages. Numerical results verify the effectiveness of the proposed optimization, the correctness of the synthesized wave-domain functionalities, and the strong end-to-end MU-MIMO OFDM performance of the resulting architecture.

eess.SP

Task-Oriented Wave Processing with Stacked Intelligent Metasurfaces: Framework, Fusion, and Challenges

The deep integration of diverse services in sixth-generation (6G) networks poses significant challenges to conventional task-agnostic channels, often resulting in performance conflicts. To resolve these bottlenecks, this article introduces a physical-layer computing paradigm enabled by stacked intelligent metasurfaces (SIMs), transforming the wireless environment from a passive medium into a programmable signal processor. Specifically, we establish a unified framework to map high-level service requirements directly to wave-domain synthesis. We then investigate the fusion of diverse services, demonstrating how the deep computational architecture of SIMs resolves resource conflicts in integrated sensing and communication (ISAC) and integrated communication and computation (ICC) scenarios. Furthermore, we critically analyze fundamental challenges, including diffractive channel modeling and inverse task-to-phase mapping, while validating through numerical results that this approach elevates the system from simple coexistence to true service symbiosis. Finally, we discuss key research directions to pave the way for service-native 6G architectures.

cs.IT

CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning

Structured pruning compresses large language models (LLMs) by removing whole computational units, such as attention heads and feed-forward (FFN) channel groups. Most training-free methods, however, rank these units independently, implicitly treating the loss from pruning a set as the sum of its individual losses. This view fails for Transformers, whose sublayers are coupled through a shared residual stream. Two individually weak units can thus be jointly indispensable, yet independent scoring is blind to such dependence and removes them together. We introduce CoCurve (Cross-Module Co-Pruning Curvature), a calibration-only, fine-tuning-free method that prunes attention and FFN units jointly. A second-order Taylor expansion of the token-level KL between the frozen model and its masked copy yields a single Fisher matrix whose diagonal is classical node saliency and whose off-diagonal entries are co-pruning curvature edges: the extra damage of removing two units together. Under a single-ablation additivity approximation this matrix reduces to a Gram product of single-unit ablation features, so the full M x M interaction is recovered from M forward passes, with no pairwise sweeps or gradients. Pruning then reduces to one budgeted quadratic program, solved in a single shot under a shared attention--FFN budget, with no labels, fine-tuning, or recovery.

cs.LG

Task-Oriented Sensing and Covert Transmissions for Collaborative Multi-AUV Systems

In underwater covert cooperative missions, autonomous underwater vehicles (AUVs) often cannot rely on active sonar to continuously obtain complete information, since active sensing and frequent communications increase the risk of exposure. As a result, AUVs primarily rely on passive observation, an approach that yields incomplete local perception and limited task efficiency. Although underwater acoustic communications can mitigate this limitation through information sharing, they are simultaneously constrained by long delays, severe interference, low reliability, and the risk of covert exposure. Existing communications-oriented multi-agent reinforcement learning (MARL) studies often model communication as an ideal information flow, whereas traditional communication optimization primarily focuses on link-level performance. However, both are insufficient to characterize the actual contribution of perceptual information to cooperative tasks under realistic conditions of covert physical communications. This paper proposes a Sensed Information Value Realization Multi-Agent Reinforcement Learning (SVR-MARL) framework that leverages practical information to characterize the utility of information for cooperative tasks and learns distributed cooperative policies under realistic communication and covert constraints. Through a case study of covert multi-AUV cooperative localization and tracking, the potential of the proposed framework to improve collaborative task efficiency while reducing unnecessary communication and exposure risks is demonstrated.

cs.LG

CarbonCLIP: Enhance Carbon Prediction from Satellite Imagery via Integrated Street-View Semantics and Temporal Context Training

Accurately estimating urban carbon emissions is critical for sustainable urban planning, yet many existing approaches remain difficult to apply consistently across cities due to data-source heterogeneity and the lack of fine-grained semantic-temporal context in remote sensing data. We propose CarbonCLIP, a task-oriented multimodal distillation framework that improves satellite-based carbon emission prediction by transferring contextual knowledge into a unified satellite representation through dual-branch contrastive learning. Unlike conventional methods that rely on static visual features, CarbonCLIP explicitly bridges the gap between top-down satellite views and ground-level human activities. Specifically, the spatial branch uses fine-grained textual descriptions automatically generated from street-view images by Large Multimodal Models (LMMs) to provide semantic priors reflecting building functions, infrastructure, and urban activities, while the temporal branch employs a month encoder to encode temporal priors associated with monthly emission variation. CarbonCLIP requires multimodal data only during the pretraining phase; during inference, it relies solely on satellite imagery, thereby supporting scalable deployment when ground-level data are unavailable at inference. Experiments on Beijing and Singapore demonstrate that CarbonCLIP outperforms baselines in both study cities. The results validate that our method effectively transfers multimodal knowledge into satellite representations, offering a robust solution for satellite-based urban carbon modeling.

cs.CV

Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits

Mechanistic interpretability often relies on component-level interventions to discover how a model produces a behavior. This guides attribution, capability knockout, and model pruning downstream to operate by scoring each unit by the effect of ablation in isolation. Such first-order scoring is natural when component importance is additive, but becomes misleading when a transformer self-repairs: after a primary component is removed, a dormant backup can take over, muting the primary's measured effect while the backup itself appears irrelevant on the intact model. We recast this failure as a recovery task, conditional circuit completion, and introduce Conditional Co-Ablation (CoAx), a label-free, output-grounded score that asks how much each remaining unit's ablation effect grows once a primary set has been removed. This conditional growth exposes the second-order interaction that single-unit scores discard. On the GPT-2-small IOI circuit, CoAx raises backup-head recovery from 0.33 to 0.91 ROC-AUC, outperforming all baselines, including self-repair-aware gradient scores (best 0.82); counterfactual patching verifies that the recovered heads causally carry the repair. The same label-free procedure transfers to induction across eight models. Beyond discovery, the recovered backups correct self-repair-masked attribution, identify the components required for capability knockout, and yield repair-aware structured pruning scaling from 124M to 7B. Component importance is therefore not merely an isolated-unit property: in robust circuits, the components that matter can become visible only under the interventions that make them necessary.

cs.LG

RobustMAD: Evaluating Real-World Robustness of Multimodal Small Language Models for Deployable Anomaly Detection Assistants

Multimodal industrial anomaly inspection assistants are a critical component of next-generation smart factories, enabling interactive vision-language-based querying. However, multimodal large language models remain impractical for on-site deployment due to prohibitive computational demands and privacy risks from cloud-based inference. Compact multimodal small language models (MSLMs) offer a deployable alternative, yet progress is constrained by the lack of comprehensive robustness analyses and meaningfully challenging benchmarks that reflect real-world industrial conditions. To address this gap, we develop RobustMAD, the first deployment-motivated benchmark, designed to comprehensively evaluate model robustness through diverse open-ended queries spanning object understanding, anomaly detection, unanswerable problems, and visual quality degradations. Contrary to conventional assumptions, top-performing MSLMs exhibit promising capabilities, surprisingly outperforming even the larger GPT-5 Nano. However, they still fall short of safety-critical requirements, and RobustMAD reveals critical robustness gaps that pose operational risks. In particular, three recurring failure modes emerge: (i) fragile multimodal grounding under fine-grained distinctions or degraded visual conditions, (ii) insufficiently comprehensive responses, and (iii) weak logical grounding on unanswerable or ill-posed queries, leading to hallucinated outputs. Grounded in these insights, we provide actionable guidance for the design of next-generation multimodal industrial inspection assistants that leverage their promising competence. Code is available at https://github.com/en-research/RobustMAD.

cs.LG

XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge Composition

Large Language Models (LLMs) are increasingly deployed for knowledge synthesis, yet their capacity for compositional generalization in scientific knowledge remains under-characterized. Existing benchmarks primarily focus on single-turn restricted scenarios, failing to capture the capability boundaries exposed by real-world interactive scientific workflows. To address this, we introduce XDomainBench, a diagnostic benchmark for interactive interdisciplinary scientific reasoning. We formalize the composition order and mixture structure to enable systematic stress-testing from single-discipline to inter-disciplinary, comprising 8,598 interactive sessions across 20 domains and 4 task categories, with 8 realistic trajectory patterns covering difficulty and domain-mixture dynamics, simulating real AI4S scenarios. Large-scale evaluation of LLMs reveals a systematic reasoning collapse as composition order increases, stemming from two root causes: (i) direct difficulty increases induced by domain composition, and (ii) indirect interaction-amplified failures where trajectory patterns trigger error accumulation, reasoning breaks, and domain confusion, ultimately leading to session collapse.

cs.AI

Signal-Dependent Shot Noise Modeling of Rydberg Atomic Quantum Receivers: A Design Perspective

In this paper, we develop a communication-oriented complex baseband equivalent model for superheterodyne Rydberg atomic quantum receivers (RAQRs). The model explicitly captures photodetection-induced signal-dependent shot noise and its coupling with the optical operating point. By leveraging an atomic superheterodyne architecture and a strong local oscillator, we construct a complex baseband representation for both the received signal and the signal-dependent shot noise under both direct incoherent optical detection and balanced coherent optical detection. The derived model reveals that the optical operating point jointly determines the normalized effective receive gain and the equivalent noise background, thereby establishing a traceable gain-noise tradeoff governed by system design. More importantly, the proposed model shows that neglecting signal-dependent shot noise may lead to inaccurate operating-point design. Finally, by extending to the multiple-input-multiple-output (MIMO) case, we derive a lower bound on the achievable rate while considering the signal-dependent shot noise. Our analysis \textcolor{black}{reveals} that the non-zero asymptotic rate of RAQ-MIMO and its superiority over conventional RF-MIMO hinge on the normalized noise floor of the RAQ receive chain falling below that of RF MIMO. Simulation results validate our analysis and yield practical, closed-form design guidelines for RAQR front ends, revealing parameter regimes in which RAQ-MIMO outperforms conventional MIMO systems.

eess.SP

The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs

On-policy distillation (OPD) is widely used for LLM post-training. When pushed with a reward-extrapolation coefficient lambda > 1, the student can lift past the teacher in domain, but past a threshold lambda* the same step violates the output contract on structured-output tasks. In a single-position Bernoulli reduction, we derive a closed-form base-relative clip-safety threshold lambda*(p,b,c) determined by three measurable quantities: the teacher modal probability, the warm-start mass, and the importance-sampling clip strength. Above lambda*, the extrapolated fixed point exits the clip-safe region, changing training from format-preserving to format-collapsing. We extend the rule to calibrated K-ary listwise JSON tasks where a single binding equivalence class dominates the output contract and SFT retains parse headroom. On Amazon Fashion, three pre-registered tests--a fine-grid cliff interval, a budget-extension test, and a small-clip cross-prediction--fall within their locked prediction windows, with the small-clip value matching the closed-form prediction below grid resolution. Operating just below lambda*, ListOPD brings a 1.7B Qwen3 student to in-domain parity with an 8B-SFT baseline at one-fifth the parameters. The gain is driven primarily by format adherence: NDCG@1 on parsed outputs remains flat across lambda, while parse validity sharply changes at the predicted boundary. The cliff diagnostic is rubric-independent, whereas the parity claim uses a Gemini-graded rubric and inherits that evaluator's exposure.

cs.LG

Uplink Signal Detection For Large-Scale MIMO-ISAC Systems

Next-generation wireless communication systems are unifying large-scale multiple-input multiple-output (MIMO) and integrated sensing and communication (ISAC) to enhance sensing and communication performance. In this paper, the signal detection problem for MIMO-ISAC systems is modeled as a mixed-integer least squares (MILS) problem. To solve it efficiently, we propose a projection-based neighborhood search-aided alternating direction method of multipliers (P-NS-ADMM) detection scheme. By theoretical analysis, we demonstrate that P-NS-ADMM achieves the same received diversity order as maximum likelihood (ML) detection. For further complexity reduction, an iteration-based NS-ADMM (I-NS-ADMM) is proposed to remove the complex projection operation. Complexity analysis shows its complexity advantage compared with P-NS-ADMM. Moreover, to better estimate the sensing signals for I-NS-ADMM, a flexible mechanism of ADMM iterations is given. Finally, simulations demonstrate the proposed NS-aided ADMM detection schemes have significant performance advantages in terms of both BER and NMSE.

cs.IT

Tag-based Physical-Layer Authentication Against Message Interference

Tag-based Physical-Layer Authentication (PLA) has attracted significant attention in recent years due to its low complexity, high security, and low latency. Traditional tag-based PLA schemes typically estimate tags by decoding the message and then subtracting the estimation of the message from the received signal. However, these approaches suffer from two main limitations. First, decoding errors introduce message interference that degrades authentication performance. Second, the analytical complexity of decoding errors leads to sub-optimal threshold settings, thereby limiting detection probability. To address these limitations, this paper proposes a Tag-Based Challenge-Response (TBCR) scheme and a Series Cancellation Authentication (SCA) scheme. Specifically, in the TBCR scheme, the tags are superimposed on a forwarded challenge signal, enabling the receiver to estimate tags by removing the known challenge signal rather than relying on decoding. However, the challenge-response mechanism introduces extra noise. Here, we propose the SCA scheme without the noise interference, where both the series signal generation and cancellation modules are well-designed to generate authentication signals and estimate tags, respectively. Furthermore, we derive the closed-form expressions to evaluate the robustness and security of both proposed schemes. Notably, on one hand, the optimal threshold and detection probability are derived, which theoretically reveal that the SCA scheme always achieves the ideal detection performance, while the TBCR scheme does so in the absence of noise at Alice. On the other hand, the TBCR scheme provides enhanced security at high Signal-to-Noise Ratio (SNR) regions with fewer keys. Theoretical analysis and simulation demonstrate that both proposed schemes significantly outperform the benchmarks in detection probability with reduced time complexity.

cs.IT

Antenna Elements' Trajectory Optimization for Throughput Maximization in Continuous-Trajectory Fluid Antenna-Aided Wireless Communications

Fluid antenna (FA) systems offer novel spatial degrees of freedom (DoFs) with the potential for significant performance gains. Compared to existing works focusing solely on optimizing FA positions at discrete time instants, we introduce the concept of continuous-trajectory fluid antenna (CTFA), which explicitly considers the antenna element's movement trajectory across continuous time intervals and incorporates the inherent kinematic constraints present in practical FA implementations. Accordingly, we formulate the total throughput maximization problem in CTFA-aided wireless communication systems, addressing the joint optimization of continuous antenna trajectories in conjunction with the transmit covariance matrices under kinematic constraints. To effectively solve this non-convex problem with highly coupled optimization variables, we develop an iterative algorithm based on block coordinate descent (BCD) and majorization-minimization (MM) principles with the aid of the weighted minimum mean square error (WMMSE) method. Finally, numerical results are presented to validate the efficacy of the proposed algorithms and to quantify the substantial total throughput advantages afforded by the conceived CTFA-aided system compared to conventional fixed-position antenna (FPA) benchmarks and alternative approaches employing simplified trajectories.

eess.SP

Rydberg Atomic Quantum Receivers for Wireless Communications: Two-Color vs. Three-Color Excitation

An efficient three-color (3C) laser excitation-based Rydberg atomic quantum receiver (RAQR) architecture is investigated for wireless communications, utilizing a five-level (5L) electronic transition mechanism. Specifically, the conventional two-color (2C) RAQR with the four-level (4L) excitation faces three fundamental obstacles: 1) high cost and engineering challenges due to the reliance on unstable short-wavelength lasers; 2) a fundamental sensitivity limit in thermal atoms caused by residual Doppler broadening; and 3) the inability to detect low-frequency bands due to the energy-level constraint of two-photon resonance. To address these challenges, this paper analyzes a 3C5L-RAQR architecture with all-red/infrared lasers, which not only solves the engineering cost issues but also enables effective Doppler cancellation and low-frequency detection by exploiting the three-photon resonance. Bridging atomic physics and communication theory, an end-to-end equivalent baseband signal model is derived. Furthermore, the performance of different RAQR architectures is evaluated in terms of sensitivity, achievable rate and spectrum access range. Moreover, we provide an exact numerical solution for practical RAQRs by employing the Liouvillian superoperator formalism. Numerical results demonstrate that the exhibited 3C5L-RAQR achieves superior sensitivity compared to the conventional 2C4L-RAQR and a classical antenna-based radio frequency receiver for weak-signal detection. Finally, the inherent sensitivity-rate trade-off is revealed, showing that the 3C5L-RAQR is more suitable for deployment in power-limited communication scenarios demanding broad spectrum access.

cs.IT

Indirect and Direct Multiuser Hybrid Beamforming for Far-Field and Near-Field Communications: A Deep Learning Approach

Hybrid beamforming for extremely large-scale multiple-input multiple-output (XL-MIMO) systems is challenging in the near field because the channel depends jointly on angle and distance, and the multiuser interference (MUI) is strong. Existing deep learning methods typically follow either a decoupled design that optimizes analog beamforming without explicitly accounting for MUI, or an end-to-end (E2E) joint analog-digital optimization that can be unstable under nonconvex constant-modulus (CM), pronounced analog-digital coupling, and gradient pattern of sum-rate loss. To address both issues, we develop a complex-valued E2E framework based on a variant minimum mean square error (variant-MMSE) criterion, where the digital precoder is eliminated in closed form via Karush-Kuhn-Tucker (KKT) conditions so that analog learning is trained with a stable objective. The network employs a grouped complex-convolution sensing front-end for uplink (UL) measurements, a shared complex multi-layer perceptron (MLP) for per-user feature extraction, and a merged constant-modulus head to output the analog precoder. In the indirect mode, the network designs hybrid beamformers from estimated channel state information (CSI). In the direct mode where explicit CSI is unavailable, the network learns the sensing operator and the analog mapping from short pilots, after which additional pilots estimate the equivalent channel and enable a KKT closed-form digital precoder. Simulations show that the indirect mode approaches the performance of iterative variant-MMSE optimization with a complexity reduction proportional to the antenna number. In the direct mode, the proposed method improves spectral efficiency over sparse-recovery pipelines and recent deep learning baselines under the same pilot budget.

eess.SP