SearcharxivSearch

arXiv subjects

Wei Yu

Publications and source records attributed to Wei Yu.

At least 37 records · Page 2Linked to original sources

Uplink-Downlink Duality for Beamforming in Integrated Sensing and Communications

This paper considers the beamforming and power optimization problem for a class of integrated sensing and communications (ISAC) problems that utilize the communication signals simultaneously for sensing. We formulate the problem of minimizing the Bayesian Cramér-Rao bound (BCRB) on the mean-squared error of estimating a vector of parameters, while satisfying downlink signal-to-interference-and-noise-ratio constraints for a set of communication users at the same time. The proposed optimization framework comprises two key new ingredients. First, we show that the BCRB minimization problem corresponds to maximizing beamforming power along certain sensing directions of interest. Second, the classical uplink-downlink duality for multiple-input multiple-output communications can be extended to the ISAC setting, but unlike the classical communication problem, the dual uplink problem for ISAC may entail negative noise power and needs to include an extra condition on the uplink beamformers. This new duality theory opens doors for efficient iterative algorithm for optimizing power and beamformers for ISAC.

cs.IT

SID: Sliding into Distribution for Robust Few-Demonstration Manipulation

Generalizing robotic manipulation across object poses, viewpoints, and dynamic disturbances is difficult, especially with only a few demonstrations. End-to-end visuomotor policies are expressive but data-hungry, while planning and optimization satisfy explicit constraints but do not directly capture the interaction strategies demonstrated by humans. We propose Sliding into Distribution (SID), a structured framework that learns an object-centric motion field from canonicalized demonstrations to iteratively slide the system toward the demonstrated manifold and into the reliable operating region of a lightweight egocentric execution policy, mitigating out-of-distribution (OOD) execution. The motion field provides large corrective motions when far from the demonstration manifold and naturally vanishes near convergence, enabling robust reaching under substantial pose and viewpoint shifts. Within the reached regime, an egocentric policy trained with conditioned flow matching performs task-specific manipulation, supported by kinematically consistent point-cloud reprojection augmentation that preserves action-observation consistency. Across six real-world tasks, SID achieves approximately 90% success under OOD initializations with only two demonstrations, with under a 10% drop under distractors and external disturbances. Overall, SID provides a new paradigm for few-shot manipulation: explicitly managing distribution shift via online distribution recovery.

cs.RO

EmambaIR: Efficient Visual State Space Model for Event-guided Image Reconstruction

Recent event-based image reconstruction methods predominantly rely on Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) to process complementary event information. However, these architectures face fundamental limitations: CNNs often fail to capture global feature correlations, whereas ViTs incur quadratic computational complexity (e.g., $O(n^2)$), hindering their application in high-resolution scenarios. To address these bottlenecks, we introduce EmambaIR, an Efficient visual State Space Model designed for image reconstruction using spatially sparse and temporally continuous event streams. Our framework introduces two key components: the cross-modal Top-k Sparse Attention Module (TSAM) and the Gated State-Space Module (GSSM). TSAM efficiently performs pixel-level top-k sparse attention to guide cross-modal interactions, yielding rich yet sparse fusion features. Subsequently, GSSM utilizes a nonlinear gated unit to enhance the temporal representation of vanilla linear-complexity ($O(n)$) SSMs, effectively capturing global contextual dependencies without the typical computational overhead. Extensive experiments on six datasets across three diverse image reconstruction tasks - motion deblurring, deraining, and High Dynamic Range (HDR) enhancement - demonstrate that EmambaIR significantly outperforms state-of-the-art methods while offering substantial reductions in memory consumption and computational cost. The source code and data are publicly available at: https://github.com/YunhangWickert/EmambaIR

cs.CV

Active MIMO Sensing With Exploration-Exploitation Tradeoff

This paper develops an active sensing framework for designing the transmit and receive beamformers of a multiple-input multiple-output (MIMO) radar system. In the proposed technique, the beamformers are adaptively designed in each sensing stage based on the measurements made in the previous sensing stages. The beamformers are determined by minimizing the Bayesian Cram{é}r-Rao bound (BCRB) for the estimation of the unknown sensing parameters at each stage via Lagrangian dual optimization. To address the exploration-exploitation tradeoff that is inherent to such an adaptive design, this paper proposes two variants of the BCRB optimization problem: an exploration-centric variant, that ensures that multiple orthogonal beamforming directions are probed in each sensing stage, and an exploitation-centric variant, that does not restrict the number of optimal beamformers. Each variant of the optimization problem is solved via an alternating optimization algorithm that alternates between solving for the transmit beamformers and solving for the receive beamformers. The algorithm is shown to converge to a stationary point provided that each optimization problem is solved to global optimality. Moreover, this paper studies each of the two BCRB optimization sub-problems in the Lagrangian dual domain and shows that despite the non-convexity, global optimality is guaranteed provided that certain sufficient conditions hold. The conditions pertain to the multiplicity of the eigenvalues of a specific direction matrix that can be analytically written in terms of the optimal dual variables. These conditions further imply the tightness of the semidefinite relaxation of the optimization problems. Simulation results demonstrate the benefits of the proposed BCRB-based design compared to state-of-the-art adaptive beamforming strategies.

eess.SP

Multi-Carrier Modulation: An Evolution from Time-Frequency Domain to Delay-Doppler Domain

The recently proposed orthogonal delay-Doppler division multiplexing (ODDM) modulation, which is a delay-Doppler (DD) domain multi-carrier (DDMC) modulation scheme based on the DD domain orthogonal pulse (DDOP), is studied. We first revisit the linear time-varying (LTV) channel model for the wireless channel, and review the conventional multi-carrier (MC) modulation schemes and their design guidelines for both linear time-invariant (LTI) and LTV channels. We then focus on the representation of the LTV channel in an equivalent sampled DD (ESDD) domain, and propose an impulse-function-based transmission strategy for the ESDD channel. Next, we take an in-depth look into the DDOP and show that it achieves orthogonality with respect to the fine time and frequency resolutions in the ESDD domain thus behaves like an impulse function. This allows us to unveil the unique input-output relation of the resultant ODDM modulation over the ESDD channel. We point out that the conventional MC modulation design guidelines based on the Weyl-Heisenberg (WH) frame theory can be relaxed without compromising its orthogonality or violating the WH frame theory. More specifically, for a practical communication system with bandwidth and duration constraints, MC modulation signals can be designed considering so-called local or sufficient (bi)orthogonality, which refers to the (bi)orthogonality among a WH subset for the MC signal within a specific bandwidth and duration. This novel design guideline could potentially open up opportunities for developing future waveforms required by new applications such as communication systems associated with high delay and/or Doppler shifts, as well as integrated sensing and communications.

cs.IT

Reverberation lags viewed in hard X-rays from an accreting stellar-mass black hole

Accreting black holes are thought to swallow matter in the form of a disk and a hot cloud of plasma that glows brightly in X-rays, known as the corona. The X-ray emitting region is far too small to be directly imaged, but rapid variability of the X-ray signal can be used to infer the geometry by measuring time lags caused by material propagating towards the black hole and by coronal X-rays reflecting off the disk to imprint a reverberation lag. Reverberation lags can be recognized by characteristic spectral features, including an iron emission line at $\sim 6.4$ keV and a broad Compton hump peaking at $\sim 30$ keV. These reverberation features have both previously been detected for a few supermassive black holes in active galactic nuclei (AGNs). However, it is much more challenging to detect reverberation lags from stellar-mass black holes because they are more than a million times smaller. Previous reverberation lag measurements for stellar-mass black holes in X-ray binary systems have thus been limited to energies below 10 keV. Here we report on the first detection of the Compton hump reverberation feature from an X-ray binary, achieved by measuring lags in the broad energy range of $\sim 1-150$ keV. The accompanying detection of an iron line feature confirms the scenario of X-ray reverberation and provides strong evidence that the accretion flows in AGNs and X-ray binaries are governed by an ubiquitous process. Reverberation lags are prominent only in the most rapid variability, whereas lags in the slower variability are commonly attributed to propagating mass accretion rate perturbations. Our lag measurements up to the highest energy to date reveal that this lag in the slower variability evolves dramatically on timescales of days.

astro-ph.HE

Site-Specific Channel Modeling and Optimization of RIS-Assisted Multiuser MISO Systems

This paper presents a physics-based channel modeling and optimization framework for reconfigurable intelligent surface (RIS)-assisted downlink multi-user multiple-input single-output (MU-MISO) communication systems in site-specific environments. A hybrid ray-tracing (RT) and full-wave electromagnetic analysis approach is developed to construct a deterministic channel model that explicitly captures multipath propagation, RIS scattering behavior, and mutual coupling effects through a non-diagonal load impedance representation. Based on this model, an alternating optimization scheme jointly updates the base-station (BS) beamformer and RIS load impedances to maximize the minimum achievable rate under a total transmit power constraint and practical capacitance limits. The objective of the proposed framework is to provide a reliable initial assessment of the system-level impact of RIS deployment in realistic propagation scenarios. To evaluate this capability, the RIS is operated in a column-paired 1-bit control mode that enables exhaustive evaluation of all realizable configurations in both simulation and measurement. Performance is compared at the distribution level through achievable-rate histograms across all configurations and further examined under small user-location variations. The observed agreement between simulation and measurement demonstrates that the proposed framework reliably captures practical performance trends and provides useful guidance for the design and deployment of RIS-assisted MU-MISO systems in site-specific environments.

eess.SP

Refining Few-Step Text-to-Multiview Diffusion via Reinforcement Learning

Text-to-multiview (T2MV) diffusion models have shown great promise in generating multiple views of a scene from a single text prompt. While few-step backbones enable real-time T2MV generation, they often compromise key aspects of generation quality, such as per-view fidelity and cross-view consistency. Reinforcement learning (RL) finetuning offers a potential solution, yet existing approaches designed for single-image diffusion do not readily extend to the few-step T2MV setting, as they neglect cross-view coordination and suffer from weak learning signals in few-step regimes. To address this, we propose MVC-ZigAL, a tailored RL finetuning framework for few-step T2MV diffusion models. Specifically, its core insights are: (1) a new MDP formulation that jointly models all generated views and assesses their collective quality via a joint-view reward; (2) a novel advantage learning strategy that exploits the performance gains of a self-refinement sampling scheme over standard sampling, yielding stronger learning signals for effective RL finetuning; and (3) a unified RL framework that extends advantage learning with a Lagrangian dual formulation for multiview-constrained optimization, balancing single-view and joint-view objectives through adaptive primal-dual updates under a self-paced threshold curriculum that harmonizes exploration and constraint enforcement. Collectively, these designs enable robust and balanced RL finetuning for few-step T2MV diffusion models, yielding substantial gains in both per-view fidelity and cross-view consistency. Code is available at https://github.com/ZiyiZhang27/MVC-ZigAL.

cs.LG

MosaicMem: Hybrid Spatial Memory for Controllable Video World Models

Video diffusion models are moving beyond short, plausible clips toward world simulators that must remain consistent under camera motion, revisits, and intervention. Yet spatial memory remains a key bottleneck: explicit 3D structures can improve reprojection-based consistency but struggle to depict moving objects, while implicit memory often produces inaccurate camera motion even with correct poses. We propose Mosaic Memory (MosaicMem), a hybrid spatial memory that lifts patches into 3D for reliable localization and targeted retrieval, while exploiting the model's native conditioning to preserve prompt-following generation. MosaicMem composes spatially aligned patches in the queried view via a patch-and-compose interface, preserving what should persist while allowing the model to inpaint what should evolve. With PRoPE camera conditioning and two new memory alignment methods, experiments show improved pose adherence compared to implicit memory and stronger dynamic modeling than explicit baselines. MosaicMem further enables minute-level navigation, memory-based scene editing, and autoregressive rollout.

cs.CV

Rationale-Grounded In-Context Learning for Time Series Reasoning with Multimodal Large Language Models

The underperformance of existing multimodal large language models for time series reasoning lies in the absence of rationale priors that connect temporal observations to their downstream outcomes, which leads models to rely on superficial pattern matching rather than principled reasoning. We therefore propose the rationale-grounded in-context learning for time series reasoning, where rationales work as guiding reasoning units rather than post-hoc explanations, and develop the RationaleTS method. Specifically, we firstly induce label-conditioned rationales, composed of reasoning paths from observable evidence to the potential outcomes. Then, we design the hybrid retrieval by balancing temporal patterns and semantic contexts to retrieve correlated rationale priors for the final in-context inference on new samples. We conduct extensive experiments to demonstrate the effectiveness and efficiency of our proposed RationaleTS on three-domain time series reasoning tasks. We will release our code for reproduction.

cs.AI

TwinSegNet: A Digital Twin-Enabled Federated Learning Framework for Brain Tumor Analysis

Brain tumor segmentation is critical in diagnosis and treatment planning for the disease. Yet, current deep learning methods rely on centralized data collection, which raises privacy concerns and limits generalization across diverse institutions. In this paper, we propose TwinSegNet, which is a privacy-preserving federated learning framework that integrates a hybrid ViT-UNet model with personalized digital twins for accurate and real-time brain tumor segmentation. Our architecture combines convolutional encoders with Vision Transformer bottlenecks to capture local and global context. Each institution fine-tunes the global model of private data to form its digital twin. Evaluated on nine heterogeneous MRI datasets, including BraTS 2019-2021 and custom tumor collections, TwinSegNet achieves high Dice scores (up to 0.90%) and sensitivity/specificity exceeding 90%, demonstrating robustness across non-independent and identically distributed (IID) client distributions. Comparative results against centralized models such as TumorVisNet highlight TwinSegNet's effectiveness in preserving privacy without sacrificing performance. Our approach enables scalable, personalized segmentation for multi-institutional clinical settings while adhering to strict data confidentiality requirements.

cs.CV

WaterWave: Bridging Underwater Image Enhancement into Video Streams via Wavelet-based Temporal Consistency Field

Underwater video pairs are fairly difficult to obtain due to the complex underwater imaging. In this case, most existing video underwater enhancement methods are performed by directly applying the single-image enhancement model frame by frame, but a natural issue is lacking temporal consistency. To relieve the problem, we rethink the temporal manifold inherent in natural videos and observe a temporal consistency prior in dynamic scenes from the local temporal frequency perspective. Building upon the specific prior and no paired-data condition, we propose an implicit representation manner for enhanced video signals, which is conducted in the wavelet-based temporal consistency field, WaterWave. Specifically, under the constraints of the prior, we progressively filter and attenuate the inconsistent components while preserving motion details and scenes, achieving a natural-flowing video. Furthermore, to represent temporal frequency bands more accurately, an underwater flow correction module is designed to rectify estimated flows considering the transmission in underwater scenes. Extensive experiments demonstrate that WaterWave significantly enhances the quality of videos generated using single-image underwater enhancements. Additionally, our method demonstrates high potential in downstream underwater tracking tasks, such as UOSTrack and MAT, outperforming the original video by a large margin, i.e., 19.7% and 9.7% on precise respectively.

cs.CV

RIS-Assisted Joint Sensing and Communications via Fractionally Constrained Fractional Programming

This paper studies an uplink dual-functional sensing and communication system aided by a reconfigurable intelligent surface (RIS), whose reflection pattern is configured to trade-off sensing and communication functionalities. Specifically, the Bayesian Cramér-Rao lower bound (BCRLB) for estimating the azimuth angle of a sensing user is minimized while ensuring the signal-to-interference-plus-noise ratio constraints for communication users. We show that this problem can be formulated as a novel fractionally constrained fractional programming (FCFP) problem. To deal with this nontrivial optimization problem, we extend a quadratic transform technique, originally proposed to handle optimization problems containing fractional structures only in objectives, to the scenario where the constraints also include ratios. First, we consider the case where the fading coefficient is known. Using the quadratic transform, the FCFP problem can be turned into a sequence of subproblems that are convex except for the constant-modulus constraints which can be tackled using a penalty-based approach. To further reduce the computational complexity, we leverage the constant-modulus conditions and propose a novel linear transform. This new transform enables the FCFP problem to be turned into a sequence of linear programming (LP) subproblems, which can be solved efficiently. Then, we consider the case where the fading coefficient is unknown. A modified BCRLB is used to make the problem more tractable, and the proposed quadratic transform based algorithm is used to solve the problem. Numerical results unveil nontrivial and effective reflection patterns that can be synthesized by the RIS to facilitate both communication and sensing functionalities.

eess.SP

Capacity Bounds for Broadcast Channels with Bidirectional Conferencing Decoders

The two-user broadcast channel (BC) with receivers connected by bidirectional cooperation links of finite capacities, known as conferencing decoders, is considered. A novel capacity region outer bound is established based on multiple applications of the Csiszár-Körner identity. Achievable rate regions are derived by using Marton's coding as the transmission scheme, together with different combinations of decode-and-forward and quantize-bin-and-forward strategies at the receivers. It is shown that the outer bound coincides with the achievable rate region for a new class of semi-deterministic BCs with degraded message sets; for this class of channels, one-round cooperation is sufficient to achieve the capacity. Capacity result is also derived for a class of more capable semi-deterministic BCs with both common and private messages and one-sided conferencing. For the Gaussian BC with conferencing decoders, if the noises at the decoders are perfectly correlated (i.e., correlation is either 1 or -1), the new outer bound yields exact capacity region for two cases: i) BC with degraded message sets; ii) BC with one-sided conferencing from the weaker receiver to the stronger receiver. When the noises have arbitrary correlation, the outer bound is shown to be within half a bit from the capacity region for these same two cases. Finally, for the general Gaussian BC, a one-sided cooperation scheme based on decode-and-forward from the stronger receiver to the weaker receiver is shown to achieve the capacity region to within $\frac{1}{2}\log (\frac{2}{1-|λ|})$ bits, where $λ$ is the noise correlation. An interesting implication of these results is that for a Gaussian BC with perfectly negatively correlated noises and conferencing decoders with finite cooperation link capacities, it is possible to achieve a strictly positive rate using only an infinitesimal amount of transmit power.

cs.IT

Multimodal Visual Image Based User Association and Beamforming Using Graph Neural Networks

This paper proposes an approach that leverages multimodal data by integrating visual images with radio frequency (RF) pilots to optimize user association and beamforming in a downlink wireless cellular network under a max-min fairness criterion. Traditional methods typically optimize wireless system parameters based on channel state information (CSI). However, obtaining accurate CSI requires extensive pilot transmissions, which lead to increased overhead and latency. Moreover, the optimization of user association and beamforming is a discrete and non-convex optimization problem, which is challenging to solve analytically. In this paper, we propose to incorporate visual camera data in addition to the RF pilots to perform the joint optimization of user association and beamforming. The visual image data helps enhance channel awareness, thereby reducing the dependency on extensive pilot transmissions for system optimization. We employ a learning-based approach based on using first a detection neural network that estimates user locations from images, and subsequently two graph neural networks (GNNs) that extract features for system optimization based on the location information and the received pilots, respectively. Then, a multimodal GNN is constructed to integrate the features for the joint optimization user association and beamforming. Simulation results demonstrate that the proposed method achieves superior performance, while having low computational complexity and being interpretable and generalizable, making it an effective solution as compared to traditional methods based only on RF pilots.

eess.SP

Adaptive Dropout: Unleashing Dropout across Layers for Generalizable Image Super-Resolution

Blind Super-Resolution (blind SR) aims to enhance the model's generalization ability with unknown degradation, yet it still encounters severe overfitting issues. Some previous methods inspired by dropout, which enhances generalization by regularizing features, have shown promising results in blind SR. Nevertheless, these methods focus solely on regularizing features before the final layer and overlook the need for generalization in features at intermediate layers. Without explicit regularization of features at intermediate layers, the blind SR network struggles to obtain well-generalized feature representations. However, the key challenge is that directly applying dropout to intermediate layers leads to a significant performance drop, which we attribute to the inconsistency in training-testing and across layers it introduced. Therefore, we propose Adaptive Dropout, a new regularization method for blind SR models, which mitigates the inconsistency and facilitates application across intermediate layers of networks. Specifically, for training-testing inconsistency, we re-design the form of dropout and integrate the features before and after dropout adaptively. For inconsistency in generalization requirements across different layers, we innovatively design an adaptive training strategy to strengthen feature propagation by layer-wise annealing. Experimental results show that our method outperforms all past regularization methods on both synthetic and real-world benchmark datasets, also highly effective in other image restoration tasks. Code is available at \href{https://github.com/xuhang07/Adpative-Dropout}{https://github.com/xuhang07/Adpative-Dropout}.

cs.CV

Non-Intrusive Load Monitoring Based on Image Load Signatures and Continual Learning

Non-Intrusive Load Monitoring (NILM) identifies the operating status and energy consumption of each electrical device in the circuit by analyzing the electrical signals at the bus, which is of great significance for smart power management. However, the complex and changeable load combinations and application environments lead to the challenges of poor feature robustness and insufficient model generalization of traditional NILM methods. To this end, this paper proposes a new non-intrusive load monitoring method that integrates "image load signature" and continual learning. This method converts multi-dimensional power signals such as current, voltage, and power factor into visual image load feature signatures, and combines deep convolutional neural networks to realize the identification and classification of multiple devices; at the same time, self-supervised pre-training is introduced to improve feature generalization, and continual online learning strategies are used to overcome model forgetting to adapt to the emergence of new loads. This paper conducts a large number of experiments on high-sampling rate load datasets, and compares a variety of existing methods and model variants. The results show that the proposed method has achieved significant improvements in recognition accuracy.

cs.LG

RAW Image Reconstruction from RGB on Smartphones. NTIRE 2025 Challenge Report

Numerous low-level vision tasks operate in the RAW domain due to its linear properties, bit depth, and sensor designs. Despite this, RAW image datasets are scarce and more expensive to collect than the already large and public sRGB datasets. For this reason, many approaches try to generate realistic RAW images using sensor information and sRGB images. This paper covers the second challenge on RAW Reconstruction from sRGB (Reverse ISP). We aim to recover RAW sensor images from smartphones given the corresponding sRGB images without metadata and, by doing this, ``reverse" the ISP transformation. Over 150 participants joined this NTIRE 2025 challenge and submitted efficient models. The proposed methods and benchmark establish the state-of-the-art for generating realistic RAW data.

eess.IV