SearcharxivSearch

arXiv subjects

Wei Yu

Publications and source records attributed to Wei Yu.

At least 55 records · Page 3Linked to original sources

Quadratic Transform for Fractional Programming in Signal Processing and Machine Learning

Fractional programming (FP) is a branch of mathematical optimization that deals with the optimization of ratios. It is an invaluable tool for signal processing and machine learning, because many key metrics in these fields are fractionally structured, e.g., the signal-to-interference-plus-noise ratio (SINR) in wireless communications, the Cramér-Rao bound (CRB) in radar sensing, the normalized cut in graph clustering, and the margin in support vector machine (SVM). This article provides a comprehensive review of both the theory and applications of a recently developed FP technique known as the quadratic transform, which can be applied to a wide variety of FP problems, including both the minimization and the maximization of the sum of functions of ratios as well as matrix-ratio problems.

cs.IT

Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction

Video virtual try-on aims to seamlessly dress a subject in a video with a specific garment. The primary challenge involves preserving the visual authenticity of the garment while dynamically adapting to the pose and physique of the subject. While existing methods have predominantly focused on image-based virtual try-on, extending these techniques directly to videos often results in temporal inconsistencies. Most current video virtual try-on approaches alleviate this challenge by incorporating temporal modules, yet still overlook the critical spatiotemporal pose interactions between human and garment. Effective pose interactions in videos should not only consider spatial alignment between human and garment poses in each frame but also account for the temporal dynamics of human poses throughout the entire video. With such motivation, we propose a new framework, namely Dynamic Pose Interaction Diffusion Models (DPIDM), to leverage diffusion models to delve into dynamic pose interactions for video virtual try-on. Technically, DPIDM introduces a skeleton-based pose adapter to integrate synchronized human and garment poses into the denoising network. A hierarchical attention module is then exquisitely designed to model intra-frame human-garment pose interactions and long-term human pose dynamics across frames through pose-aware spatial and temporal attention mechanisms. Moreover, DPIDM capitalizes on a temporal regularized attention loss between consecutive frames to enhance temporal consistency. Extensive experiments conducted on VITON-HD, VVT and ViViD datasets demonstrate the superiority of our DPIDM against the baseline methods. Notably, DPIDM achieves VFID score of 0.506 on VVT dataset, leading to 60.5% improvement over the state-of-the-art GPD-VVTO approach.

cs.CV

Coded Downlink Massive Random Access and a Finite de Finetti Theorem

This paper considers a massive connectivity setting in which a base-station (BS) aims to communicate sources $(X_1,\cdots,X_k)$ to a randomly activated subset of $k$ users, among a large pool of $n$ users, via a common message in the downlink. Although the identities of the $k$ active users are assumed to be known at the BS, each active user only knows whether itself is active and does not know the identities of the other active users. A naive coding strategy is to transmit the sources alongside the identities of the users for which the source information is intended. This requires $H(X_1,\cdots,X_k) + k\log(n)$ bits, because the cost of specifying the identity of one out of $n$ users is $\log(n)$ bits. For large $n$, this overhead can be significant. This paper shows that it is possible to develop coding techniques that eliminate the dependency of the overhead on $n$, if the source distribution follows certain symmetry. Specifically, if the source distribution is independently and identically distributed (i.i.d.) then the overhead can be reduced to at most $O(\log(k))$ bits, and in case of uniform i.i.d. sources, the overhead can be further reduced to $O(1)$ bits. For sources that follow a more general exchangeable distribution, the overhead is at most $O(k)$ bits, and in case of finite-alphabet exchangeable sources, the overhead can be further reduced to $O(\log(k))$ bits. The downlink massive random access problem is closely connected to the study of finite exchangeable sequences. The proposed coding strategy allows bounds on the Kullback-Leibler (KL) divergence between finite exchangeable distributions and i.i.d. mixture distributions to be developed and gives a new KL divergence version of the finite de Finetti theorem, which is scaling optimal.

cs.IT

Rate-Distortion-Perception Tradeoff Based on the Conditional-Distribution Perception Measure

This paper studies the rate-distortion-perception (RDP) tradeoff for a memoryless source model in the asymptotic limit of large block-lengths. The perception measure is based on a divergence between the distributions of the source and reconstruction sequences \emph{conditioned} on the encoder output, first proposed by Mentzer et al. We consider the case when there is no shared randomness between the encoder and the decoder and derive a single-letter characterization of the RDP function for the case of discrete memoryless sources. This is in contrast to the marginal-distribution metric case (introduced by Blau and Michaeli), whose RDP characterization remains open when there is no shared randomness. The achievability scheme is based on lossy source coding with a posterior reference map. For the case of continuous valued sources under the squared error distortion measure and the squared quadratic Wasserstein perception measure, we also derive a single-letter characterization and show that the decoder can be restricted to a noise-adding mechanism. Interestingly, the RDP function characterized for the case of zero perception loss coincides with that of the marginal metric, and further zero perception loss can be achieved with a 3-dB penalty in minimum distortion. Finally we specialize to the case of Gaussian sources, and derive the RDP function for Gaussian vector case and propose a reverse water-filling type solution. We also partially characterize the RDP function for a mixture of Gaussian vector sources.

cs.IT

PID-GM: PID Control with Gain Mapping

Proportional-Integral-Differential (PID) control is widely used in industrial control systems. However, up to now there are at least two open problems related with PID control. One is to have a comprehensive understanding of its robustness with respect to model uncertainties and disturbances. The other is to build intuitive, explicit and mathematically provable guidelines for PID gain tuning. In this paper, we introduce a simple nonlinear mapping to determine PID gains from three auxiliary parameters. By the mapping, PID control is shown to be equivalent to a new PD control (serving as a nominal control) plus an uncertainty and disturbance compensator (to recover the nominal performance). Then PID control can be understood, designed and tuned in a Two-Degree-of-Freedom (2-DoF) control framework. We discuss some basic properties of the mapping, including the existence, uniqueness and invertibility. Taking as an example the PID control applied to a general uncertain second-order plant, we prove by the singular perturbation theory that the closed-loop steady-state and transient performance depends explicitly on one auxiliary parameter which can be viewed as the virtual singular perturbation parameter (SPP) of PID control. All the three PID gains are monotonically decreasing functions of the SPP, indicating that the smaller the SPP is, the higher the PID gains are, and the better the robustness of PID control is. Simulation and experimental examples are provided to demonstrate the properties of the mapping as well as the effectiveness of the mapping based PID gain turning.

eess.SY

NTIRE 2025 Challenge on Event-Based Image Deblurring: Methods and Results

This paper presents an overview of NTIRE 2025 the First Challenge on Event-Based Image Deblurring, detailing the proposed methodologies and corresponding results. The primary goal of the challenge is to design an event-based method that achieves high-quality image deblurring, with performance quantitatively assessed using Peak Signal-to-Noise Ratio (PSNR). Notably, there are no restrictions on computational complexity or model size. The task focuses on leveraging both events and images as inputs for single-image deblurring. A total of 199 participants registered, among whom 15 teams successfully submitted valid results, offering valuable insights into the current state of event-based image deblurring. We anticipate that this challenge will drive further advancements in event-based vision research.

cs.CV

Learning Beamforming Codebooks for Active Sensing with Reconfigurable Intelligent Surface

This paper explores the design of beamforming codebooks for the base station (BS) and for the reconfigurable intelligent surfaces (RISs) in an active sensing scheme for uplink localization, in which the mobile user transmits a sequence of pilots to the BS through reflection at the RISs, and the BS and the RISs are adaptively configured by carefully choosing BS beamforming codeword and RIS codewords from their respective codebooks in a sequential manner to progressively focus onto the user. Most existing codebook designs for RIS are not tailored for active sensing, by which we mean the choice of the next codeword should depend on the measurements made so far, and the sequence of codewords should dynamically focus reflection toward the user. Moreover, most existing codeword selection methods rely on exhaustive search in beam training to identify the codeword with the highest signal-to-noise ratio (SNR), thus incurring substantial pilot overhead as the size of the codebook scales. This paper proposes a learning-based approach for codebook construction and for codeword selection for active sensing. The proposed learning approach aims to locate a target in the service area by recursively selecting a sequence of BS beamforming codewords and RIS codewords from the respective codebooks as more measurements become available without exhaustive beam training. The codebook design and the codeword selection fuse key ideas from the vector quantized variational autoencoder (VQ-VAE) and the long short-term memory (LSTM) network to learn respectively the discrete function space of the codebook and the temporal dependencies between measurements.

eess.SP

Rate-Distortion-Perception Tradeoff for Gaussian Vector Sources

This paper studies the rate-distortion-perception (RDP) tradeoff for a Gaussian vector source coding problem where the goal is to compress the multi-component source subject to distortion and perception constraints. Specifically, the RDP setting with either the Kullback-Leibler (KL) divergence or Wasserstein-2 metric as the perception loss function is examined, and it is shown that for Gaussian vector sources, jointly Gaussian reconstructions are optimal. We further demonstrate that the optimal tradeoff can be expressed as an optimization problem, which can be explicitly solved. An interesting property of the optimal solution is as follows. Without the perception constraint, the traditional reverse water-filling solution for characterizing the rate-distortion (RD) tradeoff of a Gaussian vector source states that the optimal rate allocated to each component depends on a constant, called the water level. If the variance of a specific component is below the water level, it is assigned a zero compression rate. However, with active distortion and perception constraints, we show that the optimal rates allocated to the different components are always positive. Moreover, the water levels that determine the optimal rate allocation for different components are unequal. We further treat the special case of perceptually perfect reconstruction and study its RDP function in the high-distortion and low-distortion regimes to obtain insight to the structure of the optimal solution.

cs.IT

Spectral analysis of the X-ray flares in the 2023 outburst of the new black binary transient Swift J1727.8--1613 observed with Insight-HXMT

The new black hole transient Swift J1727.8--1613 exhibited a series of X-ray flares during its 2023 outburst extensively observed with Insight-HXMT. We analyze the spectra of the flaring period using a series of models consisting of a multi-color disk and several different non-thermal components, and several consistent conclusions are obtained among these models. First, Swift J1727.8--1613 was in the transition process from the hard intermediate state (HIMS) to the very high state (VHS) during the first flaring period (MJD 60197--60204), and afterwards it exhibited typical VHS parameter characteristics, such as high temperature of the disk inner radius and a steep power-law spectrum with a photon index of 2.6. Second, the flares in the VHS are characterized by a rapid increase in the flux of accretion disk, accompanied by a simultaneous rapid expansion of the inner radius, which could be apparent if the accretion disk hardening factor varies significantly. The strong power-law component during the VHS is likely produced by synchrotron self-Compton process in the relativistic jets, in agreement with the observed weak reflection component and lack of correlation with the disk component.

astro-ph.HE

First frequency phase transfer from the 3 mm to the 1 mm band on an Earth-sized baseline

Frequency Phase Transfer (FPT) is a technique designed to increase coherence and sensitivity in radio interferometry by making use of the non-dispersive nature of the troposphere to calibrate high-frequency data using solutions derived at a lower frequency. While the Korean VLBI Network has pioneered the use of simultaneous multi-band systems for routine FPT up to an observing frequency of 130 GHz, this technique remains largely untested in the (sub)millimeter regime. A recent effort has been made to outfit dual-band systems at (sub)millimeter observatories participating in the Event Horizon Telescope (EHT) and to test the feasibility and performance of FPT up to the observing frequencies of the EHT. We present the results of simultaneous dual-frequency observations conducted in January 2024 on an Earth-sized baseline between the IRAM 30-m in Spain and the JCMT and SMA in Hawai`i. We performed simultaneous observations at 86 and 215 GHz on the bright sources J0958+6533 and OJ287, with strong detections obtained at both frequencies. We observe a strong correlation between the interferometric phases at the two frequencies, matching the trend expected for atmospheric fluctuations and demonstrating for the first time the viability of FPT for VLBI at a wavelength of $\sim$1 millimeter. We show that the application of FPT systematically increases the 215 GHz coherence on all averaging timescales. In addition, the use of the co-located JCMT and SMA as a single dual-frequency station demonstrates the feasibility of paired-antenna FPT for VLBI for the first time, with implications for future array capabilities (e.g., ALMA sub-arraying and ngVLA calibration strategies).

astro-ph.IM

On Self-Adaptive Perception Loss Function for Sequential Lossy Compression

We consider causal, low-latency, sequential lossy compression, with mean squared-error (MSE) as the distortion loss, and a perception loss function (PLF) to enhance the realism of reconstructions. As the main contribution, we propose and analyze a new PLF that considers the joint distribution between the current source frame and the previous reconstructions. We establish the theoretical rate-distortion-perception function for first-order Markov sources and analyze the Gaussian model in detail. From a qualitative perspective, the proposed metric can simultaneously avoid the error-permanence phenomenon and also better exploit the temporal correlation between high-quality reconstructions. The proposed metric is referred to as self-adaptive perception loss function (PLF-SA), as its behavior adapts to the quality of reconstructed frames. We provide a detailed comparison of the proposed perception loss function with previous approaches through both information theoretic analysis as well as experiments involving moving MNIST and UVG datasets.

cs.LG

Correlated spectro-polarimetric study along the Z track in XTE J1701-462 puts constraints on its coronal geometry

Context. In September 2022, the transient neutron star low-mass X-ray binary XTE J1701-462 went into a new outburst. Aims. The objective of this work is to examine the evolution of the accretion geometry of XTE J1701-462 by studying the spectro-polarimetric properties along the Z track of this source. The simultaneous observations archived by the Insight-Hard X-ray Modulation Telescope (HXMT) and the Imaging X-ray Polarimetry Explorer (IXPE) give us the opportunity. Methods. We present a comprehensive X-ray spectro-polarimetric analysis of XTE J1701-462, using simultaneous observations from IXPE, Insight-HXMT and NuSTAR. For IXPE observations, two methods are employed to measure the polarization: a model-independent measurement with PCUBE and a model-dependent polarization-spectral analysis with XSPEC. The corresponding spectra from Insight-HXMT and NuSTAR are studied with two configurations that correspond to a slab-like corona and a spherical shell-like corona, respectively. Results. Significant polarization characteristics are detected in XTE J1701-462. The polarization degree shows a decreasing trend along the Z track, reducing from (4.84 $\pm$ 0.37)% to (3.76 $\pm$ 0.43)% on the horizontal branch and jumping to less than 1% on the normal branch. The simultaneous spectral analysis from Insight-HXMT and NuSTAR suggests that the evolution of the PD is closely linked to changes in the flux of the Comptonized component and its covering factor along the Z track, supporting a shrinking corona.

astro-ph.HE

MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-Reading

Lip-reading is to utilize the visual information of the speaker's lip movements to recognize words and sentences. Existing event-based lip-reading solutions integrate different frame rate branches to learn spatio-temporal features of varying granularities. However, aggregating events into event frames inevitably leads to the loss of fine-grained temporal information within frames. To remedy this drawback, we propose a novel framework termed Multi-view Temporal Granularity aligned Aggregation (MTGA). Specifically, we first present a novel event representation method, namely time-segmented voxel graph list, where the most significant local voxels are temporally connected into a graph list. Then we design a spatio-temporal fusion module based on temporal granularity alignment, where the global spatial features extracted from event frames, together with the local relative spatial and temporal features contained in voxel graph list are effectively aligned and integrated. Finally, we design a temporal aggregation module that incorporates positional encoding, which enables the capture of local absolute spatial and global temporal information. Experiments demonstrate that our method outperforms both the event-based and video-based lip-reading counterparts.

cs.CV

What Kind of Visual Tokens Do We Need? Training-free Visual Token Pruning for Multi-modal Large Language Models from the Perspective of Graph

Recent Multimodal Large Language Models(MLLMs) often use a large number of visual tokens to compensate their visual shortcoming, leading to excessive computation and obvious visual redundancy. In this paper, we investigate what kind of visual tokens are needed for MLLMs, and reveal that both foreground and background tokens are critical for MLLMs given the varying difficulties of examples. Based on this observation, we propose a graph-based method towards training-free visual token pruning, termed G-Prune.In particular, G-Prune regards visual tokens as nodes, and construct their connections based on their semantic similarities. Afterwards, the information flow is propagated via weighted links, and the most important tokens after iterations are kept for MLLMs, which can be front or background.To validate G-Prune, we apply it to a recent MLLM called LLaVA-NeXT, and conduct extensive experiments on a set of benchmarks.The experiment results show that G-Prune can greatly reduce computation overhead while retaining high performance on both coarse- and fine-grained tasks. For instance, G-Prune can reduce 63.57\% FLOPs of LLaVA-NeXT on VQA2.0 and TextVQA with only 0.95\% and 2.34\% accuracy drops, respectively.

cs.CV

Wireless 6G Connectivity for Massive Number of Devices and Critical Services

Compared to the generations up to 4G, whose main focus was on broadband and coverage aspects, 5G has expanded the scope of wireless cellular systems towards embracing two new types of connectivity: massive machine-type communication (mMTC) and ultra-reliable low-latency communications (URLLC). This paper discusses the possible evolution of these two types of connectivity within the umbrella of 6G wireless systems. The paper consists of three parts. The first part deals with the connectivity for a massive number of devices. While mMTC research in 5G predominantly focuses on the problem of uncoordinated access in the uplink for a large number of devices, the traffic patterns in 6G may become more symmetric, leading to closed-loop massive connectivity. One of the drivers for this is distributed learning/inference. The second part of the paper discusses the evolution of wireless connectivity for critical services. While latency and reliability are tightly coupled in 5G, 6G will support a variety of safety critical control applications with different types of timing requirements, as evidenced by the emergence of metrics related to information freshness and information value. Additionally, ensuring ultra-high reliability for safety critical control applications requires modeling and estimation of the tail statistics of the wireless channel, queue length, and delay. The fulfillment of these stringent requirements calls for the development of novel AI-based techniques, incorporating optimization theory, explainable AI, generative AI and digital twins. The third part analyzes the coexistence of massive connectivity and critical services. We will consider scenarios in which a massive number of devices need to support traffic patterns of mixed criticality. This is followed by a discussion about the management of wireless resources shared by services with different criticality.

cs.IT

Active Sensing for Multiuser Beam Tracking with Reconfigurable Intelligent Surface

This paper studies a beam tracking problem in which an access point (AP), in collaboration with a reconfigurable intelligent surface (RIS), dynamically adjusts its downlink beamformers and the reflection pattern at the RIS in order to maintain reliable communications with multiple mobile user equipments (UEs). Specifically, the mobile UEs send uplink pilots to the AP periodically during the channel sensing intervals, the AP then adaptively configures the beamformers and the RIS reflection coefficients for subsequent data transmission based on the received pilots. This is an active sensing problem, because channel sensing involves configuring the RIS coefficients during the pilot stage and the optimal sensing strategy should exploit the trajectory of channel state information (CSI) from previously received pilots. Analytical solution to such an active sensing problem is very challenging. In this paper, we propose a deep learning framework utilizing a recurrent neural network (RNN) to automatically summarize the time-varying CSI obtained from the periodically received pilots into state vectors. These state vectors are then mapped to the AP beamformers and RIS reflection coefficients for subsequent downlink data transmissions, as well as the RIS reflection coefficients for the next round of uplink channel sensing. The mappings from the state vectors to the downlink beamformers and the RIS reflection coefficients for both channel sensing and downlink data transmission are performed using graph neural networks (GNNs) to account for the interference among the UEs. Simulations demonstrate significant and interpretable performance improvement of the proposed approach over the existing data-driven methods with nonadaptive channel sensing schemes.

eess.SP

Covariance-Based Activity Detection in Cooperative Multi-Cell Massive MIMO: Scaling Law and Efficient Algorithms

This paper focuses on the covariance-based activity detection problem in a multi-cell massive multiple-input multiple-output (MIMO) system. In this system, active devices transmit their signature sequences to multiple base stations (BSs), and the BSs cooperatively detect the active devices based on the received signals. While the scaling law for the covariance-based activity detection in the single-cell scenario has been extensively analyzed in the literature, this paper aims to analyze the scaling law for the covariance-based activity detection in the multi-cell massive MIMO system. Specifically, this paper demonstrates a quadratic scaling law in the multi-cell system, under the assumption that the path-loss exponent of the fading channel $γ> 2.$ This finding shows that, in the multi-cell massive MIMO system, the maximum number of active devices that can be correctly detected in each cell increases quadratically with the length of the signature sequence and decreases logarithmically with the number of cells (as the number of antennas tends to infinity). Moreover, in addition to analyzing the scaling law for the signature sequences randomly and uniformly distributed on a sphere, the paper also establishes the scaling law for signature sequences based on a finite alphabet, which are easier to generate and store. Finally, this paper proposes two efficient accelerated coordinate descent (CD) algorithms with a convergence guarantee for solving the device activity detection problem. The first algorithm reduces the complexity of CD by using an inexact coordinate update strategy. The second algorithm avoids unnecessary computations of CD by using an active set selection strategy. Simulation results show that the proposed algorithms exhibit excellent performance in terms of computational efficiency and detection error probability.

cs.IT

Wake structures and performance of wind turbine rotor with harmonic surging motions under laminar and turbulent inflows

This study presents a comprehensive numerical analysis of a full-scale horizontal-axis Floating Offshore Wind Turbine (FOWT) subjected to harmonic surging motions under both laminar and turbulent inflow conditions. Utilizing high-fidelity Computational Fluid Dynamics (CFD) simulations, namely Large-Eddy Simulation (LES) with Actuator Line Model (ALM), this research investigates the rotor performance, wake characteristics, and wake structures of a surging FOWT in detail. The study delves into the influence of varying inflow turbulence intensities, surging settings, and their interplay on the aerodynamic performance and the wake aerodynamics of a FOWT rotor. The results show that, through employing the phase-locking technique, Surging Induced Periodic Coherent Structures (SIPCS) can be identified in the wake of all the surging cases studied, irrespective of the inflow conditions and the surging settings. Additionally, the findings show that the faster wake recovery observed in the surging-laminar cases is not caused by facilitating instability-induced faster wake breakdown, a previously accepted hypothesis. Instead, it is the enhanced advection process resulting from the induction fields of SIPCS that causes the wake to recover faster. The analysis of rotor performance shows that the time-averaged rotor performances are affected by the intricate dynamics arising from the surging motions. With certain surging settings, the time-averaged thrust and the time-averaged power of a surging rotor are found to be simultaneously lower and higher compared to those of a fixed rotor. Furthermore, the study underscores the importance of considering both the magnitude of surging and the rate of surging ($\mathcal{V}$ and $\mathcal{W}$) simultaneously to fully characterize the hysteresis load on a surging rotor.

physics.flu-dyn