SearcharxivSearch

arXiv subjects

Juan Cheng

Publications and source records attributed to Juan Cheng.

12 recordsLinked to original sources

FCUS-rPPG: A Fast-Converging Unsupervised Framework for Remote Photoplethysmography via Gradient Oscillation Suppression

Remote photoplethysmography (rPPG) enables non-contact extraction of blood volume pulse (BVP) signals using consumer-grade cameras. Recent unsupervised rPPG methods learn BVP representations without requiring ground-truth physiological annotations, yet their optimization is often hindered by noisy and unstable gradients, resulting in slow convergence and limited cross-domain generalization. In this paper, we propose FCUS-rPPG, a fast-converging unsupervised rPPG framework with strong generalization capability. Motivated by the observation that BVP representations exhibit both multi-spectral covariation and low-dimensional manifold structure, we design a spectrally shared backbone that facilitates BVP feature disentanglement while improving optimization efficiency. To jointly enhance convergence stability and generalization performance, we further develop a unified optimization framework operating at the gradient, loss-landscape, and feature-representation levels. Specifically, a post-verification masking mechanism filters out misleading gradients according to the weak-amplitude physiological prior of BVP signals; a perturbation-based loss landscape smoothing strategy steers optimization toward more generalizable flat minima; and a noise-aware null-space regularization constrains feature updates to the orthogonal complement of the noise subspace, thereby mitigating noise-induced representation drift. Extensive experiments on five datasets demonstrate that FCUS-rPPG requires only one training epoch, whereas existing methods typically require tens to hundreds of epochs. Notably, FCUS-rPPG consistently achieves state-of-the-art (SOTA) performance in cross-dataset evaluations. This study provides an efficient and robust solution to the real-world deployment of unsupervised rPPG. The source code will be publicly available at https://github.com/JiaJieLee/FCUS-rPPG.

cs.CV

Adaptive Physical-Facial Representation Fusion via Subject-Invariant Cross-Modal Prompt Tuning for Video-Based Emotion Recognition

Emotion recognition from facial videos enables non-contact inference of human emotional states. Although facial expressions are widely used cues, they cannot fully reflect intrinsic affective states. Remote photoplethysmography (rPPG) provides complementary physiological information, but it is highly susceptible to noise and inter-subject variability, limiting generalization to unseen individuals. Existing multimodal methods combine facial and rPPG features, yet their fusion strategies often disrupt pretrained facial representations and lack explicit mechanisms to suppress subject-specific variations. To address these issues, we propose a subject-invariant cross-modal prompt-tuning framework for video-based emotion recognition. Specifically, rPPG waveforms are transformed into noise-robust time-frequency representations (TFRs), from which modality-complementary prompts are generated to modulate facial tokens within a frozen Vision Transformer (ViT). This design enables effective cross-modal interaction while preserving the generalizable facial representations learned by the pretrained backbone. In addition, we introduce a decoupled shared-specific adapter (DSSA) into each ViT layer to explicitly separate subject-shared and subject-specific components, thereby improving cross-subject generalization. Experiments on the MAHNOB-HCI and DEAP benchmarks demonstrate that the proposed method consistently outperforms strong baselines in both recognition accuracy and generalization ability, highlighting its effectiveness for video-based emotion recognition.

cs.CV

Degradation-Robust Fusion: An Efficient Degradation-Aware Diffusion Framework for Multimodal Image Fusion in Arbitrary Degradation Scenarios

Complex degradations like noise, blur, and low resolution are typical challenges in real world image fusion tasks, limiting the performance and practicality of existing methods. End to end neural network based approaches are generally simple to design and highly efficient in inference, but their black-box nature leads to limited interpretability. Diffusion based methods alleviate this to some extent by providing powerful generative priors and a more structured inference process. However, they are trained to learn a single domain target distribution, whereas fusion lacks natural fused data and relies on modeling complementary information from multiple sources, making diffusion hard to apply directly in practice. To address these challenges, this paper proposes an efficient degradation aware diffusion framework for image fusion under arbitrary degradation scenarios. Specifically, instead of explicitly predicting noise as in conventional diffusion models, our method performs implicit denoising by directly regressing the fused image, enabling flexible adaptation to diverse fusion tasks under complex degradations with limited steps. Moreover, we design a joint observation model correction mechanism that simultaneously imposes degradation and fusion constraints during sampling to ensure high reconstruction accuracy. Experiments on diverse fusion tasks and degradation configurations demonstrate the superiority of the proposed method under complex degradation scenarios.

cs.CV

Customized Fusion: A Closed-Loop Dynamic Network for Adaptive Multi-Task-Aware Infrared-Visible Image Fusion

Infrared-visible image fusion aims to integrate complementary information for robust visual understanding, but existing fusion methods struggle with simultaneously adapting to multiple downstream tasks. To address this issue, we propose a Closed-Loop Dynamic Network (CLDyN) that can adaptively respond to the semantic requirements of diverse downstream tasks for task-customized image fusion. Specifically, CLDyN introduces a closed-loop optimization mechanism that establishes a semantic transmission chain to achieve explicit feedback from downstream tasks to the fusion network through a Requirement-driven Semantic Compensation (RSC) module. The RSC module leverages a Basis Vector Bank (BVB) and an Architecture-Adaptive Semantic Injection (A2SI) block to customize the network architecture according to task requirements, thereby enabling task-specific semantic compensation and allowing the fusion network to actively adapt to diverse tasks without retraining. To promote semantic compensation, a reward-penalty strategy is introduced to reward or penalize the RSC module based on task performance variations. Experiments on the M3FD, FMB, and VT5000 datasets demonstrate that CLDyN not only maintains high fusion quality but also exhibits strong multi-task adaptability. The code is available at https://github.com/YR0211/CLDyN.

cs.CV

A high-order, conservative and positivity-preserving intersection-based remapping method between meshes with isoparametric curvilinear cells

This paper presents a novel intersection-based remapping method for isoparametric curvilinear meshes within the indirect arbitrary Lagrangian-Eulerian (ALE) framework, addressing the challenges of transferring physical quantities between high-order curved-edge meshes. Our method leverages the Weiler-Atherton clipping algorithm to compute intersections between curved-edge quadrangles, enabling robust handling of arbitrary order isoparametric curves. By integrating multi-resolution weighted essentially non-oscillatory (WENO) reconstruction, we achieve high-order accuracy while suppressing numerical oscillations near discontinuities. A positivity-preserving limiter is further applied to ensure physical quantities such as density remain non-negative without compromising conservation or accuracy. Notably, the computational cost of handling higher-order curved meshes, such as cubic or even higher-degree parametric curves, does not significantly increase compared to secondorder curved meshes. This ensures that our method remains efficient and scalable, making it applicable to arbitrary high-order isoparametric curvilinear cells without compromising performance. Numerical experiments demonstrate that the proposed method achieves highorder accuracy, strict conservation (with errors approaching machine precision), essential non-oscillation and positivity-preserving.

math.NA

An asymptotic-preserving IMEX PN method for the gray model of the radiative transfer equation

An asymptotic-preserving (AP) implicit-explicit PN numerical scheme is proposed for the gray model of the radiative transfer equation, where the first- and second-order numerical schemes are discussed for both the linear and nonlinear models. The AP property of this numerical scheme is proved theoretically and numerically, while the numerical stability of the linear model is verified by Fourier analysis. Several classical benchmark examples are studied to validate the efficiency of this numerical scheme.

math.NA

Super-resolution enabled widefield quantum diamond microscopy

Widefield quantum diamond microscopy (WQDM) based on Kohler-illumination has been widely adopted in the field of quantum sensing, however, practical applications are still limited by issues such as unavoidable photodamage and unsatisfied spatial-resolution. Here, we design and develop a super-resolution enabled WQDM using a digital micromirror device (DMD)-based structured illumination microscopy. With the rapidly programmable illumination patterns, we have firstly demonstrated how to mitigate phototoxicity when imaging nanodiamonds in cell samples. As a showcase, we have performed the super-resolved quantum sensing measurements of two individual nanodiamonds not even distinguishable with conventional WQDM. The DMD-powered WQDM presents not only excellent compatibility with quantum sensing solutions, but also strong advantages in high imaging speed, high resolution, low phototoxicity, and enhanced signal-to-background ratio, making it a competent tool to for applications in demanding fields such as biomedical science.

physics.optics

FaFCNN: A General Disease Classification Framework Based on Feature Fusion Neural Networks

There are two fundamental problems in applying deep learning/machine learning methods to disease classification tasks, one is the insufficient number and poor quality of training samples; another one is how to effectively fuse multiple source features and thus train robust classification models. To address these problems, inspired by the process of human learning knowledge, we propose the Feature-aware Fusion Correlation Neural Network (FaFCNN), which introduces a feature-aware interaction module and a feature alignment module based on domain adversarial learning. This is a general framework for disease classification, and FaFCNN improves the way existing methods obtain sample correlation features. The experimental results show that training using augmented features obtained by pre-training gradient boosting decision tree yields more performance gains than random-forest based methods. On the low-quality dataset with a large amount of missing data in our setup, FaFCNN obtains a consistently optimal performance compared to competitive baselines. In addition, extensive experiments demonstrate the robustness of the proposed method and the effectiveness of each component of the model\footnote{Accepted in IEEE SMC2023}.

cs.LG

Landslide Surface Displacement Prediction Based on VSXC-LSTM Algorithm

Landslide is a natural disaster that can easily threaten local ecology, people's lives and property. In this paper, we conduct modelling research on real unidirectional surface displacement data of recent landslides in the research area and propose a time series prediction framework named VMD-SegSigmoid-XGBoost-ClusterLSTM (VSXC-LSTM) based on variational mode decomposition, which can predict the landslide surface displacement more accurately. The model performs well on the test set. Except for the random item subsequence that is hard to fit, the root mean square error (RMSE) and the mean absolute percentage error (MAPE) of the trend item subsequence and the periodic item subsequence are both less than 0.1, and the RMSE is as low as 0.006 for the periodic item prediction module based on XGBoost\footnote{Accepted in ICANN2023}.

cs.LG

PulseGAN: Learning to generate realistic pulse waveforms in remote photoplethysmography

Remote photoplethysmography (rPPG) is a non-contact technique for measuring cardiac signals from facial videos. High-quality rPPG pulse signals are urgently demanded in many fields, such as health monitoring and emotion recognition. However, most of the existing rPPG methods can only be used to get average heart rate (HR) values due to the limitation of inaccurate pulse signals. In this paper, a new framework based on generative adversarial network, called PulseGAN, is introduced to generate realistic rPPG pulse signals through denoising the chrominance signals. Considering that the cardiac signal is quasi-periodic and has apparent time-frequency characteristics, the error losses defined in time and spectrum domains are both employed with the adversarial loss to enforce the model generating accurate pulse waveforms as its reference. The proposed framework is tested on the public UBFC-RPPG database in both within-database and cross-database configurations. The results show that the PulseGAN framework can effectively improve the waveform quality, thereby enhancing the accuracy of HR, the heart rate variability (HRV) and the interbeat interval (IBI). The proposed method achieves the best performance compared to the denoising autoencoder (DAE) and CHROM, with the mean absolute error of AVNN (the average of all normal-to-normal intervals) improving 20.85% and 41.19%, and the mean absolute error of SDNN (the standard deviation of all NN intervals) improving 20.28% and 37.53%, respectively, in the cross-database test. This framework can be easily extended to other existing deep learning based rPPG methods, which is expected to expand the application scope of rPPG techniques.

eess.IV

An adaptive moving mesh discontinuous Galerkin method for the radiative transfer equation

The radiative transfer equation models the interaction of radiation with scattering and absorbing media and has important applications in various fields in science and engineering. It is an integro-differential equation involving time, space and angular variables and contains an integral term in angular directions while being hyperbolic in space. The challenges for its numerical solution include the needs to handle with its high dimensionality, the presence of the integral term, and the development of discontinuities and sharp layers in its solution along spatial directions. Its numerical solution is studied in this paper using an adaptive moving mesh discontinuous Galerkin method for spatial discretization together with the discrete ordinate method for angular discretization. The former employs a dynamic mesh adaptation strategy based on moving mesh partial differential equations to improve computational accuracy and efficiency. Its mesh adaptation ability, accuracy, and efficiency are demonstrated in a selection of one- and two-dimensional numerical examples.

math.NA

Bipedal nanowalker by pure physical mechanisms

Artificial nanowalkers are inspired by biomolecular counterparts from living cells, but remain far from comparable to the latter in design principles. The walkers reported to date mostly rely on chemical mechanisms to gain a direction; they all produce chemical wastes. Here we report a light-powered DNA bipedal walker based on a design principle derived from cellular walkers. The walker has two identical feet and the track has equal binding sites; yet the walker gains a direction by pure physical mechanisms that autonomously amplify an intra-site asymmetry into a ratchet effect. The nanowalker is free of any chemical waste. It has a distinct thermodynamic feature that it possesses the same equilibrium before and after operation, but generates a truly non-equilibrium distribution during operation. The demonstrated design principle exploits mechanical effects and is adaptable for use in other nanomachines.

physics.bio-ph