SearcharxivSearch

arXiv subjects

Aiping Liu

Publications and source records attributed to Aiping Liu.

16 recordsLinked to original sources

EmoFeedback$^2$: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback

Continuous emotional image content generation (C-EICG) is emerging rapidly due to its ability to produce images aligned with both user descriptions and continuous emotional values. However, existing approaches lack emotional feedback from generated images, limiting the control of emotional continuity. Additionally, their simple emotion-text alignment fails to adaptively adjust emotional prompts according to image content, leading to insufficient emotional fidelity. To address these concerns, we propose a novel generation-understanding-feedback reinforcement paradigm (EmoFeedback$^2$) for C-EICG, which exploits the reasoning capability of the fine-tuned large vision-language model (LVLM) to provide reward and textual feedback for generating high-quality images with continuous emotions. Specifically, we introduce an emotion-aware reward feedback strategy, where the LVLM evaluates the emotional values of generated images and computes the reward against target emotions, guiding the reinforcement fine-tuning of the generative model and enhancing the emotional continuity of images. Furthermore, we design a self-promotion textual feedback framework, in which the LVLM iteratively analyzes the emotional content of generated images and adaptively produces refinement suggestions for the next-round prompt, improving the emotional fidelity with fine-grained content. Extensive experimental results demonstrate that our approach effectively generates high-quality images with the desired emotions, outperforming existing state-of-the-art methods on both our custom dataset and public dataset.

cs.CV

Dual Graph Attention based Disentanglement Multiple Instance Learning for Brain Age Estimation

Deep learning techniques have demonstrated great potential for accurately estimating brain age by analyzing Magnetic Resonance Imaging (MRI) data from healthy individuals. However, current methods for brain age estimation often directly utilize whole input images, overlooking two important considerations: 1) the heterogeneous nature of brain aging, where different brain regions may degenerate at different rates, and 2) the existence of age-independent redundancies in brain structure. To overcome these limitations, we propose a Dual Graph Attention based Disentanglement Multi-instance Learning (DGA-DMIL) framework for improving brain age estimation. Specifically, the 3D MRI data, treated as a bag of instances, is fed into a 2D convolutional neural network backbone, to capture the unique aging patterns in MRI. A dual graph attention aggregator is then proposed to learn the backbone features by exploiting the intra- and inter-instance relationships. Furthermore, a disentanglement branch is introduced to separate age-related features from age-independent structural representations to ameliorate the interference of redundant information on age prediction. To verify the effectiveness of the proposed framework, we evaluate it on two datasets, UK Biobank and ADNI, containing a total of 35,388 healthy individuals. Our proposed model demonstrates exceptional accuracy in estimating brain age, achieving a remarkable mean absolute error of 2.12 years in the UK Biobank. The results establish our approach as state-of-the-art compared to other competing brain age estimation models. In addition, the instance contribution scores identify the varied importance of brain areas for aging prediction, which provides deeper insights into the understanding of brain aging.

cs.CV

Data Augmentation for Seizure Prediction with Generative Diffusion Model

Data augmentation (DA) can significantly strengthen the electroencephalogram (EEG)-based seizure prediction methods. However, existing DA approaches are just the linear transformations of original data and cannot explore the feature space to increase diversity effectively. Therefore, we propose a novel diffusion-based DA method called DiffEEG. DiffEEG can fully explore data distribution and generate samples with high diversity, offering extra information to classifiers. It involves two processes: the diffusion process and the denoised process. In the diffusion process, the model incrementally adds noise with different scales to EEG input and converts it into random noise. In this way, the representation of data can be learned. In the denoised process, the model utilizes learned knowledge to sample synthetic data from random noise input by gradually removing noise. The randomness of input noise and the precise representation enable the synthetic samples to possess diversity while ensuring the consistency of feature space. We compared DiffEEG with original, down-sampling, sliding windows and recombination methods, and integrated them into five representative classifiers. The experiments demonstrate the effectiveness and generality of our method. With the contribution of DiffEEG, the Multi-scale CNN achieves state-of-the-art performance, with an average sensitivity, FPR, AUC of 95.4%, 0.051/h, 0.932 on the CHB-MIT database and 93.6%, 0.121/h, 0.822 on the Kaggle database.

eess.SP

Proposal of a free-space-to-chip pipeline for transporting single atoms

A free-space-to-chip pipeline is proposed to efficiently transport single atoms from a magneto-optical trap to an on-chip evanescent field trap. Due to the reflection of the dipole laser on the chip surface, the conventional conveyor belt approach can only transport atoms close to the chip surface but with a distance of about one wavelength, which prevents efficient interaction between the atom and the on-chip waveguide devices. Here, based on a two-layer photonic chip architecture, a diffraction beam of the integrated grating with an incident angle of the Brewster angle is utilized to realize free-space-to-chip atom pipeline. Numerical simulation verified that the reflection of the dipole laser is suppressed and that the atoms can be brought to the chip surface with a distance of only 100nm. Therefore, the pipeline allows a smooth transport of atoms from free space to the evanescent field trap of waveguides and promises a reliable atom source for a hybrid photonic-atom chip.

physics.optics

PanFlowNet: A Flow-Based Deep Network for Pan-sharpening

Pan-sharpening aims to generate a high-resolution multispectral (HRMS) image by integrating the spectral information of a low-resolution multispectral (LRMS) image with the texture details of a high-resolution panchromatic (PAN) image. It essentially inherits the ill-posed nature of the super-resolution (SR) task that diverse HRMS images can degrade into an LRMS image. However, existing deep learning-based methods recover only one HRMS image from the LRMS image and PAN image using a deterministic mapping, thus ignoring the diversity of the HRMS image. In this paper, to alleviate this ill-posed issue, we propose a flow-based pan-sharpening network (PanFlowNet) to directly learn the conditional distribution of HRMS image given LRMS image and PAN image instead of learning a deterministic mapping. Specifically, we first transform this unknown conditional distribution into a given Gaussian distribution by an invertible network, and the conditional distribution can thus be explicitly defined. Then, we design an invertible Conditional Affine Coupling Block (CACB) and further build the architecture of PanFlowNet by stacking a series of CACBs. Finally, the PanFlowNet is trained by maximizing the log-likelihood of the conditional distribution given a training set and can then be used to predict diverse HRMS images. The experimental results verify that the proposed PanFlowNet can generate various HRMS images given an LRMS image and a PAN image. Additionally, the experimental results on different kinds of satellite datasets also demonstrate the superiority of our PanFlowNet compared with other state-of-the-art methods both visually and quantitatively.

cs.CV

Transporting cold atoms towards a GaN-on-sapphire chip via an optical conveyor belt

Trapped atoms on photonic structures inspire many novel quantum devices for quantum information processing and quantum sensing. Here, we have demonstrated a hybrid photonic-atom chip platform based on a GaN-on-sapphire chip and the transport of an ensemble of atoms from free space towards the chip with an optical conveyor belt. The maximum transport efficiency of atoms is about 50% with a transport distance of 500 $\mathrm{μm}$. Our results open up a new route toward the efficiently loading of cold atoms into the evanescent-field trap formed by the photonic integrated circuits, which promises strong and controllable interactions between single atoms and single photons.

physics.atom-ph

Model-Guided Multi-Contrast Deep Unfolding Network for MRI Super-resolution Reconstruction

Magnetic resonance imaging (MRI) with high resolution (HR) provides more detailed information for accurate diagnosis and quantitative image analysis. Despite the significant advances, most existing super-resolution (SR) reconstruction network for medical images has two flaws: 1) All of them are designed in a black-box principle, thus lacking sufficient interpretability and further limiting their practical applications. Interpretable neural network models are of significant interest since they enhance the trustworthiness required in clinical practice when dealing with medical images. 2) most existing SR reconstruction approaches only use a single contrast or use a simple multi-contrast fusion mechanism, neglecting the complex relationships between different contrasts that are critical for SR improvement. To deal with these issues, in this paper, a novel Model-Guided interpretable Deep Unfolding Network (MGDUN) for medical image SR reconstruction is proposed. The Model-Guided image SR reconstruction approach solves manually designed objective functions to reconstruct HR MRI. We show how to unfold an iterative MGDUN algorithm into a novel model-guided deep unfolding network by taking the MRI observation matrix and explicit multi-contrast relationship matrix into account during the end-to-end optimization. Extensive experiments on the multi-contrast IXI dataset and BraTs 2019 dataset demonstrate the superiority of our proposed model.

eess.IV

Proposal for stable atom trapping on a GaN-on-Sapphire chip

The hybrid photon-atom integrated circuits, which include photonic microcavities and trapped single neutral atom in their evanescent field, are of great potential for quantum information processing. In this platform, the atoms provide the single-photon nonlinearity and long-lived memory, which are complementary to the excellent passive photonics devices in conventional quantum photonic circuits. In this work, we propose a stable platform for realizing the hybrid photon-atom circuits based on an unsuspended photonic chip. By introducing high-order modes in the microring, a feasible evanescent-field trap potential well $\sim0.3\,\mathrm{mK}$ could be obtained by only $10\,\mathrm{mW}$-level power in the cavity, compared with $100\,\mathrm{mW}$-level power required in the scheme based on fundamental modes. Based on our scheme, stable single atom trapping with relatively low laser power is feasible for future studies on high-fidelity quantum gates, single-photon sources, as well as many-body quantum physics based on a controllable atom array in a microcavity.

physics.optics

Multi-grating design for integrated single-atom trapping, manipulation, and readout

An on-chip multi-grating device is proposed to interface single-atoms and integrated photonic circuits, by guiding and focusing lasers to the area with ~10um above the chip for trapping, state manipulation, and readout of single Rubidium atoms. For the optical dipole trap, two 850 nm laser beams are diffracted and overlapped to form a lattice of single-atom dipole trap, with the diameter of optical dipole trap around 2.7um. Similar gratings are designed for guiding 780 nm probe laser to excite and also collect the fluorescence of 87Rb atoms. Such a device provides a compact solution for future applications of single atoms, including the single photon source, single-atom quantum register, and sensor.

physics.optics

Toward Open-World Electroencephalogram Decoding Via Deep Learning: A Comprehensive Survey

Electroencephalogram (EEG) decoding aims to identify the perceptual, semantic, and cognitive content of neural processing based on non-invasively measured brain activity. Traditional EEG decoding methods have achieved moderate success when applied to data acquired in static, well-controlled lab environments. However, an open-world environment is a more realistic setting, where situations affecting EEG recordings can emerge unexpectedly, significantly weakening the robustness of existing methods. In recent years, deep learning (DL) has emerged as a potential solution for such problems due to its superior capacity in feature extraction. It overcomes the limitations of defining `handcrafted' features or features extracted using shallow architectures, but typically requires large amounts of costly, expertly-labelled data - something not always obtainable. Combining DL with domain-specific knowledge may allow for development of robust approaches to decode brain activity even with small-sample data. Although various DL methods have been proposed to tackle some of the challenges in EEG decoding, a systematic tutorial overview, particularly for open-world applications, is currently lacking. This article therefore provides a comprehensive survey of DL methods for open-world EEG decoding, and identifies promising research directions to inspire future studies for EEG decoding in real-world applications.

eess.SP

Automated assessment of disease severity of COVID-19 using artificial intelligence with synthetic chest CT

Background: Triage of patients is important to control the pandemic of coronavirus disease 2019 (COVID-19), especially during the peak of the pandemic when clinical resources become extremely limited. Purpose: To develop a method that automatically segments and quantifies lung and pneumonia lesions with synthetic chest CT and assess disease severity in COVID-19 patients. Materials and Methods: In this study, we incorporated data augmentation to generate synthetic chest CT images using public available datasets (285 datasets from "Lung Nodule Analysis 2016"). The synthetic images and masks were used to train a 2D U-net neural network and tested on 203 COVID-19 datasets to generate lung and lesion segmentations. Disease severity scores (DL: damage load; DS: damage score) were calculated based on the segmentations. Correlations between DL/DS and clinical lab tests were evaluated using Pearson's method. A p-value < 0.05 was considered as statistical significant. Results: Automatic lung and lesion segmentations were compared with manual annotations. For lung segmentation, the median values of dice similarity coefficient, Jaccard index and average surface distance, were 98.56%, 97.15% and 0.49 mm, respectively. The same metrics for lesion segmentation were 76.95%, 62.54% and 2.36 mm, respectively. Significant (p << 0.05) correlations were found between DL/DS and percentage lymphocytes tests, with r-values of -0.561 and -0.501, respectively. Conclusion: An AI system that based on thoracic radiographic and data augmentation was proposed to segment lung and lesions in COVID-19 patients. Correlations between imaging findings and clinical lab tests suggested the value of this system as a potential tool to assess disease severity of COVID-19.

eess.IV

Unfolding Taylor's Approximations for Image Restoration

Deep learning provides a new avenue for image restoration, which demands a delicate balance between fine-grained details and high-level contextualized information during recovering the latent clear image. In practice, however, existing methods empirically construct encapsulated end-to-end mapping networks without deepening into the rationality, and neglect the intrinsic prior knowledge of restoration task. To solve the above problems, inspired by Taylor's Approximations, we unfold Taylor's Formula to construct a novel framework for image restoration. We find the main part and the derivative part of Taylor's Approximations take the same effect as the two competing goals of high-level contextualized information and spatial details of image restoration respectively. Specifically, our framework consists of two steps, correspondingly responsible for the mapping and derivative functions. The former first learns the high-level contextualized information and the later combines it with the degraded input to progressively recover local high-order spatial details. Our proposed framework is orthogonal to existing methods and thus can be easily integrated with them for further improvement, and extensive experiments demonstrate the effectiveness and scalability of our proposed framework.

cs.CV

MLBF-Net: A Multi-Lead-Branch Fusion Network for Multi-Class Arrhythmia Classification Using 12-Lead ECG

Automatic arrhythmia detection using 12-lead electrocardiogram (ECG) signal plays a critical role in early prevention and diagnosis of cardiovascular diseases. In the previous studies on automatic arrhythmia detection, most methods concatenated 12 leads of ECG into a matrix, and then input the matrix to a variety of feature extractors or deep neural networks for extracting useful information. Under such frameworks, these methods had the ability to extract comprehensive features (known as integrity) of 12-lead ECG since the information of each lead interacts with each other during training. However, the diverse lead-specific features (known as diversity) among 12 leads were neglected, causing inadequate information learning for 12-lead ECG. To maximize the information learning of multi-lead ECG, the information fusion of comprehensive features with integrity and lead-specific features with diversity should be taken into account. In this paper, we propose a novel Multi-Lead-Branch Fusion Network (MLBF-Net) architecture for arrhythmia classification by integrating multi-loss optimization to jointly learning diversity and integrity of multi-lead ECG. MLBF-Net is composed of three components: 1) multiple lead-specific branches for learning the diversity of multi-lead ECG; 2) cross-lead features fusion by concatenating the output feature maps of all branches for learning the integrity of multi-lead ECG; 3) multi-loss co-optimization for all the individual branches and the concatenated network. We demonstrate our MLBF-Net on China Physiological Signal Challenge 2018 which is an open 12-lead ECG dataset. The experimental results show that MLBF-Net obtains an average $F_1$ score of 0.855, reaching the highest arrhythmia classification performance. The proposed method provides a promising solution for multi-lead ECG analysis from an information fusion perspective.

eess.SP

Reconfigurable vortex beam generator based on the Fourier transformation principle

A method to generate the optical vortex beam with arbitrary superposition of different orders of orbital angular momentum (OAM) on a photonic chip is proposed. The distributed Fourier holographic gratings are proposed to convert the propagating wave in waveguides to a vortex beam in the free space, and the components of different OAMs can be controlled by the amplitude and phases of on-chip incident light based on the principle of Fourier transformation. As an example, we studied a typical device composed of nine Fourier holographic gratings on fan-shaped waveguides. By scalar diffraction calculation, the OAM of the optical beam from the reconfigurable vortex beam generator can be controlled on-demand from -2nd to 2nd by adjusting the phase of input light fields, which is demonstrated numerically with the fidelity of generated optical vortex beam above 0.69. The working bandwidth of the Fourier holographic grating is about 80 nm with a fidelity above 0.6. Our work provides an feasible method to manipulate the vortex beam or detect arbitrary superposition of OAMs, which can be used in integrated photonics structures for optical trapping, signal processing, and imaging.

physics.optics

On-chip generation and control of the vortex beam

A new method to generate and control the amplitude and phase distributions of a optical vortex beam is proposed. By introducing a holographic grating on top of the dielectric waveguide, the free space vortex beam and the in-plane guiding wave can be converted to each other. This microscale holographic grating is very robust against the variation of geometry parameters. The designed vortex beam generator can produce the target beam with a fidelity up to 0.93, and the working bandwidth is about 175 nm with the fidelity larger than 0.80. In addition, a multiple generator composed of two holographic gratings on two parallel waveguides are studied, which can perform an effective and flexible modulation on the vortex beam by controlling the phase of the input light. Our work opens a new avenue towards the integrated OAM devices with multiple degrees of optical freedom, which can be used for optical tweezers, micronano imaging, information processing, and so on.

physics.optics

Interference of surface plasmon polaritons from a "point" source

The interference patterns of the surface plasmon polaritons(SPPs) on the metal surface from a "point" source are observed. These interference patterns come from the forward SPPs and the reflected one from the obstacles, such as straightedge,corner, and ring groove structure. Innovation to the previous works, a "point" SPPs source with diameter of 100 nm is generated at the freely chosen positions on Au/air interface using near field excitation method. Such a "point" source provides good enough coherence to generate obvious interference phenomenon. The constructive and destructive interference patterns of the SPPs agree well with the numerical caculation. This "point" SPPs source may be useful in the investigation of plasmonics for its high coherence, deterministic position and minimum requirement for the initial light source.

physics.optics