SearcharxivSearch

arXiv subjects

Xingye Chen

Publications and source records attributed to Xingye Chen.

6 recordsLinked to original sources

KD-CVG: A Knowledge-Driven Approach for Creative Video Generation

Creative Generation (CG) leverages generative models to automatically produce advertising content that highlights product features, and it has been a significant focus of recent research. However, while CG has advanced considerably, most efforts have concentrated on generating advertising text and images, leaving Creative Video Generation (CVG) relatively underexplored. This gap is largely due to two major challenges faced by Text-to-Video (T2V) models: (a) \textbf{ambiguous semantic alignment}, where models struggle to accurately correlate product selling points with creative video content, and (b) \textbf{inadequate motion adaptability}, resulting in unrealistic movements and distortions. To address these challenges, we develop a comprehensive Advertising Creative Knowledge Base (ACKB) as a foundational resource and propose a knowledge-driven approach (KD-CVG) to overcome the knowledge limitations of existing models. KD-CVG consists of two primary modules: Semantic-Aware Retrieval (SAR) and Multimodal Knowledge Reference (MKR). SAR utilizes the semantic awareness of graph attention networks and reinforcement learning feedback to enhance the model's comprehension of the connections between selling points and creative videos. Building on this, MKR incorporates semantic and motion priors into the T2V model to address existing knowledge gaps. Extensive experiments have demonstrated KD-CVG's superior performance in achieving semantic alignment and motion adaptability, validating its effectiveness over other state-of-the-art methods. The code and dataset will be open source at https://kdcvg.github.io/KDCVG/.

cs.CV

Quantum Sensing MRI for Noninvasive Detection of Neuronal Electrical Activity in Human Brains

Neuronal electrical activity underlies human cognition including perception, attention, memory, language, and decision-making. Yet its direct, noninvasive measurement in the living human brain remains a fundamental challenge. Existing neuroimaging techniques, including electroencephalography (EEG), magnetoencephalography (MEG), and functional magnetic resonance imaging (fMRI), are limited by trade-offs in sensitivity and spatial or temporal resolution. Here we propose quantum sensing MRI (qsMRI), a noninvasive approach that enables direct detection of neuronal firing-induced magnetic fields using a clinical MRI system. qsMRI exploits endogenous proton (1H) nuclear spins in water molecules as intrinsic quantum sensors and decodes time-resolved phase information from the free induction decay signals to infer neuronal magnetic fields. We validate qsMRI through simulations, phantom experiments, and human studies at rest and during motor tasks, and provide open experimental procedures to facilitate independent rigorous validation. We further present a case study demonstrating potential applications to neurological disorders. qsMRI represents, to our knowledge, the first-in-human application of quantum sensing on a clinical MRI platform and may lay the foundation for a non-BOLD functional imaging modality capable of probing neuronal firing dynamics in both cortical and deep brain regions.

physics.med-ph

CTR-Driven Advertising Image Generation with Multimodal Large Language Models

In web data, advertising images are crucial for capturing user attention and improving advertising effectiveness. Most existing methods generate background for products primarily focus on the aesthetic quality, which may fail to achieve satisfactory online performance. To address this limitation, we explore the use of Multimodal Large Language Models (MLLMs) for generating advertising images by optimizing for Click-Through Rate (CTR) as the primary objective. Firstly, we build targeted pre-training tasks, and leverage a large-scale e-commerce multimodal dataset to equip MLLMs with initial capabilities for advertising image generation tasks. To further improve the CTR of generated images, we propose a novel reward model to fine-tune pre-trained MLLMs through Reinforcement Learning (RL), which can jointly utilize multimodal features and accurately reflect user click preferences. Meanwhile, a product-centric preference optimization strategy is developed to ensure that the generated background content aligns with the product characteristics after fine-tuning, enhancing the overall relevance and effectiveness of the advertising images. Extensive experiments have demonstrated that our method achieves state-of-the-art performance in both online and offline metrics. Our code and pre-trained models are publicly available at: https://github.com/Chenguoz/CAIG.

cs.LG

Multi-TE Single-Quantum Sodium (23Na) MRI: A Clinically Translatable Technique for Separation of Mono- and Bi-T2 Sodium Signals

Sodium magnetic resonance imaging (MRI) is sensitive and specific to ionic balance of cells owing to 10 fold difference in sodium concentration across membrane actively maintained by sodium potassium (Na+ K+) pump. Disruption of the pump and membrane integrity, as seen in neurological disorders such as epilepsy, multiple sclerosis, bipolar disease, and mild traumatic brain injury, leads to a large increase in intracellular sodium. Such a cellular level alteration is however overshadowed by large signal from extracellular sodium, leaving behind a long standing pursuit to separate signals from sodium exhibiting mono vs biexponential transverse (T2) decay under the inherent constraint of low signal to noise ratio even at advanced clinical field of 3 Tesla. Here we propose a novel technique that exploits intrinsic difference in their T2 decays by simply acquiring single quantum images at multiple echo times (TEs) and performing accurate matrix inversion at voxel. This approach was then investigated using numerical models, agar phantoms and human subjects, showing high accuracy of the separation in phantoms (95.8 percent for monoT2 and 72.5 to 80.4 percent for biT2) and clinical feasibility in humans. Thus, sodium MRI at 3T can now facilitate detection of neurological disorders early at cellular level and response to treatment as well. Keywords. sodium MRI, single quantum MRI, triple quantum MRI, neuroimaging, neurodegeneration

physics.med-ph

PointSCNet: Point Cloud Structure and Correlation Learning Based on Space Filling Curve-Guided Sampling

Geometrical structures and the internal local region relationship, such as symmetry, regular array, junction, etc., are essential for understanding a 3D shape. This paper proposes a point cloud feature extraction network named PointSCNet, to capture the geometrical structure information and local region correlation information of a point cloud. The PointSCNet consists of three main modules: the space-filling curve-guided sampling module, the information fusion module, and the channel-spatial attention module. The space-filling curve-guided sampling module uses Z-order curve coding to sample points that contain geometrical correlation. The information fusion module uses a correlation tensor and a set of skip connections to fuse the structure and correlation information. The channel-spatial attention module enhances the representation of key points and crucial feature channels to refine the network. The proposed PointSCNet is evaluated on shape classification and part segmentation tasks. The experimental results demonstrate that the PointSCNet outperforms or is on par with state-of-the-art methods by learning the structure and correlation of point clouds effectively.

cs.CV

Super-resolution Imaging of the Fluorescent Dipole Assembly with Polarized Structured Illumination Microscopy

Fluorescence polarization microscopy images both the intensity and orientation of fluorescent dipoles, which plays a vital role in studying the molecular structure and dynamics of bio-complex. However, it is difficult to resolve the dipole assemblies on the subcellular structure and their dynamics in living cells with super-resolution. Here we report polarized structured illumination microscopy (pSIM), which decouples the entangled spatial and angular structured illumination through interpreting the dipoles in spatio-angular hyperspace. We demonstrate its application on a series of biological filamentous systems such as cytoskeleton networks and lambda-DNA, and report the dynamics of short actin sliding through myosin-coated surface. Further, pSIM reveals "side-by-side" organization of the actin ring structure in the membrane-associated periodic skeleton in hippocampal neurons. It also images the dipole dynamics of green fluorescent proteins labeled to the microtubules in live U2OS cells. pSIM can be applied directly to a large variety of commercial or home-built SIM systems.

physics.optics