SearcharxivSearch

arXiv subjects

Zhihao Qian

Publications and source records attributed to Zhihao Qian.

8 recordsLinked to original sources

A Novel Hierarchy of Quantum Kernel Networks on Smoothed Particle Hydrodynamics

This study proposed the hierarchy of quantum kernel networks by combing multi quantum networks with smoothed particle hydrodynamics (SPH). The Lagrangian quantum network model was further developed based on an improved quantum multilayer perceptron (QMLP). A sequential hybrid quantum-classical framework was constructed to ensure robust particle gradient-based optimization and mitigate barren plateaus for computational particle dynamics. This approach combines smoothing kernels with quantum learning, establishing a novel quantum intelligent particle paradigm. The framework was validated through some benchmarks on multifarious quantum neural networks, static multi-level vortex reconstructions and transient scalar advective transports. Numerical results show that while elementary quantum circuits struggle with generalization in unstructured domains, the hybrid crossed-QMLP matches the fitting accuracy of classical SPH in quantum optimized space. Despite current limitations in computational efficiency and hardware implementation, this work paves the way for a new investigation on quantum-particle approach by mapping unstructured Lagrangian particle topologies into integrated quantum networks.

quant-ph

Multi-Partitioned Computing Quantum-Particle Approach: A Hybrid Quantum Framework for Fluid Flow

This study established a quantum-classical hybrid framework that integrates quantum computing paradigm with meshfree finite particle method. By harnessing quantum superposition and entanglement, it hybridized the critical computational kernels (termed as quantum finite particle method). A resource-efficient quantum computational strategy on multi-partitioned zones was proposed, which leverages a fixed small-scale quantum circuit as a fundamental processing unit to handle inner product for arbitrarily sized arrays. This approach employs iterative nesting of the quantum-core operation to accommodate varying input dimensions while maintaining hardware feasibility throughout. Motivated with developed quantum framework, the novel numerical discretization for hybrid quantum computational particle dynamics can be derived commonly and applied in fluid flows. Through a sequence of numerical experiments purposefully, the proposed numerical model was thoroughly validated and analyzed. Results demonstrate that integrating quantum computing to hybridize conventional linear combinations of particle dynamics serves as a novel computing paradigm. By further extending into the numerical investigation of viscoelastic, highly elastic, and purely elastic fluids under high Weissenberg number conditions, the applicability of simulation framework is broadened. Despite bottlenecks in quantum hardware and computational efficiency on this process, these advances offer critical insights for transitioning quantum-enhanced fluid simulation to practical engineering applications.

physics.flu-dyn

Water evaporation-driven dynamic diode for direct electricity generation

Harnessing energy from ubiquitous water resources via molecular-scale mechanisms remains a critical frontier in sustainable energy research. Herein, we present a novel evaporation-driven power generator based on a dynamic diode architecture that continuously harvests direct current (DC) electricity by leveraging the flipping of the strong built-in electric field (up to 10E10 V/cm) generated by polar molecules such as water to drive directional carrier migration. In our system, water molecules undergo sequential polarization and depolarization at the graphene-water-silicon interface, triggering cycles of charge trapping and release. This nonionic mechanism is driven primarily by the Fermi level difference between graphene and silicon, augmented by the intrinsic dipole moment of water molecules. Structural optimization using graphene enhances evaporation kinetics and interfacial contact, yielding an open-circuit voltage of 0.35 V from a 2 cm * 1 cm device. When four units are connected in series, the system delivers a stable 1.2V output. Unlike ion-mediated energy harvesters, this corrosion-free architecture ensures long-term stability and material compatibility. Our work introduces a fundamentally new approach to water-based power generation, establishing interfacial polarization engineering as a scalable strategy for low-cost, sustainable electricity production from ambient water.

physics.atom-ph

D$^3$: Scaling Up Deepfake Detection by Learning from Discrepancy

The boom of Generative AI brings opportunities entangled with risks and concerns. Existing literature emphasizes the generalization capability of deepfake detection on unseen generators, significantly promoting the detector's ability to identify more universal artifacts. This work seeks a step toward a universal deepfake detection system with better generalization and robustness. We do so by first scaling up the existing detection task setup from the one-generator to multiple-generators in training, during which we disclose two challenges presented in prior methodological designs and demonstrate the divergence of detectors' performance. Specifically, we reveal that the current methods tailored for training on one specific generator either struggle to learn comprehensive artifacts from multiple generators or sacrifice their fitting ability for seen generators (i.e., In-Domain (ID) performance) to exchange the generalization for unseen generators (i.e., Out-Of-Domain (OOD) performance). To tackle the above challenges, we propose our Discrepancy Deepfake Detector (D$^3$) framework, whose core idea is to deconstruct the universal artifacts from multiple generators by introducing a parallel network branch that takes a distorted image feature as an extra discrepancy signal and supplement its original counterpart. Extensive scaled-up experiments demonstrate the effectiveness of D$^3$, achieving 5.3% accuracy improvement in the OOD testing compared to the current SOTA methods while maintaining the ID performance. The source code will be updated in our GitHub repository: https://github.com/BigAandSmallq/D3

cs.CV

LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion Space

While existing one-shot talking head generation models have achieved progress in coarse-grained emotion editing, there is still a lack of fine-grained emotion editing models with high interpretability. We argue that for an approach to be considered fine-grained, it needs to provide clear definitions and sufficiently detailed differentiation. We present LES-Talker, a novel one-shot talking head generation model with high interpretability, to achieve fine-grained emotion editing across emotion types, emotion levels, and facial units. We propose a Linear Emotion Space (LES) definition based on Facial Action Units to characterize emotion transformations as vector transformations. We design the Cross-Dimension Attention Net (CDAN) to deeply mine the correlation between LES representation and 3D model representation. Through mining multiple relationships across different feature and structure dimensions, we enable LES representation to guide the controllable deformation of 3D model. In order to adapt the multimodal data with deviations to the LES and enhance visual quality, we utilize specialized network design and training strategies. Experiments show that our method provides high visual quality along with multilevel and interpretable fine-grained emotion editing, outperforming mainstream methods.

cs.CV

Visible-Infrared Person Re-Identification via Patch-Mixed Cross-Modality Learning

Visible-infrared person re-identification (VI-ReID) aims to retrieve images of the same pedestrian from different modalities, where the challenges lie in the significant modality discrepancy. To alleviate the modality gap, recent methods generate intermediate images by GANs, grayscaling, or mixup strategies. However, these methods could introduce extra data distribution, and the semantic correspondence between the two modalities is not well learned. In this paper, we propose a Patch-Mixed Cross-Modality framework (PMCM), where two images of the same person from two modalities are split into patches and stitched into a new one for model learning. A part-alignment loss is introduced to regularize representation learning, and a patch-mixed modality learning loss is proposed to align between the modalities. In this way, the model learns to recognize a person through patches of different styles, thereby the modality semantic correspondence can be inferred. In addition, with the flexible image generation strategy, the patch-mixed images freely adjust the ratio of different modality patches, which could further alleviate the modality imbalance problem. On two VI-ReID datasets, we report new state-of-the-art performance with the proposed method.

cs.CV

Diffusion in Diffusion: Cyclic One-Way Diffusion for Text-Vision-Conditioned Generation

Originating from the diffusion phenomenon in physics that describes particle movement, the diffusion generative models inherit the characteristics of stochastic random walk in the data space along the denoising trajectory. However, the intrinsic mutual interference among image regions contradicts the need for practical downstream application scenarios where the preservation of low-level pixel information from given conditioning is desired (e.g., customization tasks like personalized generation and inpainting based on a user-provided single image). In this work, we investigate the diffusion (physics) in diffusion (machine learning) properties and propose our Cyclic One-Way Diffusion (COW) method to control the direction of diffusion phenomenon given a pre-trained frozen diffusion model for versatile customization application scenarios, where the low-level pixel information from the conditioning needs to be preserved. Notably, unlike most current methods that incorporate additional conditions by fine-tuning the base text-to-image diffusion model or learning auxiliary networks, our method provides a novel perspective to understand the task needs and is applicable to a wider range of customization scenarios in a learning-free manner. Extensive experiment results show that our proposed COW can achieve more flexible customization based on strict visual conditions in different application settings. Project page: https://wangruoyu02.github.io/cow.github.io/.

cs.CV

EmoSpeaker: One-shot Fine-grained Emotion-Controlled Talking Face Generation

Implementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately and comprehensively capture and express various nuanced emotional states, thereby improving the emotional quality and personalization of generated content. Generating fine-grained facial animations that accurately portray emotional expressions using only a portrait and an audio recording presents a challenge. In order to address this challenge, we propose a visual attribute-guided audio decoupler. This enables the obtention of content vectors solely related to the audio content, enhancing the stability of subsequent lip movement coefficient predictions. To achieve more precise emotional expression, we introduce a fine-grained emotion coefficient prediction module. Additionally, we propose an emotion intensity control method using a fine-grained emotion matrix. Through these, effective control over emotional expression in the generated videos and finer classification of emotion intensity are accomplished. Subsequently, a series of 3DMM coefficient generation networks are designed to predict 3D coefficients, followed by the utilization of a rendering network to generate the final video. Our experimental results demonstrate that our proposed method, EmoSpeaker, outperforms existing emotional talking face generation methods in terms of expression variation and lip synchronization. Project page: https://peterfanfan.github.io/EmoSpeaker/

cs.CV