SearcharxivSearch

arXiv subjects

He Feng

Publications and source records attributed to He Feng.

12 recordsLinked to original sources

Manipulating non-intrinsic outputs of a non-Hermitian coupled system by weak external driving

Measurement provides access to exploring and understanding nature and finds widespread applications. It is generally believed that the weak external driving used in measurement introduces negligible disturbance to the intrinsic output of the probed system. Here, we reveal that dissipation in a non-Hermitian quantum coupled system induces interference between self-excitation and coupling-excitation spectral functions, giving rise to non-intrinsic outputs. These non-intrinsic outputs can be effectively manipulated by the ratio of weak external driving strengths applied to two channels, leading to significant phenomena: extreme suppression of dissipation by convergence effect of high-loss mode towards low-loss mode, and giant shift of resonant frequency instead of original spectral Rabi splitting. Crucially, we discover the Pythagorean relation linking eigenlevel splitting, spectral Rabi splitting and average total decay rate, enabling precise measurement of the eigenlevel splitting and coupling strength. Numerical simulations confirm that the interference outputs also exist in classical coupled systems. Our work provides a profound insight into measurement and non-Hermitian physics.

quant-ph

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning

Portrait animation methods have achieved substantial visual quality and lip synchronization, but fine-grained manipulation of the eye region still faces a trade-off between input granularity and motion accuracy. Existing methods using emotion labels or coarse text prompts are insufficient for describing subtle ocular dynamics, whereas approaches based on Action Units or driving videos provide higher fidelity at the cost of a heavier input burden. These limitations are still restrictive for beyond-emotion states (e.g., thinking) and drowsiness. In light of the above, we propose CogPortrait, a two-stage framework that generates portrait animations from high-level labels. In the first stage, three chain-of-thought Multimodal Large Language Models (MLLMs) agents compile high-level labels into facial keypoints through temporal event planning, prototype retrieval, and composition from a real-behavior library, and semantic-physiological constraint enforcement. In the second stage, a DiT-based video generation backbone synthesizes the final animation conditioned on the keypoints, reference portrait, audio, and text prompt, enhanced by a dynamic classifier-free guidance strategy with eye-region-aware reweighting and KTO-based refinement for boundary cases. We further introduce the EMH benchmark covering diverse emotions and beyond-emotion categories with two AU-level metrics for evaluating fine-grained eye-region and head-motion control. Extensive experiments on HDTF and the EMH benchmark demonstrate that CogPortrait achieves more precise eye-region control than existing methods while maintaining supe- rior visual quality and identity consistency

cs.CV

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation

Portrait animation aims to synthesize talking videos from a static reference face, conditioned on audio and style frame cues (e.g., emotion and head poses), while ensuring precise lip synchronization and faithful reproduction of speaking styles. Existing diffusion-based portrait animation methods primarily focus on lip synchronization or static emotion transformation, often overlooking dynamic styles such as head movements. Moreover, most of these methods rely on a dual U-Net architecture, which preserves identity consistency but incurs additional computational overhead. To this end, we propose DiTalker, a unified DiT-based framework for speaking style-controllable portrait animation. We design a Style-Emotion Encoding Module that employs two separate branches: a style branch extracting identity-specific style information (e.g., head poses and movements), and an emotion branch extracting identity-agnostic emotion features. We further introduce an Audio-Style Fusion Module that decouples audio and speaking styles via two parallel cross-attention layers, using these features to guide the animation process. To enhance the quality of results, we adopt and modify two optimization constraints: one to improve lip synchronization and the other to preserve fine-grained identity and background details. Extensive experiments demonstrate the superiority of DiTalker in terms of lip synchronization and speaking style controllability. Project Page: https://thenameishope.github.io/DiTalker/

cs.CV

DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation

Human-centric generative models are becoming increasingly popular, giving rise to various innovative tools and applications, such as talking face videos conditioned on text or audio prompts. The core of these capabilities lies in powerful pre-trained foundation models, trained on large-scale, high-quality datasets. However, many advanced methods rely on in-house data subject to various constraints, and other current studies fail to generate high-resolution face videos, which is mainly attributed to the significant lack of large-scale, high-quality face video datasets. In this paper, we introduce a human face video dataset, \textbf{DH-FaceVid-1K}. Our collection spans 1,200 hours in total, encompassing 270,043 video clips from over 20,000 individuals. Each sample includes corresponding speech audio, facial keypoints, and text annotations. Compared to other publicly available datasets, ours distinguishes itself through its multi-ethnic coverage and high-quality, comprehensive individual attributes. We establish multiple face video generation models supporting tasks such as text-to-video and image-to-video generation. In addition, we develop comprehensive benchmarks to validate the scaling law when using different proportions of proposed dataset. Our primary aim is to contribute a face video dataset, particularly addressing the underrepresentation of Asian faces in existing curated datasets and thereby enriching the global spectrum of face-centric data and mitigating demographic biases. \textbf{Project Page:} https://luna-ai-lab.github.io/DH-FaceVid-1K/

cs.CV

One-Shot Pose-Driving Face Animation Platform

The objective of face animation is to generate dynamic and expressive talking head videos from a single reference face, utilizing driving conditions derived from either video or audio inputs. Current approaches often require fine-tuning for specific identities and frequently fail to produce expressive videos due to the limited effectiveness of Wav2Pose modules. To facilitate the generation of one-shot and more consecutive talking head videos, we refine an existing Image2Video model by integrating a Face Locator and Motion Frame mechanism. We subsequently optimize the model using extensive human face video datasets, significantly enhancing its ability to produce high-quality and expressive talking head videos. Additionally, we develop a demo platform using the Gradio framework, which streamlines the process, enabling users to quickly create customized talking head videos.

cs.CV

DialogUSR: Complex Dialogue Utterance Splitting and Reformulation for Multiple Intent Detection

While interacting with chatbots, users may elicit multiple intents in a single dialogue utterance. Instead of training a dedicated multi-intent detection model, we propose DialogUSR, a dialogue utterance splitting and reformulation task that first splits multi-intent user query into several single-intent sub-queries and then recovers all the coreferred and omitted information in the sub-queries. DialogUSR can serve as a plug-in and domain-agnostic module that empowers the multi-intent detection for the deployed chatbots with minimal efforts. We collect a high-quality naturally occurring dataset that covers 23 domains with a multi-step crowd-souring procedure. To benchmark the proposed dataset, we propose multiple action-based generative models that involve end-to-end and two-stage training, and conduct in-depth analyses on the pros and cons of the proposed baselines.

cs.CL

A Mathematica code for calculating massless spectrum of (0,2) Landau-Ginzburg orbifold

In this short paper, we try to explain how to use our program which has been written in Wolfram Mathematica to get the massless spectrum of any Landau-Ginzburg orbifold. The technique has been developed by Witten-Kachru theoretically, but calculating it for an explicit Landau-Ginzburg model is exhausting and in general, beyond human ability to calculate using pen and paper.

hep-th

Heterotic/Heterotic and Heterotic/F-theory Duality

We consider heterotic target space dual (0,2) GLSMs on elliptically fibered Calabi-Yau manifolds. In this context, each half of the "dual" heterotic theories must in turn have an F-theory dual. Moreover, the apparent relationship between two heterotic compactifications seen in (0,2) heterotic target space dual pairs should, in principle, induce some putative correspondence between the dual F-theory geometries. It has previously been conjectured in the literature that (0,2) target space duality might manifest in F-theory as multiple K3-fibrations of the same elliptically fibered Calabi-Yau manifold. In this work we investigate this conjecture in the context of both 6-dimensional and 4-dimensional effective theories and demonstrate that in general, (0,2) target space duality cannot be explained by such a simple phenomenon alone. In all cases, we provide evidence that non-geometric data in F-theory must play at least some role in the induced F-theory correspondence, while leaving the full determination of the putative new F-theory duality to future work.

hep-th

Applicability of coupling strength estimation for linear chains of restricted access

The characterization of an unknown quantum system requires the Hamiltonian identification. The full access to the system, however, is usually restricted, hindering the direct retrieval of relevant parameters, and a reliable indirect estimation is usually required. In this work, the algorithm proposed by Burgarth et al. [Phys. Rev. A 79, 020305 (2009)], which allows estimating the coupling strengths in a linear chain by addressing only one end site, is further investigated. The scheme is numerically studied for states with chain structure, exploring its applicability against observational errors including the limited signal-noise ratio and the finite spectral width. The spectral distribution of the end state is shown to determine the applicability of the method, and reducing the loss from truncated spectral components is critical to realizing the robust reconstruction of coupling strengths.

quant-ph

Dimerized Decomposition of Quantum Evolution on an Arbitrary Graph

The study of quantum evolution on graphs for diversified topologies is beneficial to modeling various realistic systems. A systematic method, the dimerized decomposition, is proposed to analyze the dynamics on an arbitrary network. By introducing global "flows" among interlinked dimerized subsystems, each of which locally consists of an input and a output port, the method provides an intuitive picture that the local properties of the subsystem are separated from the global structure of the network. The pictorial interpretation of quantum evolution as multiple flows through the graph allows for the analysis of the complex network dynamics supplementary to the conventional spectral method.

quant-ph

New Evidence for (0,2) Target Space Duality

In the context of (0,2) gauged linear sigma models, we explore chains of perturbatively dual heterotic string compactifications. The notion of target space duality originates in non-geometric phases and can be used to generate distinct GLSMs with shared geometric phases leading to apparently identical target space theories. To date, this duality has largely been studied at the level of counting states in the effective theories. We extend this analysis to the effective potential and loci of enhanced symmetry in dual theories. By engineering vector bundles with non-trivial constraints arising from slope-stability (i.e. D-terms) and holomorphy (i.e. F-terms) the detailed structure of the vacuum space of the dual theories can be explored. Our results give new evidence that GLSM target space duality may provide important hints towards a more complete understanding of (0,2) string dualities.

hep-th