SearcharxivSearch

arXiv subjects

Zhipeng Li

Publications and source records attributed to Zhipeng Li.

At least 19 recordsLinked to original sources

EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis

Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotional text-to-speech (TTS) systems, however, condition on a single discrete label or static embedding per utterance, fundamentally misaligning with the temporal nature of affect. While recent LLM-based TTS systems may implicitly vary prosody through text understanding, such variation is neither explicitly controllable nor precise enough for targeted intra-utterance transitions. We address three challenges: (1) a multi-pass flow blending pipeline synthesizes frame-aligned transition audio, circumventing the scarcity of natural intra-utterance transitions; (2) dual-stage Valence-Arousal-Dominance (VAD) conditioning guides prosodic planning in the LLM and acoustic realization in the flow decoder via frame-level VAD embeddings; (3) direction-magnitude decoupled injection structurally separates emotion direction from injection magnitude, preventing content degradation. EmoTra-TTS adds only +0.43% parameters with no latency overhead, achieves 30%-87% relative improvement on emotion transition quality, corroborated by 64.4%-79.5% overall win rates in pairwise preference tests against four SOTA baselines and two commercial systems.

eess.AS

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, but a spoken answer still leaves the agent visually absent. We introduce \textbf{Ex-Omni-2D}, a framework that answers a multimodal query with coordinated text, personalized speech, and reference-conditioned video. The dialogue model first writes a structured \textit{Visual Thought Plan} (VTP) for scene, emotion, and motion, then generates the response text and multi-codebook speech units. These speech units are decoded into audio and aligned with video frames, giving the speech and avatar modules a common timing signal while allowing them to learn from different data sources. The video module is trained as a full-sequence Teacher conditioned on reference appearance, VTP semantics, and frame-aligned speech units. We further explore to distill it into a few-step block-causal \emph{Streaming Student}; its Prefix Streaming mechanism carries the previous clean latent into the next chunk and is analyzed as a partial mitigation for late-chunk subject drift. At $400\times720$/$720\times400$, the four-step four-GPU Student provides incremental output with lower startup latency than the full-sequence Teacher.

cs.AI

Educational Short Videos: Bibliometric Trends, Thematic Structure, and Operationalisation

Educational short-video research spans disciplines, platforms and learning contexts, but the same label is applied to resources differing in function, activity, context and evaluation, complicating comparison and evidence synthesis. This study mapped the development, thematic organisation and operationalisation of educational short-video research. We analysed 2,169 records indexed in Web of Science and Scopus up to 11 June 2026 using bibliometric analysis, non-negative matrix factorisation topic modelling and structured content analysis. Publication output increased sharply from the mid-2010s but remained dispersed across outlets. Among 16 first-level topics, Skill Development in Educational Contexts was the largest, forming the structural core of the field, while topic overlap was predominantly pairwise. Video-Based Health Interventions for Attitude Change and Cognitive Load and Engagement in Instructional Video Design combined positive recent growth with comparatively high citation visibility, whereas Social Media Engagement Strategies for Education emerged as a rapidly growing direction. Knowledge/Achievement and Engagement/Motivation were the most widely represented outcome domains, and experimental and synthesis designs were more common among identifiable records in health-related topics. Indexed descriptions most often foregrounded intended users, educational uses, learning content, and interactivity. A representative duration was available for 511 records, but no consistently applied numerical threshold was evident. These findings support a working definition of educational short videos as discrete multimedia messages combining words and visuals and designed or used to promote learning. Shortness should be interpreted relative to educational purpose, content unit and surrounding activity rather than duration alone.

cs.CY

Retrofitting Existing 3D Objects with Surface-Conforming Capacitive Sensing

Augmenting the surface of 3D objects with capacitive sensing is challenging when their volumes cannot be modified. In this paper, we present a generative computational fabrication pipeline that retrofits surface-only sensor layouts to 3D geometries for multi-touch interaction. Our method scans a real-world object to obtain its 3D mesh, generates and optimizes a 3D sensor design of drive and sense lines for mutual-capacitance sensing under physical and hardware constraints, and unfolds the design into individual 2D stencils that can be cut from conductive material. Our fabrication pipeline cuts these stencils from thin copper foil with a vinyl cutter and then assists manual sensor attachment by projecting the sensor design onto the dynamically registered real-world object. We connect the resulting electrode mesh to a mutual-capacitance scanning controller and resolve touch interaction in real time. We demonstrate our approach with four 3D geometries and evaluate our method and fabrication pipeline on them.

cs.HC

M\"obius-like Real-Space Topology Reshapes Spectral Winding Topology in Hatano-Nelson Rings

The spectral winding number serves as a bulk topological invariant in non-Hermitian systems, governing the emergence of skin modes and encoding the non-Hermitian bulk-boundary correspondence. However, most existing studies are built on conventional lattice geometries such as linear chains, rings, or planar arrays, leaving the role of real-space topological connectivity as an independent degree of freedom largely unexplored. Here, we construct a M\"obius ring system by cutting two parallel Hatano-Nelson (HN) rings and reconnecting them with a half-twist, without altering any local hopping parameter. This topological reconstruction transforms the periodic-boundary spectrum from two disjoint ellipses into a multi-petalled rose curve, and leads to distinct decay lengths for different eigenstates under open boundary conditions. Moreover, the spectral winding number can be driven through discrete winding-number jumps by tuning the coupling strength, with critical values obtained analytically. Our results demonstrate that real-space M\"obius connectivity, mediated by the coupling strength, provides an independent and tunable foundation for the systematic control of non-Hermitian topology, with implications for the design of topological devices and sensing schemes.

physics.optics

Probing mesoscopic nonlocal screening in van der Waals heterostructures with polaritons

Predictive optical modelling of van der Waals (vdW) heterostructures is critical for meta-optics, near-field photonics and quantum technologies. At their buried interfaces, charge transfer and spatially extended screening challenge local descriptions based on layer-by-layer stacking of fixed permittivity tensors. However, such nonlocal corrections have been established mainly for plasmonic systems at {\aa}ngstr\"om-nanometre scales and are often assumed negligible on optical-wavelength scales. Here we challenge this view by uncovering a mesoscopic nonlocal screening regime, extending up to ~140 nm, at buried charge-transfer interfaces in transition-metal dichalcogenide/{\alpha}-molybdenum trioxide (TMDC/{\alpha}-MoO3) phonon-polaritonic heterostructures. Using phonon polaritons as an ultrasensitive probe, we quantify charge transfer from polariton-wavelength shifts and find a thickness-independent saturated response as {\alpha}-MoO3 is thinned. Rather than merely complicating optical modelling, this nonlocal saturation turns a design-level correction into an opportunity by yielding a transferable cross-material metric. Across more than 120 devices, this metric scales linearly with the work-function difference between the TMDC and {\alpha}-MoO3. We further identify a lattice-mismatch-set energy threshold for charge transfer, revising Anderson-type band alignment for vdW interfaces.

physics.optics

Automating UI Optimization through Multi-Agentic Reasoning

We present AutoOptimization, a novel multi-objective optimization framework for adapting user interfaces. From a user's verbal preferences for changing a UI, our framework guides a prioritization-based Pareto frontier search over candidate layouts. It selects suitable objective functions for UI placement while simultaneously parameterizing them according to the user's instructions to define the optimization problem. A solver then generates a series of optimal UI layouts, which our framework validates against the user's instructions to adapt the UI with the final solution. Our approach thus overcomes the previous need for manual inspection of layouts and the use of population averages for objective parameters. We integrate multiple agents sequentially within our framework, enabling the system to leverage their reasoning capabilities to interpret user preferences, configure the optimization problem, and validate optimization outcomes.

cs.HC

Preference-Guided Prompt Optimization for Text-to-Image Generation

Generative models are increasingly powerful, yet users struggle to guide them through prompts. The generative process is difficult to control and unpredictable, and user instructions may be ambiguous or under-specified. Prior prompt refinement tools heavily rely on human effort, while prompt optimization methods focus on numerical functions and are not designed for human-centered generative tasks, where feedback is better expressed as binary preferences and demands convergence within few iterations. We present APPO, a preference-guided prompt optimization algorithm. Instead of iterating prompts, users only provide binary preferential feedback. APPO adaptively balances its strategies between exploiting user feedback and exploring new directions, yielding effective and efficient optimization. We evaluate APPO on image generation, and the results show APPO enables achieving satisfactory outcomes in fewer iterations with lower cognitive load than manual prompt editing. We anticipate APPO will advance human-AI collaboration in generative tasks by leveraging user preferences to guide complex content creation.

cs.HC

Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models

Omni-modal large language models (OLLMs) aim to unify multimodal understanding and generation, yet extending them to jointly produce speech and 3D facial animation remains largely underexplored. A key challenge is the mismatch between the discrete semantic reasoning of LLMs and the dense temporal dynamics required for 3D facial motion. We propose Expressive Omni (Ex-Omni), a framework that augments OLLMs with speech-accompanied 3D facial animation. Ex-Omni decouples semantic reasoning from temporal generation through a speech-unit generator with blendshape co-supervision and a non-autoregressive blendshape decoder, where speech units provide temporal scaffolding and hidden speech representations carry facially relevant cues. We further introduce a token-as-query gated fusion (TQGF) interface for controlled semantic injection, as well as InstructS2SF-1200K, a 1.2M-sample weakly supervised dataset for speech-accompanied facial animation. Extensive experiments show that Ex-Omni retains competitive speech QA capability while natively generating coordinated text, speech, and 3D facial animation, and approaches the Audio2Face-3D teacher cascade in synchronization and human preference.

cs.CV

A Globally Convergent Variational Framework for Mode Number Detection via Spectral Cutting Curves

Automatically determining the number of intrinsic mode functions (IMFs) and their center frequencies in Variational Mode Decomposition (VMD) remains an open mathematical challenge. Existing methods rely on heuristic settings, trial-and-error, or recursive extraction lacking theoretical convergence guarantees. We propose a variational framework that endogenously determines the number of modes. Any curve below the spectral amplitude divides the area under the spectrum into 2 parts and generate the connected intervals where spectrum locates above it, whose count defines the modal number K[g] -- a topological functional induced by the cutting curve. Since K[g] is discontinuous and intractable for direct optimization, we seek the optimal cutting curve as a continuous variational surrogate: it separates distinct spectral peaks into individual regions above it while merging noise-induced fragments below. This surrogate adversarially maximizes the integral of g while penalizing its curvature, transforming the problem into iteratively solving a fourth-order boundary value problem via Lagrangian duality. We establish a rigorous proof of global convergence for the dual ascent algorithm in function space. Comprehensive numerical experiments on artificial and real-world signals including ECG data show accurate estimates of IMFs and center frequencies, avoiding redundant modes while ensuring recovery of necessary components, providing a robust, theoretically grounded initialization routine for VMD.

math-ph

Rate-Optimal Streaming Codes Under an Extended Delay Profile for Three-Node Relay Networks With Burst Erasures

This paper investigates streaming codes for three-node relay networks under burst packet erasures with a delay constraint $T$. In any sliding window of $T+1$ consecutive packets, the source-to-relay and relay-to-destination channels may introduce burst erasures of lengths at most $b_1$ and $b_2$, respectively. Let $u = \max\{b_1, b_2\}$ and $v = \min\{b_1, b_2\}$. Singhvi et al. proposed a construction achieving the optimal rate when $u\mid (T-u-v)$. In this paper, we present an extended delay profile method that attains the optimal rate under a relaxed constraint $\frac{T - u - v}{2u - v} \leq \left\lfloor \frac{T - u - v}{u} \right\rfloor$ and it strictly cover restriction $u\mid (T-u-v)$. %Furthermore, we demonstrate that the optimal rate for streaming codes is not achievable when $0< T-u-v<v$ under the convolutional code framework.

cs.IT

BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis

Autoregressive (AR) frameworks have recently achieved remarkable progress in zero-shot text-to-speech (TTS) by leveraging discrete speech tokens and large language model techniques. Despite their success, existing AR-based zero-shot TTS systems face two critical limitations: (i) an inherent speed-quality trade-off, as sequential token generation either reduces frame rates at the cost of expressiveness or enriches tokens at the cost of efficiency, and (ii) a text-oriented supervision mismatch, as cross-entropy loss penalizes token errors uniformly without considering the fine-grained acoustic similarity among adjacent tokens. To address these challenges, we propose BridgeTTS, a novel AR-TTS framework built upon the dual speech representation paradigm BridgeCode. BridgeTTS reduces AR iterations by predicting sparse tokens while reconstructing rich continuous features for high-quality synthesis. Joint optimization of token-level and feature-level objectives further enhances naturalness and intelligibility. Experiments demonstrate that BridgeTTS achieves competitive quality and speaker similarity while significantly accelerating synthesis. Speech demos are available at https://test1562.github.io/demo/.

cs.SD

Rate-Optimal Streaming Codes over Three-Node Relay Networks with Burst Erasures

This paper investigates streaming codes over three-node relay networks under burst packet erasures with a delay constraint $T$. In any sliding window of $T+1$ consecutive packets, the source-to-relay and relay-to-destination channels may introduce burst erasures of lengths at most $b_1$ and $b_2$, respectively. Singhvi et al. proposed a construction achieving the optimal code rate when $\max\{b_1,b_2\}\mid (T-b_1-b_2)$. We construct streaming codes with the optimal rate under the condition $T\geq b_1+b_2+\frac{b_1b_2}{|b_1-b_2|}$, thereby enriching the family of rate-optimal streaming codes for three-node relay networks.

cs.IT

Long-Context Speech Synthesis with Context-Aware Memory

In long-text speech synthesis, current approaches typically convert text to speech at the sentence-level and concatenate the results to form pseudo-paragraph-level speech. These methods overlook the contextual coherence of paragraphs, leading to reduced naturalness and inconsistencies in style and timbre across the long-form speech. To address these issues, we propose a Context-Aware Memory (CAM)-based long-context Text-to-Speech (TTS) model. The CAM block integrates and retrieves both long-term memory and local context details, enabling dynamic memory updates and transfers within long paragraphs to guide sentence-level speech synthesis. Furthermore, the prefix mask enhances the in-context learning ability by enabling bidirectional attention on prefix tokens while maintaining unidirectional generation. Experimental results demonstrate that the proposed method outperforms baseline and state-of-the-art long-context methods in terms of prosody expressiveness, coherence and context inference cost across paragraph-level speech.

eess.AS

Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech

Advances in speech representation and large language models have enhanced zero-shot text-to-speech (TTS) performance. However, existing zero-shot TTS models face challenges in capturing the complex correlations between acoustic and semantic features, resulting in a lack of expressiveness and similarity. The primary reason lies in the complex relationship between semantic and acoustic features, which manifests independent and interdependent aspects.This paper introduces a TTS framework that combines both autoregressive (AR) and non-autoregressive (NAR) modules to harmonize the independence and interdependence of acoustic and semantic information. The AR model leverages the proposed Parallel Tokenizer to synthesize the top semantic and acoustic tokens simultaneously. In contrast, considering the interdependence, the Coupled NAR model predicts detailed tokens based on the general AR model's output. Parallel GPT, built on this architecture, is designed to improve zero-shot text-to-speech synthesis through its parallel structure. Experiments on English and Chinese datasets demonstrate that the proposed model significantly outperforms the quality and efficiency of the synthesis of existing zero-shot TTS models. Speech demos are available at https://t1235-ch.github.io/pgpt/.

eess.AS

Efficient Visual Appearance Optimization by Learning from Prior Preferences

Adjusting visual parameters such as brightness and contrast is common in our everyday experiences. Finding the optimal parameter setting is challenging due to the large search space and the lack of an explicit objective function, leaving users to rely solely on their implicit preferences. Prior work has explored Preferential Bayesian Optimization (PBO) to address this challenge, involving users to iteratively select preferred designs from candidate sets. However, PBO often requires many rounds of preference comparisons, making it more suitable for designers than everyday end-users. We propose Meta-PO, a novel method that integrates PBO with meta-learning to improve sample efficiency. Specifically, Meta-PO infers prior users' preferences and stores them as models, which are leveraged to intelligently suggest design candidates for the new users, enabling faster convergence and more personalized results. An experimental evaluation of our method for appearance design tasks on 2D and 3D content showed that participants achieved satisfactory appearance in 5.86 iterations using Meta-PO when participants shared similar goals with a population (e.g., tuning for a ``warm'' look) and in 8 iterations even generalizes across divergent goals (e.g., from ``vintage'', ``warm'', to ``holiday''). Meta-PO makes personalized visual optimization more applicable to end-users through a generalizable, more efficient optimization conditioned on preferences, with the potential to scale interface personalization more broadly.

cs.HC

Dynamic Parameter Memory: Temporary LoRA-Enhanced LLM for Long-Sequence Emotion Recognition in Conversation

Recent research has focused on applying speech large language model (SLLM) to improve speech emotion recognition (SER). However, the inherently high frame rate in speech modality severely limits the signal processing and understanding capabilities of SLLM. For example, a SLLM with a 4K context window can only process 80 seconds of audio at 50Hz feature sampling rate before reaching its capacity limit. Input token compression methods used in SLLM overlook the continuity and inertia of emotions across multiple conversation turns. This paper proposes a Dynamic Parameter Memory (DPM) mechanism with contextual semantics and sentence-level emotion encoding, enabling processing of unlimited-length audio with limited context windows in SLLM. Specifically, DPM progressively encodes sentence-level information and emotions into a temporary LoRA module during inference to effectively "memorize" the contextual information. We trained an emotion SLLM as a backbone and incorporated our DPM into inference for emotion recognition in conversation (ERC). Experimental results on the IEMOCAP dataset show that DPM significantly improves the emotion recognition capabilities of SLLM when processing long audio sequences, achieving state-of-the-art performance.

cs.CL

Exploring Remote Physiological Signal Measurement under Dynamic Lighting Conditions at Night: Dataset, Experiment, and Analysis

Remote photoplethysmography (rPPG) is a non-contact technique for measuring human physiological signals. Due to its convenience and non-invasiveness, it has demonstrated broad application potential in areas such as health monitoring and emotion recognition. In recent years, the release of numerous public datasets has significantly advanced the performance of rPPG algorithms under ideal lighting conditions. However, the effectiveness of current rPPG methods in realistic nighttime scenarios with dynamic lighting variations remains largely unknown. Moreover, there is a severe lack of datasets specifically designed for such challenging environments, which has substantially hindered progress in this area of research. To address this gap, we present and release a large-scale rPPG dataset collected under dynamic lighting conditions at night, named DLCN. The dataset comprises approximately 13 hours of video data and corresponding synchronized physiological signals from 98 participants, covering four representative nighttime lighting scenarios. DLCN offers high diversity and realism, making it a valuable resource for evaluating algorithm robustness in complex conditions. Built upon the proposed Happy-rPPG Toolkit, we conduct extensive experiments and provide a comprehensive analysis of the challenges faced by state-of-the-art rPPG methods when applied to DLCN. The dataset and code are publicly available at https://github.com/dalaoplan/Happp-rPPG-Toolkit.

cs.CV