SearcharxivSearch

arXiv subjects

Li Dai

Publications and source records attributed to Li Dai.

At least 19 recordsLinked to original sources

CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation

Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing methods often rely on discrete emotion categories, which fail to capture the continuous evolution of affect. They also overlook the temporal frequency mismatch between audio articulation and emotional expression. In this paper, we propose CETalk, an audio-driven 3D facial animation framework conditioned on continuous Valence--Arousal (VA) representations for fine-grained emotion control. CETalk predicts a sequence of FLAME parameters through three key components: a Dynamic Emotion Modulation Module that adaptively scales emotional intensity using audio-derived cues; a Multi-Scale Temporal Modeling mechanism that employs parallel branches to decouple high-frequency articulatory movements from low-frequency emotional dynamics; and a Dynamic Fusion Mechanism that integrates these multi-scale features via an adaptive gating network. To support training and evaluation, we construct 3D-VA-MEAD, a large-scale dataset with automatically estimated VA annotations and reconstructed 3D facial motions. Extensive experiments demonstrate that CETalk outperforms state-of-the-art methods in both lip-sync accuracy and emotional expressiveness, while enabling smooth and controllable emotion transitions.

cs.CV

Missing-Token Prompted Reliability-Aware Fusion for Robust Polyglot Speaker Identification

Accurate and robust multimodal speaker identification is essential for multimedia understanding and biometric authentication. However, real-world polyglot scenarios pose two key challenges: speaker-discriminative representations should generalize across languages, and the model should remain reliable when face information is unavailable. To address these challenges, we propose MRAF, a Missing-Token Prompted Reliability-Aware Fusion framework for polyglot speaker identification across complete-modality, missing-face, and cross-lingual scenarios. MRAF represents unavailable face inputs with a learnable missing token instead of fixed zero-valued features, providing a trainable representation of the missing visual state. This design reduces the distribution gap caused by missing inputs and allows subsequent reliability estimation and cross-modal fusion to operate within a unified token space. To adaptively integrate modalities with different reliability, MRAF further introduces a reliability-aware cross-attention fusion module, which estimates face and audio reliability scores, normalizes them into modality weights, and applies these weights to token representations before bidirectional cross-attention. In this way, the model can emphasize reliable modality cues while suppressing unreliable ones. During training, MRAF jointly optimizes multi-branch classification losses, audio-only knowledge distillation, and center loss to improve speaker discrimination and missing-modality robustness. Experiments on the official POLY-SIM 2026 test set demonstrate the effectiveness of the proposed framework. In the final evaluation, MRAF achieves 100% accuracy on P3 and P5, and obtains competitive results on the more challenging missing-face settings P4 and P6. The source code will be released at https://github.com/MSA-LMC/MRAF.

cs.SD

Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention

Accurate engagement estimation is essential for adaptive human-computer interaction systems, yet robust deployment is hindered by poor generalizability across diverse domains and challenges in modeling complex interaction dynamics.To tackle these issues, we propose DAPA (Domain-Adaptive Parallel Attention), a novel framework for generalizable conversational engagement modeling. DAPA introduces a Domain Prompting mechanism by prepending learnable domain-specific vectors to the input, explicitly conditioning the model on the data's origin to facilitate domain-aware adaptation while preserving generalizable engagement representations. To capture interactional synchrony, the framework also incorporates a Parallel Cross-Attention module that explicitly aligns reactive (forward BiLSTM) and anticipatory (backward BiLSTM) states between participants.Extensive experiments demonstrate that DAPA establishes a new state-of-the-art performance on several cross-cultural and cross-linguistic benchmarks, notably achieving an absolute improvement of 0.45 in Concordance Correlation Coefficient (CCC) over a strong baseline on the NoXi-J test set. The superiority of our method was also confirmed by winning the first place in the Multi-Domain Engagement Estimation Challenge at MultiMediate'25.

cs.CV

Stochastic Tube-based Model Predictive Control for Cyber-Physical Systems under False Data Injection Attacks with Bounded Probability

This paper addresses the challenge of amplitude-unbounded false data injection (FDI) attacks targeting the sensor-to-controller (S-C) channel in cyber-physical systems (CPSs). We introduce a resilient tube-based model predictive control (MPC) scheme. This scheme incorporates a threshold-based attack detector and a control sequence buffer to enhance system security. We mathematically model the common FDI attacks and derive the maximum duration of such attacks based on the hypothesis testing principle. Following this, the minimum feasible sequence length of the control sequence buffer is obtained. The system is proven to remain input-to-state stable (ISS) under bounded external disturbances and amplitude-unbounded FDI attacks. Moreover, the feasible region under this scenario is provided in this paper. Finally, the proposed algorithm is validated by numerical simulations and shows superior control performance compared to the existing methods.

eess.SY

CMQCIC-Bench: A Chinese Benchmark for Evaluating Large Language Models in Medical Quality Control Indicator Calculation

Medical quality control indicators are essential to assess the qualifications of healthcare institutions for medical services. With the impressive performance of large language models (LLMs) like GPT-4 in the medical field, leveraging these technologies for the Medical Quality Control Indicator Calculation (MQCIC) presents a promising approach. In this work, (1) we introduce a real-world task MQCIC and propose an open-source Chinese electronic medical records (EMRs)-based dataset (CMQCIC-Bench) comprising 785 instances and 76 indicators. (2) We propose a semi-automatic method to enhance the rule representation. Then we propose the Clinical Facts-based Inferential Rule (CF-IR) method that disentangles the clinical fact verification and inferential rule reasoning actions. (3) We conduct comprehensive experiments on 20 representative LLMs, covering general and medical models. Our findings reveal that CF-IR outperforms Chain-of-Thought methods in MQCIC tasks. (4) We conduct an error analysis and investigate the capabilities of clinical fact verification and inferential rule reasoning, providing insights to improve performance in the MQCIC further. The dataset and code is available in this repository https://github.com/YuY-2001/C-MQCIC.

cs.CL

Workflow-based Fast Data-driven Predictive Control with Disturbance Observer in Cloud-edge Collaborative Architecture

Data-driven predictive control (DPC) has been studied and used in various scenarios, since it could generate the predicted control sequence only relying on the historical input and output data. Recently, based on cloud computing, data-driven predictive cloud control system (DPCCS) has been proposed with the advantage of sufficient computational resources. However, the existing computation mode of DPCCS is centralized. This computation mode could not utilize fully the computing power of cloud computing, of which the structure is distributed. Thus, the computation delay could not been reduced and still affects the control quality. In this paper, a novel cloud-edge collaborative containerised workflow-based DPC system with disturbance observer (DOB) is proposed, to improve the computation efficiency and guarantee the control accuracy. First, a construction method for the DPC workflow is designed, to match the distributed processing environment of cloud computing. But the non-computation overheads of the workflow tasks are relatively high. Therefore, a cloud-edge collaborative control scheme with DOB is designed. The low-weight data could be truncated to reduce the non-computation overheads. Meanwhile, we design an edge DOB to estimate and compensate the uncertainty in cloud workflow processing, and obtain the composite control variable. The UUB stability of the DOB is also proved. Third, to execute the workflow-based DPC controller and evaluate the proposed cloud-edge collaborative control scheme with DOB in the real cloud environment, we design and implement a practical workflow-based cloud control experimental system based on container technology. Finally, a series of evaluations show that, the computation times are decreased by 45.19% and 74.35% for two real-time control examples, respectively, and by at most 85.10% for a high-dimension control example.

eess.SY

A Hierarchical Robust Control Strategy for Decentralized Signal-Free Intersection Management

The development of connected and automated vehicles is the key to improving urban mobility safety and efficiency. This paper focuses on cooperative vehicle management at a signal-free intersection with consideration of vehicle modeling uncertainties and sensor measurement disturbances. The problem is approached by a hierarchical robust control strategy in a decentralized traffic coordination framework where optimal control and tube-based robust model predictive control methods are designed to hierarchically solve the optimal crossing order and the velocity trajectories of a group of CAVs in terms of energy consumption and throughput. To capture the energy consumption of each vehicle, their powertrain system is modeled in line with an electric drive system. With a suitable relaxation and spatial modeling approach, the optimization problems in the proposed strategy can be formulated as convex second-order cone programs, which provide a unique and computationally efficient solution. A rigorous proof of the equivalence between the convexified and the original problems is also provided. Simulation results illustrate the effectiveness and robustness of the proposed strategy and reveal the impact of traffic density on the control solution. The study of the Pareto optimal solutions for the energy-time objective shows that a minor reduction in journey time can considerably reduce energy consumption, which emphasizes the necessity of optimizing their trade-off. Finally, the numerical comparisons carried out for different prediction horizons and sampling intervals provide insight into the control design.

math.OC

Cloud-based computational model predictive control using a parallel multi-block ADMM approach

Heavy computational load for solving nonconvex problems for large-scale systems or systems with real-time demands at each sample step has been recognized as one of the reasons for preventing a wider application of nonlinear model predictive control (NMPC). To improve the real-time feasibility of NMPC with input nonlinearity, we devise an innovative scheme called cloud-based computational model predictive control (MPC) by using an elaborately designed parallel multi-block alternating direction method of multipliers (ADMM) algorithm. This novel parallel multi-block ADMM algorithm is tailored to tackle the computational issue of solving a nonconvex problem with nonlinear constraints.

math.OC

Minimal Leader Set for Controllability of k-distant Trees

Minimal controllability problem plays an important role in the field of network control. A New concept-Minimum Perfect Critical Set (MPCS)is proposed. Four different MPCSs were found for k-distant tree graphs. Based on this concept of MPCS, an algorithm for finding the minimal leader set is provided. Numerical experiments show that these theories enable the algorithm to find a minimal leader set with a probability of more than 0.98. Further, some other numerical characteristics of the minimal leader set of k-distant trees were found.

math.OC

Design and Implementation of Data-driven Predictive Cloud Control System

Nowadays, the rapid increases of the scale and complexity of the controlled plants bring new challenges such as computing power and storage for conventional control systems. Cloud computing is concerned as a powerful solution to handle the complex large-scale control missions using sufficient computing resources. However, the developed computing ability enables more complex devices and mass data being involved and thus the applications of model-based algorithms are constrained. Motivated by the above, we propose an original data-driven predictive cloud control system. To achieve the proposed system, a practical data-driven predictive cloud control platform rather than only a numerical simulator is established and together a cloud-edge communication scheme is developed. Finally, the verification of simulations and experiments as well as discussions demonstrate the effectiveness of the proposed system.

eess.SY

Minimum leader selection for Structural Controllability of Undirected Graphs with Leader-follower Framework

The optimization problem of the minimum set of leaders for the controllability of undirected graphs are addressed. It is difficult to find not only its optimal solution but also its approximate algorithm. We propose a new concept, namely minimal perfect critical set (MPCS), to obtain an optimal solution. Some properties are presented, and on the basis of these theorems, the problem of the minimum set of leaders of two typical self-similar bipartite networks, namely deterministic scale-free networks (DSFN) and Cayley trees, is solved completely.

math.OC

On the Leaders' Graphical Characterization for Controllability of Path Related Graphs

The problem of leaders location plays an important role in the controllability of undirected graphs.The concept of minimal perfect critical vertex set is introduced by drawing support from the eigenvector of Laplace matrix. Using the notion of minimal perfect critical vertex set, the problem of finding the minimum number of controllable leader vertices is transformed into the problem of finding all minimal perfect critical vertex sets. Some necessary and sufficient conditions for special minimal perfect critical vertex sets are provided, such as minimal perfect critical 2 vertex set, and minimal perfect critical vertex set of path or path related graphs. And further, the leaders location problem for path graphs is solved completely by the algorithm provided in this paper. An interesting result that there never exist a minimal perfect critical 3 vertex set is proved, too.

math.OC

Detecting the entanglement of vortices in ultracold bosons with artificial gauge fields

The entanglement of vortices in a two-dimensional Bose-Hubbard model with artificial gauge fields is investigated using the exact diagonalization techniques. We propose an effective Hamiltonian for the spin-spin interactions between vortices responsible for this entanglement, and show that the entanglement can be detected through the quantum interference of the bosons in the vortex centers achieved using the Raman coupling and the quantum gas microscope. The strong bosonic coherence between the vortex centers originates from the charge-density wave order in the vortex core. It is robust against the varying of the pinning strength for the vortices to a wide range, and the coherent bosons can be viewed as a qubit stored in the ground state of the system. Our proposal provides a feasible scheme of quantum memory for storing qubits useful in quantum computation.

quant-ph

Non-trivially graded self-dual fusion categories of rank $4$

Let $\mathcal{C}$ be a self-dual spherical fusion categories of rank $4$ with non-trivial grading. We complete the classification of Grothendieck ring $K(\mathcal{C})$ of $\mathcal{C}$; that is, we prove that $K(\mathcal{C})\cong Fib\otimes\mathbb{Z}[\mathbb{Z}_2]$, where $Fib$ is the Fibonacci fusion ring and $\mathbb{Z}[\mathbb{Z}_2]$ is the group ring on $\mathbb{Z}_2$. In particular, if $\mathcal{C}$ is braided then it is equivalent to $\textbf{Fib}\boxtimes\textbf{Vec}_{\mathbb{Z}_2}^ω$ as fusion categories, where $\textbf{Fib}$ is a Fibonacci category and $\textbf{Vec}_{\mathbb{Z}_2}^ω$ is a rank $2$ pointed fusion category.

math.RA

On semisimple quasitriangular Hopf algebras of dimension $dq^n$

Let $q>2$ be a prime number, $d$ be an odd square-free natural number, and $n$ be a non-negative integer. We prove that a semisimple quasitriangular Hopf algebra of dimension $dq^n$ is solvable in the sense of Etingof, Nikshych and Ostrik. In particular, if $n\leq 3$ then it is either isomorphic to $k^G$ for some abelian group $G$, or twist equivalent to a Hopf algebra which fits into a cocentral abelian exact sequence.

math.RA

Robust MPC for tracking of nonholonomic robots with additive disturbances

In this paper, two robust model predictive control (MPC) schemes are proposed for tracking control of nonholonomic systems with bounded disturbances: tube-MPC and nominal robust MPC (NRMPC). In tube-MPC, the control signal consists of a control action and a nonlinear feedback law based on the deviation of the actual states from the states of a nominal system. It renders the actual trajectory within a tube centered along the optimal trajectory of the nominal system. Recursive feasibility and input-to-state stability are established and the constraints are ensured by tightening the input domain and the terminal region. While in NRMPC, an optimal control sequence is obtained by solving an optimization problem based on the current state, and the first portion of this sequence is applied to the real system in an open-loop manner during each sampling period. The state of nominal system model is updated by the actual state at each step, which provides additional a feedback. By introducing a robust state constraint and tightening the terminal region, recursive feasibility and input-to-state stability are guaranteed. Simulation results demonstrate the effectiveness of both strategies proposed.

eess.SY

Entanglement convertibility by sweeping through the quantum phases of the alternating bonds $XXZ$ chain

We study the entanglement structure and the topological edge states of the ground state of the spin-1/2 XXZ model with bond alternation. We employ parity-density matrix renormalization group with periodic boundary conditions. The finite-size scaling of Rényi entropies $S_2$ and $S_\infty$ are used to construct the phase diagram of the system. The phase diagram displays three possible phases: Haldane type (an example of symmetry protected topological ordered phases), Classical Dimer and Néel phases, the latter bounded by two continuous quantum phase transitions. The entanglement and non-locality in the ground state are studied and quantified by the entanglement convertibility. We found that, at small spatial scales, the ground state is not convertible within the topological Haldane dimer phase. The phenomenology we observe can be described in terms of correlations between edge states. We found that the entanglement spectrum also exhibits a distinctive response in the topological phase: the effective rank of the reduced density matrix displays a specifically large "susceptibility" in the topological phase. These findings support the idea that although the topological order in the ground state cannot be detected by local inspection, the ground state response at local scale can tell the topological phases apart from the non-topological phases.

quant-ph