SearcharxivSearch

arXiv subjects

Di Gao

Publications and source records attributed to Di Gao.

15 recordsLinked to original sources

$D\rightarrow \pi$ transitions from QCD Light-Cone Sum Rules with the chiral currents

We present a systematic study of the transition form factors for the semileptonic decays $D\rightarrow Pe^{+}v_{e}$ and $D\rightarrow P\mu^{+}\bar{v}_{\mu}$ using QCD light-cone sum rules with chiral currents, with emphasis on the nonperturbative structure of the pion. The distribution amplitudes of the meson $\pi$, including twist-2 and higher-twist components, are analyzed and incorporated in our computation, employing both Gegenbauer polynomial expansions and the Basis Light-Front Quantization (BLFQ) method, where the latter yields a distribution that rapidly converges to the asymptotic form. Two relations linking transition form factors are derived and found to be consistent with the chiral symmetry of QCD. Our numerical predictions for the $D$-to-pion form factors agee well with BES\uppercase\expandafter{\romannumeral3} measurements and lattice QCD calculations. Furthermore, the differential decay widths obtained thereby for the $D$-to-pion decay also show consistency with the BES\uppercase\expandafter{\romannumeral3} data in the high-momentum-transfer ($q$) region with the $q^{2}>1\ \rm{GeV^2}$, implying new opportunities for the precise extraction of the CKM matrix element and the exploration of physics beyond the Standard Model.

hep-ph

Revisiting Ripple Effects in Knowledge Editing through Pressure-Aware Joint Neighborhood Optimization

Single-edit updates in large language models can trigger ripple effects across local knowledge neighborhoods: desirable propagation to related facts and unintended perturbation of preserved ones. Existing methods address these two effects separately, without explicitly modeling their coupling. We challenge this separation through an analysis of ripple responses across typical baselines, identifying two coupled design pressures: editable-side coordination and preserved-side leakage. We propose Joint Neighborhood Optimization (JNO), a new knowledge-editing framework to formalize and jointly address both pressures at the target-planning stage. JNO instantiates this principle through Pressure-Aware Coordination (PAC), which jointly optimizes neighborhood target representations under coupled constraints, and a semantic pre-execution gate that rejects high-risk target plans before parameter execution. Experiments on RippleEdits show JNO improves propagation and preservation metrics by at least 7.0% while preserving cross-backbone editing stability.

cs.AI

Two to Tango: Coupled Task-Reference Selection for Safe LLM Fine-tuning

Fine-tuning safety aligned large language models (LLMs) on downstream data improves adaptation but may erode learned safety behavior. Existing methods use fixed safety examples, global constraints, or one-sided task filtering. Our diagnostics show task updates expose different safety constraints, motivating joint selection of relevant references and compatible task samples. We propose DualSelect, a coupled framework for task and reference selection that refreshes task conditioned safety references before filtering whole task samples compatible with the induced reference direction. Under a minimax view, DualSelect selects safety references with high preservation loss and task conflict, together with compatible task samples, through entropy-regularized scoring surrogates, lazy reference refresh, and gradient correction. On 1B-8B LLMs, DualSelect preserves safety without losing task utility; using the REDORCA judge, it improves Safety Avg. over the strongest baseline by at least 5.10 points and remains highest in Safety Avg. across judges with moderate overhead. This view extends to retention focused continual learning.

cs.LG

Transition form factors for $D/B$ to $a_{0}(980)$ from light-cone QCD sum rules

We apply the light-cone QCD sum rules with chiral currents to compute transition form factors of the semileptonic decay of the charmed and bottom scalar mesons $D/B$ to $a_{0}(980)l^{+}\nu_{l}$ $(l=e$, $\mu)$, where the scalar meson $a_{0}(980)$ are firstly regarded as two quark states, and contributions from the tetraquark component are then added via proper modulation of four distribution amplitudes considered. The obtained transition form factors and branching fraction are free of the contribution from the light-cone distribution amplitudes at the level of twist-three and two simple relations connecting their form factors are obtained. Our computations indicates that the decay branch fractions are on the margin of the measurements reported by Beijing Spectrometer \uppercase\expandafter{\romannumeral3} in the pure two-quark scenario and are in good agreement with observations when tetraquark component is considered.

hep-ph

MetaKE: Meta-Learning for Knowledge Editing Toward a Better Accuracy-Editability Trade-off

Existing locate-then-edit Knowledge Editing (KE) methods typically decompose editing into two stages: upstream target representation optimization and downstream constrained parameter optimization. The optimization across the two stages is disconnected: upstream applies uniform regularization without observing downstream realization of the planned residual, hindering a refined accuracy-editability trade-off. Since this realization is request-specific and depends on downstream constraints, uniform regularization can over-shrink high-association requests, causing insufficient editing, while it can under-regularize low-association requests, producing over-large planned residuals that reduce downstream editability. To bridge this disconnect, we propose MetaKE (Meta-learning for Knowledge Editing), a new framework that unifies upstream and downstream stages into a bi-level optimization problem. The inner level optimizes parameter updates for the target representation, while the outer level optimizes representation using feedback from downstream constraints, achieving a better semantic accuracy-editability trade-off. To avoid costly multi-layer backpropagation, we introduce a Structural Gradient Proxy to approximate and propagate this feedback. Extensive experiments show that MetaKE outperforms strong baselines, offering a new perspective on KE.

cs.CL

Radiative Decays of Vector Mesons with Light-Cone Sum Rules

Hadronic electromagnetic form factors and radiative decay properties offer a crucial window into the nonperturbative dynamics of Quantum chromodynamics (QCD). In this work, we employ the light-cone sum rules (LCSR) method to systematically investigate the M1 radiative decay of vector mesons. Our study covers processes including $K^{*-}\rightarrow K^-\gamma$, $D^*\rightarrow D\gamma$, $B^*\rightarrow B\gamma$, $D^{*+}_s\rightarrow D^+_s\gamma$, and $B_s^*\rightarrow B_s\gamma$, and further extends to the excited charmonium state $\psi(2S)$. Our calculations yield decay widths for $K^*$ and $\psi(2S)$ that are in excellent agreement with experimental data. For the charm and bottom meson decays, where precise measurements are lacking, we provide theoretical predictions and compare them with other theoretical approaches. Most notably, our analysis reveals a universal linear dependence of the decay width on a function A(x) in the logarithmic coordinate system, which originates from the two-body decay dynamics and the ratio of the initial and final state decay constants. This relationship holds for the ground state $V \rightarrow P \gamma $ processes here and suggests a broader applicability to radiative decays of ground-state vector mesons.

hep-ph

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation

Emotional talking-head generation has emerged as a pivotal research area at the intersection of computer vision and multimodal artificial intelligence, with its core value lying in enhancing human-computer interaction through immersive and empathetic engagement.With the advancement of multimodal large language models, the driving signals for emotional talking-head generation has shifted from audio and video to more flexible text. However, current text-driven methods rely on predefined discrete emotion label texts, oversimplifying the dynamic complexity of real facial muscle movements and thus failing to achieve natural emotional expressiveness.This study proposes the Think-Before-Draw framework to address two key challenges: (1) In-depth semantic parsing of emotions--by innovatively introducing Chain-of-Thought (CoT), abstract emotion labels are transformed into physiologically grounded facial muscle movement descriptions, enabling the mapping from high-level semantics to actionable motion features; and (2) Fine-grained expressiveness optimization--inspired by artists' portrait painting process, a progressive guidance denoising strategy is proposed, employing a "global emotion localization--local muscle control" mechanism to refine micro-expression dynamics in generated videos.Our experiments demonstrate that our approach achieves state-of-the-art performance on widely-used benchmarks, including MEAD and HDTF. Additionally, we collected a set of portrait images to evaluate our model's zero-shot generation capability.

cs.CV

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition

Dynamic Facial Expression Recognition (DFER) aims to identify human emotions from temporally evolving facial movements and plays a critical role in affective computing. While recent vision-language approaches have introduced semantic textual descriptions to guide expression recognition, existing methods still face two key limitations: they often underutilize the subtle emotional cues embedded in generated text, and they have yet to incorporate sufficiently effective mechanisms for filtering out facial dynamics that are irrelevant to emotional expression. To address these gaps, We propose GRACE, Granular Representation Alignment for Cross-modal Emotion recognition that integrates dynamic motion modeling, semantic text refinement, and token-level cross-modal alignment to facilitate the precise localization of emotionally salient spatiotemporal features. Our method constructs emotion-aware textual descriptions via a Coarse-to-fine Affective Text Enhancement (CATE) module and highlights expression-relevant facial motion through a motion-difference weighting mechanism. These refined semantic and visual signals are aligned at the token level using entropy-regularized optimal transport. Experiments on three benchmark datasets demonstrate that our method significantly improves recognition performance, particularly in challenging settings with ambiguous or imbalanced emotion classes, establishing new state-of-the-art (SOTA) results in terms of both UAR and WAR.

cs.CV

Data Poisoning in Deep Learning: A Survey

Deep learning has become a cornerstone of modern artificial intelligence, enabling transformative applications across a wide range of domains. As the core element of deep learning, the quality and security of training data critically influence model performance and reliability. However, during the training process, deep learning models face the significant threat of data poisoning, where attackers introduce maliciously manipulated training data to degrade model accuracy or lead to anomalous behavior. While existing surveys provide valuable insights into data poisoning, they generally adopt a broad perspective, encompassing both attacks and defenses, but lack a dedicated, in-depth analysis of poisoning attacks specifically in deep learning. In this survey, we bridge this gap by presenting a comprehensive and targeted review of data poisoning in deep learning. First, this survey categorizes data poisoning attacks across multiple perspectives, providing an in-depth analysis of their characteristics and underlying design princinples. Second, the discussion is extended to the emerging area of data poisoning in large language models(LLMs). Finally, we explore critical open challenges in the field and propose potential research directions to advance the field further. To support further exploration, an up-to-date repository of resources on data poisoning in deep learning is available at https://github.com/Pinlong-Zhao/Data-Poisoning.

cs.CR

Twist-2 distribution amplitudes of $a_{0}(980)$ and $a_{0}(1450)$

We investigate the twist-2 distribution amplitudes of the scalar mesons $a_{0}(980)$ and $a_{0}(1450)$ in the two-quark picture. The moments of these scalar mesons are obtained up to the third order with QCD sum rules method. With these moments, the first two Gegenbauer coefficients are determined and utilized to analyze the twist-2 distribution amplitudes. Our numerical results indicate that the meson $a_{0}(980)$ favors a conventional two-quark ground state. The paper concludes with an examination of the form factors for the transitions $B/D\rightarrow a_{0}$.

hep-ph

Masses of doubly heavy tetraquarks $QQ\bar{n}\bar{q}$ with $J^{P}=1^{+}$

We apply the method of QCD sum rules to study the doubly heavy tetraquark states $QQ\bar{q}\bar{n}$ with spin-parity $J^{P}=1^{+}$ and strangeness $S=0, -1$ using careful estimates of the Borel and threshold parameters involved. Masses of the doubly bottom and charmed tetraquarks with isospin $I=0,1/2, 1$ are computed precisely via taking into account multifarious condensates up to dimension $10$. Comparing with the two-heavy meson thresholds, we find that all nonstrange doubly-bottom tetraquarks and a doubly-charmed tetraquarks associted with $J_{3}$ with $J^{P}=1^{+}$ are stable against strong decay into two bottom mesons while a doubly-charmed tetraquarks associated with current $J_{2}$ is unstable against strong decay. By the way, weak decay widths of the doubly bottom tetraquarks are also given.

hep-ph

Study of form factors and branching ratios for $D\rightarrow S,Al\bar{ν_{l}}$ with light-cone sum rules

We systematically study the semileptonic decay process of $ D\rightarrow S,A l\bar{ν_{l}}(l=e,μ)$ by light-cone sum rules (LCSR) with chiral currents, calculate the form factors containing only the contribution of the leading twist light-cone distribution amplitudes (LCDAs). For scalar mesons $a_{0}(980)$ and $a_{0}(1450)$, we take them as $q\bar{q}$ states. For axial-vector meson, we study $a_{1}(1260)( 1^{3}p^{1})$ and $b_{1}(1235)( 1^{1}p^{1})$. Based on the results of these form factors, we further present the branching ratios of these semileptonic decay processes. The numerical results for $ D\rightarrow a_{0}(980), b_{1}(1235)l\bar{ν_{l}} $ are in good agreement with experiments and that for $ D\rightarrow a_{0}(1450)l\bar{ν_{l}}$ process are expected to be tested experimentally in the future.

hep-ph

Private Knowledge Transfer via Model Distillation with Generative Adversarial Networks

The deployment of deep learning applications has to address the growing privacy concerns when using private and sensitive data for training. A conventional deep learning model is prone to privacy attacks that can recover the sensitive information of individuals from either model parameters or accesses to the target model. Recently, differential privacy that offers provable privacy guarantees has been proposed to train neural networks in a privacy-preserving manner to protect training data. However, many approaches tend to provide the worst case privacy guarantees for model publishing, inevitably impairing the accuracy of the trained models. In this paper, we present a novel private knowledge transfer strategy, where the private teacher trained on sensitive data is not publicly accessible but teaches a student to be publicly released. In particular, a three-player (teacher-student-discriminator) learning framework is proposed to achieve trade-off between utility and privacy, where the student acquires the distilled knowledge from the teacher and is trained with the discriminator to generate similar outputs as the teacher. We then integrate a differential privacy protection mechanism into the learning procedure, which enables a rigorous privacy budget for the training. The framework eventually allows student to be trained with only unlabelled public data and very few epochs, and hence prevents the exposure of sensitive training data, while ensuring model utility with a modest privacy budget. The experiments on MNIST, SVHN and CIFAR-10 datasets show that our students obtain the accuracy losses w.r.t teachers of 0.89%, 2.29%, 5.16%, respectively with the privacy bounds of (1.93, 10^-5), (5.02, 10^-6), (8.81, 10^-6). When compared with the existing works \cite{papernot2016semi,wang2019private}, the proposed work can achieve 5-82% accuracy loss improvement.

cs.CR

Eva-CiM: A System-Level Performance and Energy Evaluation Framework for Computing-in-Memory Architectures

Computing-in-Memory (CiM) architectures aim to reduce costly data transfers by performing arithmetic and logic operations in memory and hence relieve the pressure due to the memory wall. However, determining whether a given workload can really benefit from CiM, which memory hierarchy and what device technology should be adopted by a CiM architecture requires in-depth study that is not only time consuming but also demands significant expertise in architectures and compilers. This paper presents an energy evaluation framework, Eva-CiM, for systems based on CiM architectures. Eva-CiM encompasses a multi-level (from device to architecture) comprehensive tool chain by leveraging existing modeling and simulation tools such as GEM5, McPAT [2] and DESTINY [3]. To support high-confidence prediction, rapid design space exploration and ease of use, Eva-CiM introduces several novel modeling/analysis approaches including models for capturing memory access and dependency-aware ISA traces, and for quantifying interactions between the host CPU and CiM modules. Eva-CiM can readily produce energy estimates of the entire system for a given program, a processor architecture, and the CiM array and technology specifications. Eva-CiM is validated by comparing with DESTINY [3] and [4], and enables findings including practical contributions from CiM-supported accesses, CiM-sensitive benchmarking as well as the pros and cons of increased memory size for CiM. Eva-CiM also enables exploration over different configurations and device technologies, showing 1.3-6.0X energy improvement for SRAM and 2.0-7.9X for FeFET-RAM, respectively.

cs.AR

A Comparative Study for Non-rigid Image Registration and Rigid Image Registration

Image registration algorithms can be generally categorized into two groups: non-rigid and rigid. Recently, many deep learning-based algorithms employ a neural net to characterize non-rigid image registration function. However, do they always perform better? In this study, we compare the state-of-art deep learning-based non-rigid registration approach with rigid registration approach. The data is generated from Kaggle Dog vs Cat Competition \url{https://www.kaggle.com/c/dogs-vs-cats/} and we test the algorithms' performance on rigid transformation including translation, rotation, scaling, shearing and pixelwise non-rigid transformation. The Voxelmorph is trained on rigidset and nonrigidset separately for comparison and we also add a gaussian blur layer to its original architecture to improve registration performance. The best quantitative results in both root-mean-square error (RMSE) and mean absolute error (MAE) metrics for rigid registration are produced by SimpleElastix and non-rigid registration by Voxelmorph. We select representative samples for visual assessment.

cs.CV