SearcharxivSearch

arXiv subjects

Xiyuan Gao

Publications and source records attributed to Xiyuan Gao.

At least 19 recordsLinked to original sources

Towards a new paradigm of scientific discovery with socialized artificial intelligence

Scientific discovery has advanced through successive transformations in the organization of knowledge. Observation and experimentation established the empirical foundations of science. Theory made it possible to derive general principles from particular phenomena. Computation extended inquiry into systems beyond direct observation, while data-intensive methods opened new spaces of pattern and prediction. Science now confronts a different frontier. The central challenge is no longer simply to produce more information, but to organize expanding knowledge, reasoning, and evidence into a coherent process of discovery. Here, we introduce Bridging Literature, Agents, and Zero-gap Experimentation (BLAZE), a paradigm of socialized scientific intelligence. BLAZE conceives AI not as an assistant for isolated research tasks, but as an organizational infrastructure for scientific discovery. It connects persistent knowledge, collective reasoning, empirical validation, and human judgment within a continuous research lifecycle, transforming fragmented activities into a cumulative process of inquiry, criticism, and revision. The central premise of BLAZE is that scientific intelligence does not arise from computation alone. It emerges from the sustained interaction among knowledge, hypotheses, experiments, and collective verification. By organizing humans and machines within a shared scientific process, BLAZE makes discovery more traceable, reproducible, and cumulative while preserving human creativity, judgment, and responsibility. Socialized scientific intelligence may provide a foundation for the next era of science. Its purpose is not to replace human discovery, but to extend the scale, depth, and continuity of collective scientific inquiry.

cs.AI

Towards the Swampland of Flavour Symmetries

In this thesis, we demonstrate the predictive power of the flavour models without explicit flavour symmetries, which we call the `swampland of flavour symmetries'. Firstly, we revisit the minimal type II seesaw model, which extends the Standard Model (SM) with a TeV-scale triplet scalar field. We find that certain flavour textures of the neutrino mass matrix can suppress the tightly constrained $μ\to e$ transition rates and allow a $5-6$ TeV effective cut-off scale, even when the relevant flavour symmetries are all strongly broken. Next, we show two examples in which some SM flavour parameters are calculable, but the underlying symmetries remain implicit. (i) In the most minimal $SO(10)$ theory, although the quark-lepton symmetry does not manifest at low energies, we find that the $b-τ$ mass ratio can be correctly predicted when the leptoquarks contained in the scalar sector lie at TeV scale, as motivated by the long-standing $B$ anomalies. (ii) We identify a new class of anomaly-free chiral symmetries in the SM (including three right-handed neutrinos), referred to as enhanced $B-L$, under which all neutrinos are massless. In this framework, the neutrino masses are not free parameters but arise solely through the dynamical symmetry breaking effects, known as neutrino condensate. Furthermore, we note that the swampland of flavour symmetries could also involve light new particles, in particular the flavourful axions. As a preliminary study, we calculate an overlooked two-loop contribution to the axion flavour violating interactions in a minimal axion model, and analyze its impact on explaining a recent excess at Belle II.

hep-ph

Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework

Sarcasm is a pragmatic phenomenon in which speakers convey meanings that diverge from literal content, relying on an interaction between semantics and prosodic expression. However, how these cues jointly contribute to the recognition of sarcasm remains poorly understood. We propose a computational framework that models sarcasm as the integration of semantic interpretation and prosodic realization. Semantic cues are derived from an LLaMA 3 model fine-tuned to capture discourse-level markers of sarcastic intent, while prosodic cues are extracted through semantically aligned utterances drawn from a database of sarcastic speech, providing prosodic exemplars of sarcastic delivery. Using a speech synthesis testbed, perceptual evaluations show that semantic and prosodic cues enhance perceived sarcasm, with the combined system achieving the best downstream F1 while maintaining high subjective sarcasm ratings. These findings highlight the complementary roles of semantics and prosody in pragmatic interpretation and illustrate how modeling can shed light on the mechanisms underlying sarcastic communication.

cs.CL

Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection

Sarcasm fundamentally alters meaning through tone and context, yet detecting it in speech remains a challenge due to data scarcity. In addition, existing detection systems often rely on multimodal data, limiting their applicability in contexts where only speech is available. To address this, we propose an annotation pipeline that leverages large language models (LLMs) to generate a sarcasm dataset. Using a publicly available sarcasm-focused podcast, we employ GPT-4o and LLaMA 3 for initial sarcasm annotations, followed by human verification to resolve disagreements. We validate this approach by comparing annotation quality and detection performance on a publicly available sarcasm dataset using a collaborative gating architecture. Finally, we introduce PodSarc, a large-scale sarcastic speech dataset created through this pipeline. The detection model achieves a 73.63% F1 score, demonstrating the dataset's potential as a benchmark for sarcasm detection research.

cs.CL

Hunting for Neutrino Texture Zeros with Muon and Tau Flavor Violation

We revisit the minimal type II seesaw mechanism generating the Majorana neutrino mass matrix $M^ν$, under the assumption that two entries of $M^ν$ vanish. Such flavor structures are known as two-zero textures. Processes with charged lepton flavor violation (CLFV), absent in the Standard Model (SM), can have sizable rates in this framework and are directly linked to the flavor structure of $M^ν$. For each allowed two-zero texture, we quantify the predicted correlations among various CLFV observables using current neutrino oscillation data and show that they lead to distinctive patterns of CLFV processes that could be discriminated between at running and upcoming experiments. In addition, together with information from colliders, the sensitivity of these correlations to renormalization group (RG) effects could shed light on the potentially ultra-high scale where new dynamics (e.g. some underlying flavor symmetry) give rise to the two-zero texture. Furthermore, we find that certain zero textures, although not third-generation specific, can suppress $μ\to e$ transitions while allowing the rate of the process $τ\to \barμee$ to be within the future experimental sensitivity, even when the RG evolution is taken into account. The lowest possible cut-off scale of the effective theory, constructed by treating the two-zero flavor structure of $M^ν$ as a CLFV spurion, can therefore reach $5-6$ TeV. Our results provide further motivation for searches for $τ$ CLFV at Belle II, as probes of new physics complementary to MEG II and the upcoming Mu3e, COMET, and Mu2e experiments, as well as for collider searches for doubly charged scalar bosons.

hep-ph

SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoning

Multimodal sarcasm detection requires resolving pragmatic incongruity across textual, acoustic, and visual cues through cross-modal reasoning. To enable robust sarcasm reasoning with foundation models, we propose SarcasmMiner, a reinforcement learning based post-training framework that resists hallucination in multimodal reasoning. We reformulate sarcasm detection as structured reasoning and adopt a dual-track distillation strategy: high-quality teacher trajectories initialize the student model, while the full set of trajectories trains a generative reward model (GenRM) to evaluate reasoning quality. The student is optimized with group relative policy optimization (GRPO) using decoupled rewards for accuracy and reasoning quality. On MUStARD++, SarcasmMiner increases F1 from 59.83% (zero-shot), 68.23% (supervised finetuning) to 70.22%. These findings suggest that reasoning-aware reward modeling enhances both performance and multimodal grounding.

cs.MM

TeV-scale scalar leptoquarks motivated by B anomalies improve Yukawa unification in SO(10) GUT

It is common practice to explain deviations between data and Standard-Model (SM) predictions by postulating new particles at the TeV scale ad-hoc. This approach becomes much more convincing, if one successfully embeds the postulated particles into a UV completion which addresses other conceptual or phenomenological shortcomings of the SM. We present a study of an SO(10) grand unified theory which contains scalar leptoquark fields employed to explain the ``flavour anomalies'' in $b\rightarrow s$ and $b\rightarrow c$ decays. We find that the additional degrees of freedom improve the renormalization-group (RG) evolution of the SM parameters. In particular, the light leptoquarks modify the RG evolution of the Yukawa couplings such that successful bottom-tau unification becomes possible in a minimal SO(10) GUT with only a $126$-plet coupling to fermions. If we amend the Yukawa interaction of the minimal one-generation model with a second fermion multiplet and small flavor-violating terms, we find the flavour violation in the leptoquark couplings growing with the RG evolution while it stays small in the Yukawa interaction of the SM Higgs boson. By employing mass splittings among the members of the $126$-plet one can increase the effect and obtain large flavor violation in leptoquark couplings from tiny perturbations at the GUT scale, because the flavour-conserving limit is an unstable initial condition for the RG equations.

hep-ph

Neutrino Masses with Enhanced $B-L$ Symmetry

Assuming all three known neutrinos are Dirac fermions, $U(1)_{B-L}$ can be an exact symmetry. We show that, if the condition of charge quantization is relaxed, the anomaly-free $B-L$ charges of two out of three right-handed neutrinos can be enhanced by arbitrarily large factors, while all other fermions retain their canonical charges. We call this setup as `enhanced $B-L$ symmetry' and promote it to be local. As long as this enhanced $B-L$ gauge symmetry remains unbroken, neutrinos stay chiral and massless at low energies. Nonzero neutrino masses then require sub-eV-scale symmetry breaking order parameters, which we associate with gravity-induced neutrino condensate. If the enhancement is large and the $B-L$ gauge boson $A'$ is lighter than the heaviest neutrino, then the neutrino decay into $A'$ directly constrains the gauge coupling, which can be significantly stronger than the baryon-based fifth-force tests. Through kinetic mixing with the photon, $A'$ can also mediate neutrino-electron and coherent neutrino-nucleus scatterings, leading to possible signatures in neutrino observatories and dark matter detectors.

hep-ph

$B\rightarrow K + invisible$ in a model with axion-like particles

The localized excess of $B\rightarrow K + \textit{invisible}$ events reported by Belle-II is commonly interpreted as a signal for $B\to Ka$, where $a$ is an axion-like particle (ALP). In these proceedings, we summarize two theoretical updates regarding the $b\to sa$ decay amplitude within a minimal UV-complete model for invisible ALPs, namely the DFSZ model. (i) Contributions from certain two-loop diagrams can dominate the one-loop ones and are thus relevant for phenomenology. (ii) Although not unique, a rare feature -- which we refer to as apparent non-decoupling -- emerges in the light effective theory. The renormalizable ALP interactions fail to capture this behavior and therefore are incomplete as low-energy effective descriptions.

hep-ph

Multimodal Negative Learning

Multimodal learning systems often encounter challenges related to modality imbalance, where a dominant modality may overshadow others, thereby hindering the learning of weak modalities. Conventional approaches often force weak modalities to align with dominant ones in "Learning to be (the same)" (Positive Learning), which risks suppressing the unique information inherent in the weak modalities. To address this challenge, we offer a new learning paradigm: "Learning Not to be" (Negative Learning). Instead of enhancing weak modalities' target-class predictions, the dominant modalities dynamically guide the weak modality to suppress non-target classes. This stabilizes the decision space and preserves modality-specific information, allowing weak modalities to preserve unique information without being over-aligned. We proceed to reveal multimodal learning from a robustness perspective and theoretically derive the Multimodal Negative Learning (MNL) framework, which introduces a dynamic guidance mechanism tailored for negative learning. Our method provably tightens the robustness lower bound of multimodal learning by increasing the Unimodal Confidence Margin (UCoM) and reduces the empirical error of weak modalities, particularly under noisy and imbalanced scenarios. Extensive experiments across multiple benchmarks demonstrate the effectiveness and generalizability of our approach against competing methods. The code will be available at https://github.com/BaoquanGong/Multimodal-Negative-Learning.git.

cs.LG

Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding

Sarcasm detection remains a challenge in natural language understanding, as sarcastic intent often relies on subtle cross-modal cues spanning text, speech, and vision. While prior work has primarily focused on textual or visual-textual sarcasm, comprehensive audio-visual-textual sarcasm understanding remains underexplored. In this paper, we systematically evaluate large language models (LLMs) and multimodal LLMs for sarcasm detection on English (MUStARD++) and Chinese (MCSD 1.0) in zero-shot, few-shot, and LoRA fine-tuning settings. In addition to direct classification, we explore models as feature encoders, integrating their representations through a collaborative gating fusion module. Experimental results show that audio-based models achieve the strongest unimodal performance, while text-audio and audio-vision combinations outperform unimodal and trimodal models. Furthermore, MLLMs such as Qwen-Omni show competitive zero-shot and fine-tuned performance. Our findings highlight the potential of MLLMs for cross-lingual, audio-visual-textual sarcasm understanding.

cs.CL

Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects

Sarcasm, a common feature of human communication, poses challenges in interpersonal interactions and human-machine interactions. Linguistic research has highlighted the importance of prosodic cues, such as variations in pitch, speaking rate, and intonation, in conveying sarcastic intent. Although previous work has focused on text-based sarcasm detection, the role of speech data in recognizing sarcasm has been underexplored. Recent advancements in speech technology emphasize the growing importance of leveraging speech data for automatic sarcasm recognition, which can enhance social interactions for individuals with neurodegenerative conditions and improve machine understanding of complex human language use, leading to more nuanced interactions. This systematic review is the first to focus on speech-based sarcasm recognition, charting the evolution from unimodal to multimodal approaches. It covers datasets, feature extraction, and classification methods, and aims to bridge gaps across diverse research domains. The findings include limitations in datasets for sarcasm recognition in speech, the evolution of feature extraction techniques from traditional acoustic features to deep learning-based representations, and the progression of classification methods from unimodal approaches to multimodal fusion techniques. In so doing, we identify the need for greater emphasis on cross-cultural and multilingual sarcasm recognition, as well as the importance of addressing sarcasm as a multimodal phenomenon, rather than a text-based challenge.

cs.CL

Integrating Feedback Loss from Bi-modal Sarcasm Detector for Sarcastic Speech Synthesis

Sarcastic speech synthesis, which involves generating speech that effectively conveys sarcasm, is essential for enhancing natural interactions in applications such as entertainment and human-computer interaction. However, synthesizing sarcastic speech remains a challenge due to the nuanced prosody that characterizes sarcasm, as well as the limited availability of annotated sarcastic speech data. To address these challenges, this study introduces a novel approach that integrates feedback loss from a bi-modal sarcasm detection model into the TTS training process, enhancing the model's ability to capture and convey sarcasm. In addition, by leveraging transfer learning, a speech synthesis model pre-trained on read speech undergoes a two-stage fine-tuning process. First, it is fine-tuned on a diverse dataset encompassing various speech styles, including sarcastic speech. In the second stage, the model is further refined using a dataset focused specifically on sarcastic speech, enhancing its ability to generate sarcasm-aware speech. Objective and subjective evaluations demonstrate that our proposed methods improve the quality, naturalness, and sarcasm-awareness of synthesized speech.

cs.CL

$B\rightarrow K + \text{axion-like particles}$: effective versus UV-complete models and enhanced two-loop contributions

An axion-like particle $a$ (ALP) can explain the excess of $B\rightarrow K + \text{invisible}$ events at Belle-II. However, many analyses of ALP scenarios are over-simplified. We revisit the $B\rightarrow K a$ transition rate in a popular minimal and UV complete model with two Higgs doublets (2HDM) and a complex singlet (DFSZ model). To this end we compare our results with previous studies which derived the $\overline{b}sa$ vertex from the $\overline{b}sA$ vertex, where $A$ is the heavy pseudo-scalar of the 2HDM, in terms of an $a-A$ mixing angle. We find this approach to work only at the leading one-loop order, while it fails at the two-loop level. Furthermore, while an approximate $Z_2$ symmetry suppresses the leading-order amplitude by a factor of $1/\tanβ$, which is the ratio of the two vacuum expectation values of the Higgs doublets, we find the two-loop contribution unsuppressed and phenomenologically relevant for $\tanβ\gtrsim 5$. We determine the allowed parameter space and underline the importance of better searches for $Υ\rightarrow γ+$invisible and for a possible excess in $B\rightarrow Kμ^+μ^-$. We further study the low-energy axion effective theory which leads to a divergent and basis-dependent amplitude. As a conceptual result, we clarify the ambiguities and identify which low-energy framework is consistent with the DFSZ model.

hep-ph

Minimal renormalizable $SO(10)$, spontaneous CP violation, and flavor implications

In these proceedings, we summarize our recent findings on a minimal renormalizable $SO(10)$ grand unified theory. With the assumption of spontaneous $CP$ violation, the low-energy theory becomes a constrained two-Higgs-doublet model, whose mass spectrum has an upper bound of 545 GeV. High-luminosity collider experiments may find its flavor violating signals and reveal hidden mixing parameters, enabling new predictions for proton decay. There is a possibility that a hint of minimal $SO(10)$ can appear soon in near future experiments.

hep-ph

Spontaneous CP Violation and Flavor Changing Neutral Currents in Minimal SO(10)

We explore spontaneous CP violation (SCPV) in the minimal non-supersymmetric SO(10) grand unified theory (GUT), with a scalar sector comprising a CP-even $45_H$, a $126_H$, and a complex $10_H$. All renormalizable couplings are real due to CP symmetry, and the Kobayashi-Maskawa phase arises solely from complex electroweak vacuum expectation values. The model requires an additional Higgs doublet fine-tuned below 500 GeV and constrains new Yukawa couplings, linking certain flavor-violating (FV) processes. Future proton decay observations may reveal correlated FV decay ratios, offering insights into minimal SO(10).

hep-ph

Asymmetric Reinforcing against Multi-modal Representation Bias

The strength of multimodal learning lies in its ability to integrate information from various sources, providing rich and comprehensive insights. However, in real-world scenarios, multi-modal systems often face the challenge of dynamic modality contributions, the dominance of different modalities may change with the environments, leading to suboptimal performance in multimodal learning. Current methods mainly enhance weak modalities to balance multimodal representation bias, which inevitably optimizes from a partialmodality perspective, easily leading to performance descending for dominant modalities. To address this problem, we propose an Asymmetric Reinforcing method against Multimodal representation bias (ARM). Our ARM dynamically reinforces the weak modalities while maintaining the ability to represent dominant modalities through conditional mutual information. Moreover, we provide an in-depth analysis that optimizing certain modalities could cause information loss and prevent leveraging the full advantages of multimodal data. By exploring the dominance and narrowing the contribution gaps between modalities, we have significantly improved the performance of multimodal learning, making notable progress in mitigating imbalanced multimodal learning.

cs.CV

AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bi-modal Data Augmentation

Detecting sarcasm effectively requires a nuanced understanding of context, including vocal tones and facial expressions. The progression towards multimodal computational methods in sarcasm detection, however, faces challenges due to the scarcity of data. To address this, we present AMuSeD (Attentive deep neural network for MUltimodal Sarcasm dEtection incorporating bi-modal Data augmentation). This approach utilizes the Multimodal Sarcasm Detection Dataset (MUStARD) and introduces a two-phase bimodal data augmentation strategy. The first phase involves generating varied text samples through Back Translation from several secondary languages. The second phase involves the refinement of a FastSpeech 2-based speech synthesis system, tailored specifically for sarcasm to retain sarcastic intonations. Alongside a cloud-based Text-to-Speech (TTS) service, this Fine-tuned FastSpeech 2 system produces corresponding audio for the text augmentations. We also investigate various attention mechanisms for effectively merging text and audio data, finding self-attention to be the most efficient for bimodal integration. Our experiments reveal that this combined augmentation and attention approach achieves a significant F1-score of 81.0% in text-audio modalities, surpassing even models that use three modalities from the MUStARD dataset.

cs.CL