SearcharxivSearch

arXiv subjects

Yang Ma

Publications and source records attributed to Yang Ma.

At least 19 recordsLinked to original sources

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation

Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external APIs, including adaptation to platform-specific data distributions, policy-specific objectives, and product-level safety constraints. As a result, platforms must undertake internal model development, naturally turning to shared public research for guidance. However, existing multimodal foundation-model studies primarily report architectures, training recipes, data scaling strategies, and benchmark results, but provide less systematic guidance on how failures should be localized and translated into targeted model-development interventions. Interventions are essential because deployment failures are rarely self-explanatory. Similar failures can originate from different causes. Without targeted interventions, improvement reduces to heuristic trial-and-error, where benchmark improvements are weakly attributable, and failures are difficult to trace to their underlying causes. To address this gap, we present a diagnostic methodology for industry-scale Audio-Visual-Language Models AVLM development. The methodology maps model failures into a taxonomy of observable failure signatures and links each class of failure to an intervention space. We instantiate this methodology across the development and alignment lifecycle of an AVLM foundation model for a large-scale video and live-streaming platform. The resulting system supports over 100 regions and is designed for noisy, ambiguous, and highly diverse content drawn from global platform traffic.

cs.LG

Probing $\tau$ lepton dipole moments at future Lepton Colliders

The electric and magnetic dipole moments of the electron and of the muon provide stringent tests of the Standard Model and sensitive probes of new physics. By contrast, the corresponding dipole moments of the $\tau$ lepton remain weakly constrained. This study explores the potential of future lepton colliders, focusing on the $e^+e^-$ Future Circular Collider and a multi-TeV muon collider, to probe $\tau$ dipole moments. We consider multiple channels, including $\ell^+\ell^- \to \tau^+\tau^-$ ($\ell=e,\mu$), associated Higgs production $\mu^+\mu^- \to \tau^+\tau^- H$, radiative Higgs decays $H \to \tau^+\tau^-\gamma$, and vector-boson scattering $\ell^+\ell^- \to \ell^+\ell^-\tau^+\tau^-$ and $\mu^+\mu^- \to \bar\nu\nu\tau^+\tau^-$. Our results show that these facilities are highly complementary and can extend existing bounds by several orders of magnitude.

hep-ph

Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization

Reinforcement Learning (RL) has proven highly effective in addressing complex control and decision-making tasks. However, in most traditional RL algorithms, the policy is typically parameterized as a diagonal Gaussian distribution, which constrains the policy from capturing multimodal distributions, making it difficult to cover the full range of optimal solutions in multi-solution problems, and the return is reduced to a mean value, losing its multimodal nature and thus providing insufficient guidance for policy updates. In response to these problems, we propose a RL algorithm termed flow-based policy with distributional RL (FP-DRL). This algorithm models the policy using flow matching, which offers both computational efficiency and the capacity to fit complex distributions. Additionally, it employs a distributional RL approach to model and optimize the entire return distribution, thereby more effectively guiding multimodal policy updates and improving agent performance. Experimental trails on MuJoCo benchmarks demonstrate that the FP-DRL algorithm achieves state-of-the-art (SOTA) performance in most MuJoCo control tasks while exhibiting superior representation capability of the flow policy.

cs.LG

Chasing the two-Higgs-doublet model via electroweak corrections at $e^+e^-$ colliders

We present a comprehensive study of Higgs boson production associated with a neutrino pair at $e^+e^-$ colliders ($e^+ e^- \to h \, \nu \bar{\nu}$) at next-to-leading-order accuracy in both the Standard Model and the two-Higgs-doublet model. We show that these new physics effects will be observable in total and differential cross sections when compared with theoretical predictions that include electroweak corrections, even in the Higgs alignment limit. This highlights the potential of precision studies at future $e^+e^-$ colliders for searching new physics.

hep-ph

Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning

Embodied AI aims to develop intelligent systems with physical forms capable of perceiving, decision-making, acting, and learning in real-world environments, providing a promising way to Artificial General Intelligence (AGI). Despite decades of explorations, it remains challenging for embodied agents to achieve human-level intelligence for general-purpose tasks in open dynamic environments. Recent breakthroughs in large models have revolutionized embodied AI by enhancing perception, interaction, planning and learning. In this article, we provide a comprehensive survey on large model empowered embodied AI, focusing on autonomous decision-making and embodied learning. We investigate both hierarchical and end-to-end decision-making paradigms, detailing how large models enhance high-level planning, low-level execution, and feedback for hierarchical decision-making, and how large models enhance Vision-Language-Action (VLA) models for end-to-end decision making. For embodied learning, we introduce mainstream learning methodologies, elaborating on how large models enhance imitation learning and reinforcement learning in-depth. For the first time, we integrate world models into the survey of embodied AI, presenting their design methods and critical roles in enhancing decision-making and learning. Though solid advances have been achieved, challenges still exist, which are discussed at the end of this survey, potentially as the further research directions.

cs.RO

Non-Standard Neutrino Interactions at a Muon Collider Neutrino Detector

In addition to their broad physics reach enabled by their high energies and precision, future multi-TeV muon colliders will also be the world's most intense sources of neutrinos. This offers the opportunity to search for new non-standard neutrino interactions, possible by installing a dedicated forward neutrino detector in the straight sections of the collision ring, which is then used to measure reactions initiated by neutrinos from the decaying beam muons. In this paper, we show that these searches can exceed current and upcoming bounds on non-standard neutrino interactions from low-energy precision experiments and the LHC. This is achieved by the large flux of high-energetic neutrinos, the precise knowledge of the neutrino flavor composition on each side of the interaction point and the chirality of the neutrinos. We further discuss the technical requirements of the proposed forward neutrino detector, \FASERmuC, to maximally exploit this physics potential.

hep-ph

Light Axion-Like Particles at Future Lepton Colliders

Axion-like particles (ALPs) are well-motivated extensions of the Standard Model (SM) that appear in many new physics scenarios, with masses spanning a broad range. In this work, we systematically study the production and detection prospects of light ALPs at future lepton colliders, including electron-positron and multi-TeV muon colliders. At lepton colliders, light ALPs can be produced in association with a photon or a $Z$ boson. For very light ALPs ($m_a < 1$ MeV), the ALPs are typically long-lived and escape detection, leading to a mono-$V$ ($V = \gamma, Z$) signature. In the long-lived limit, we find that the mono-photon channel at the Tera-$Z$ stage of future electron-positron colliders provides the strongest constraints on ALP couplings to SM gauge bosons, $g_{aVV}$, thanks to the high luminosity, low background, and resonant enhancement from on-shell $Z$ bosons. At higher energies, the mono-photon cross section becomes nearly energy-independent, and the sensitivity is governed by luminosity and background. At multi-TeV muon colliders, the mono-$Z$ channel can yield complementary constraints. For heavier ALPs ($m_a > 100$ MeV) that decay promptly, mono-$V$ signatures are no longer valid. In this case, ALPs can be probed via non-resonant vector boson scattering (VBS) processes, where the ALP is exchanged off-shell, leading to kinematic deviations from SM expectations. We analyze constraints from both light-by-light scattering and electroweak VBS, the latter only accessible at TeV-scale colliders. While generally weaker, these constraints are robust and model-independent. Our combined analysis shows that mono-$V$ and non-resonant VBS channels provide powerful and complementary probes of ALP-gauge boson interactions.

hep-ph

Distributed Resource Block Allocation for Wideband Cell-free System

This paper studies distributed resource block (RB) allocation in wideband orthogonal frequency-division multiplexing (OFDM) cell-free systems. We propose a novel distributed sequential algorithm and its two variants, which optimize RB allocation based on the information obtained through over-the-air (OTA) transmissions between access points (APs) and user equipments, enabling local decision updates at each AP. To reduce the overhead of OTA transmission, we further develop a distributed deep learning (DL)-based method to learn the RB allocation policy. Simulation results demonstrate that the proposed distributed algorithms perform close to the centralized algorithm, while the DL-based method outperforms existing baseline methods.

eess.SP

RA-BLIP: Multimodal Adaptive Retrieval-Augmented Bootstrapping Language-Image Pre-training

Multimodal Large Language Models (MLLMs) have recently received substantial interest, which shows their emerging potential as general-purpose models for various vision-language tasks. MLLMs involve significant external knowledge within their parameters; however, it is challenging to continually update these models with the latest knowledge, which involves huge computational costs and poor interpretability. Retrieval augmentation techniques have proven to be effective plugins for both LLMs and MLLMs. In this study, we propose multimodal adaptive Retrieval-Augmented Bootstrapping Language-Image Pre-training (RA-BLIP), a novel retrieval-augmented framework for various MLLMs. Considering the redundant information within vision modality, we first leverage the question to instruct the extraction of visual information through interactions with one set of learnable queries, minimizing irrelevant interference during retrieval and generation. Besides, we introduce a pre-trained multimodal adaptive fusion module to achieve question text-to-multimodal retrieval and integration of multimodal knowledge by projecting visual and language modalities into a unified semantic space. Furthermore, we present an Adaptive Selection Knowledge Generation (ASKG) strategy to train the generator to autonomously discern the relevance of retrieved knowledge, which realizes excellent denoising performance. Extensive experiments on open multimodal question-answering datasets demonstrate that RA-BLIP achieves significant performance and surpasses the state-of-the-art retrieval-augmented models.

cs.MM

Higgs-muon interactions at a multi-TeV muon collider

We establish a simple yet general parameterization of Higgs-muon interactions within the effective field theory frameworks, including both the Higgs Effective Field Theory (HEFT) and the Standard Model Effective Field Theory (SMEFT). We investigate the potential of a muon collider, operating at center-of-mass energies of 3 and 10 TeV, to probe Higgs-muon interactions. All possible processes involving the direct production of multiple electroweak bosons ($W$, $Z$, and $H$) with up to five final-state particles are considered. Our findings indicate that a muon collider can achieve greater sensitivity than the high-luminosity LHC, especially considering the independence of the Higgs decay branching fraction to muons. Notably, a 10 TeV muon collider offers exceptional sensitivity to muon-Higgs interactions, surpassing the 3 TeV option. In particular, searches based on multi-Higgs production prove highly effective for probing these couplings.

hep-ph

EW corrections and Heavy Boson Radiation at a high-energy muon collider

In this work we investigate several phenomenological and technical aspects related to electroweak (EW) corrections at a high-energy muon collider, focusing on direct production processes (no VBF configurations). We study in detail the accuracy of the Sudakov approximation, in particular the Denner-Pozzorini algorithm, comparing it with exact calculations at NLO EW accuracy. We also assess the relevance of resumming EW Sudakov logarithms (EWSL) at 3 and 10 TeV collisions. Furthermore, we scrutinise the impact of additional Heavy Boson Radiation (HBR), namely the weak emission of $W, Z$, and Higgs bosons in inclusive and semi-inclusive configurations. All results are obtained via the fully automated and publicly available code MadGraph5_aMC@NLO.

hep-ph

A new probe of dark matter-baryon interactions in compact stellar systems

We investigate the astrophysical consequences of an attractive long-range interaction between dark matter and baryonic matter. Our study highlights the role of this interaction in inducing dynamical friction between dark matter and stars, which can significantly influence the evolution of compact stellar systems. Using the star cluster in Eridanus II as a case study, we derive a new stringent upper bound on the interaction strength $\tilde{\alpha}\leq 314.5$ for the interaction range $\lambda = 1$ pc. This constraint is independent of the dark matter mass and can improve the existing model-independent limits on $\tilde{\alpha}$ by a few orders of magnitude. Furthermore, we observe that the constraint is insensitive to the mass of the stellar system and the dark matter density in the stellar system as long as the system is dark matter dominated. This new approach can be applied to many other stellar systems, and we obtain comparable constraints from compact stellar halos observed in ultrafaint dwarf galaxies.

hep-ph

Symmetry Awareness Encoded Deep Learning Framework for Brain Imaging Analysis

The heterogeneity of neurological conditions, ranging from structural anomalies to functional impairments, presents a significant challenge in medical imaging analysis tasks. Moreover, the limited availability of well-annotated datasets constrains the development of robust analysis models. Against this backdrop, this study introduces a novel approach leveraging the inherent anatomical symmetrical features of the human brain to enhance the subsequent detection and segmentation analysis for brain diseases. A novel Symmetry-Aware Cross-Attention (SACA) module is proposed to encode symmetrical features of left and right hemispheres, and a proxy task to detect symmetrical features as the Symmetry-Aware Head (SAH) is proposed, which guides the pretraining of the whole network on a vast 3D brain imaging dataset comprising both healthy and diseased brain images across various MRI and CT. Through meticulous experimentation on downstream tasks, including both classification and segmentation for brain diseases, our model demonstrates superior performance over state-of-the-art methodologies, particularly highlighting the significance of symmetry-aware learning. Our findings advocate for the effectiveness of incorporating symmetry awareness into pretraining and set a new benchmark for medical imaging analysis, promising significant strides toward accurate and efficient diagnostic processes. Code is available at https://github.com/bitMyron/sa-swin.

eess.IV

Activation Map-based Vector Quantization for 360-degree Image Semantic Communication

In virtual reality (VR) applications, 360-degree images play a pivotal role in crafting immersive experiences and offering panoramic views, thus improving user Quality of Experience (QoE). However, the voluminous data generated by 360-degree images poses challenges in network storage and bandwidth. To address these challenges, we propose a novel Activation Map-based Vector Quantization (AM-VQ) framework, which is designed to reduce communication overhead for wireless transmission. The proposed AM-VQ scheme uses the Deep Neural Networks (DNNs) with vector quantization (VQ) to extract and compress semantic features. Particularly, the AM-VQ framework utilizes activation map to adaptively quantize semantic features, thus reducing data distortion caused by quantization operation. To further enhance the reconstruction quality of the 360-degree image, adversarial training with a Generative Adversarial Networks (GANs) discriminator is incorporated. Numerical results show that our proposed AM-VQ scheme achieves better performance than the existing Deep Learning (DL) based coding and the traditional coding schemes under the same transmission symbols.

eess.IV

Probing Higgs-muon interactions at a multi-TeV muon collider

We study the capabilities of a muon collider, at 3 and 10 TeV center-of-mass energy, of probing the interactions of the Higgs boson with the muon. We consider all the possible processes involving the direct production of EW bosons ($W,Z$ and $H$) with up to five particles in the final state. We study these processes both in the HEFT and SMEFT frameworks, assuming that the dominant BSM effects originate from the muon Yukawa sector. Our study shows that a Muon Collider has sensitivity beyond the LHC, as it not only relies on the Higgs-decay branching fraction to muons. A 10 TeV muon collider provides a unique sensitivity on muon and (multi-) Higgs interactions, significantly better than the 3 TeV option. We find searches based purely on multi-Higgs production to be particularly effective in probing these couplings.

hep-ph

Precision test of the muon-Higgs coupling at a high-energy muon collider

We explore the sensitivity of directly testing the muon-Higgs coupling at a high-energy muon collider. This is strongly motivated if there exists new physics that is not aligned with the Standard Model Yukawa interactions which are responsible for the fermion mass generation. We illustrate a few such examples for physics beyond the Standard Model. With the accidentally small value of the muon Yukawa coupling and its subtle role in the high-energy production of multiple (vector and Higgs) bosons, we show that it is possible to measure the muon-Higgs coupling to an accuracy of ten percent for a 10 TeV muon collider and a few percent for a 30 TeV machine by utilizing the three boson production, potentially sensitive to a new physics scale about $\Lambda \sim$ 30-100 TeV.

hep-ph

Higgs decay to charmonia and the charm-quark Yukawa coupling

After the great triumph of the Higgs discovery in 2012, the next target at the energy frontier will be to study the Higgs properties and to search for the next scale beyond the SM. Experimentally, the $H\to c \bar{c}$ channel would be extremely difficult to dig out because of both the weak Yukawa coupling and the daunting SM di-jet background. We propose to test the charm-quark Yukawa coupling at the LHC and future hadron colliders with the Higgs boson decay to $J/\psi$ via the charm-quark fragmentation. Using the non-relativistic quantum chromodynamics (NRQCD), we study the charmonia production via the Higgs boson decay channel $ H \to c \ \bar{c} + J/\psi $(or $ \eta_c $), where both the color-singlet and color-octet contributions are considered. Our result opens another door to improve determinations at the LHC of the Higgs Yukawa couplings: the final state from this decay mode is quite distinctive with $J/\psi\to e^+e^-,\, \mu^+\mu^-$ and the branching fraction is enhanced by the charm-quark fragmentation mechanism.

hep-ph

Can the three new states around 2.2 GeV assign to $\omega(3D)$

Recently, the BESIII Collaboration reported three resonances: $X(2232)$ with $M = 2232 \pm 19 \pm 27$ MeV and $\Gamma = 93 \pm 53 \pm 20$ MeV, $X(2200)$ whose mass $M = 2200 \pm 11 \pm 17$ MeV and width $\Gamma = 74 \pm 20 \pm 24$ MeV as well as $X(2222)$ which has mass of $2222 \pm 7 \pm 2$ MeV and the width of $59 \pm 30 \pm 6$ MeV. The mass spectrum of $\omega$ meson family is studied utilizing the modified Godfrey-Isgur model, and the two-body strong decays of $X(2232)$, $X(2200)$ and $X(2222)$ within two different approaches of the $^3P_0$ model. We find that the newly discovered states $X(2232)$, $X(2200)$ and $X(2222)$ may be the same and are most likely to be the $\omega(3D)$ state. The discovery could be useful in establishing entire $\omega$ mesons.

hep-ph