SearcharxivSearch

arXiv subjects

Junjie Cao

Publications and source records attributed to Junjie Cao.

At least 19 recordsLinked to original sources

Higgsino Dark Matter Interpretation of the LZ High-Recoil Event in the GNMSSM with TeV-Scale Gauginos

The nuclear recoil event at approximately 248 keV reported by the LUX-ZEPLIN collaboration motivates an investigation of endothermic dark matter scattering. We study this within the General Next-to-Minimal Supersymmetric Standard Model (GNMSSM), with Higgsino-dominated neutralino DM undergoing the $Z$-mediated transition $\widetildeχ_1^0N\to\widetildeχ_2^0N$. In the conventional thermal Higgsino limit of the MSSM, the observed relic abundance selects a mass near 1.1 TeV, while a neutralino splitting of a few hundred keV typically requires gaugino masses of order $10^7$ GeV. In the GNMSSM, Higgsino-Singlino mixing introduces an additional contribution to the splitting that can cancel the gaugino-induced contribution, allowing sub-MeV splitting with multi-TeV gauginos. This mixing also modifies the inelastic scattering coupling and annihilation rates, while coannihilation with sleptons provides freedom in obtaining the observed relic abundance. We present six benchmark points with dark-matter masses of 0.66-1.11 TeV, neutralino splittings of 333-350 keV, and gaugino masses of 2-5 TeV. These points reproduce the observed relic abundance and satisfy direct-detection, Higgs, flavor, and collider constraints. Within the Standard Halo Model and extended-likelihood analysis, all six points yield $Δχ^2<1$ relative to the best fit. Our results show how the GNMSSM can accommodate the LZ high-recoil event without an ultraheavy gaugino sector. A quantitative assessment of solar-capture and neutrino-telescope constraints remains necessary for establishing viability.

hep-ph

Disorder Thresholds and Free Energy of Brownian Directed Polymers with Product and Radial Spatial Correlations

We study a Brownian directed polymer in a centered Gaussian environment that is white in time and colored in space having long-range spatial correlations. For product-type covariances \(Q(x)\asymp\prod_{j=1}^d(1+|x_j|)^{-α_j}\), with \(α_j\in(0,1)\) and \(κ=\sum_jα_j\), we identify the disorder transition at the marginal value \(κ=2\). For \(κ>2\), weak disorder holds at sufficiently small inverse temperature; for \(κ<2\), the quenched free energy $p(β)$ satisfies \(-p(β)\asympβ^{4/(2-κ)}\) as \(β\downarrow0\). For \(κ=2\), strong disorder holds for every \(β>0\), while \(p(β)=0\) for all sufficiently small \(β\), so \(β_c=0<\barβ_c\). We also consider the radial covariance cases, where $ Q(x)\asymp(1+|x|)^{-\vartheta}$, when \(d\ge3,\vartheta=2\) and \(d=2,\vartheta\ge2\), which was left unanswered in Lacoin~\cite{Lacoin2011}. When $d=2$, we get \(\ln(-p(β))\asymp-β^{-2}\) for \(\vartheta>2\) and \(\ln(-p(β))\asymp-β^{-1}\) for \(\vartheta=2\). The proofs consist of replica coupling, Feynman--Kac variational formula, overlap methods, and continuous-space fractional moments with ordered Wiener-chaos changes of measure.

math.PR

Scaling limit of 1+1 dimensional directed polymer with power-law tail and spatial correlated noise

We study a $(1+1)$-dimensional directed polymer in a spatially correlated random environment generated by power-law tail variables: $ω(i,x)=\sum_{y\in\mathbb Z}ψ_{y-x}ξ(i,y), ψ_y\sim λ_r |y|^{-r}, r\in(1/2,1)$, where the variables $ξ(i,y)$ are i.i.d. and have a regularly varying right tail with exponent $α>2$. The spatial covariance of the environment has long-range decay with Hurst parameter $H=\frac32-r\in(1/2,1)$. We identify the limiting fluctuations of the log-partition function in the intermediate disorder regime and show that the critical tail exponent is $α_c=\frac{3}{H}=\frac{6}{3-2r}$. When $α>α_c$, the model has the same scaling limits as the corresponding Gaussian spatially correlated polymer: if $β_NN^{H/2}\toβ\in(0,\infty)$, the centered log-partition function converges to the logarithm of the solution of the stochastic heat equation driven by fractional spatial noise; if $β_NN^{H/2}\to0$, its normalized fluctuation converges to a centered Gaussian law. In the regime $2<α\leα_c$, at the scale $β_NN^{H/2}=βN^{H/2}/l(N^{3/2})$, the log-partition function still satisfies Gaussian fluctuation. The main ingredient is a truncation comparison argument adapted to long-range moving-average environments, together with an invariance principle for polynomial chaos. Due to the non-locality of the environments, we perform a far-near field analysis, as well as multiscale analysis, to prove that the truncated version does not change the log-partition function at the corresponding scales.

math.PR

HeartMuLa: A Family of Open Sourced Music Foundation Models

We present a family of open-source Music Foundation Models designed to advance large-scale music understanding and generation across diverse tasks and modalities. Our framework consists of four major components: (1) HeartCLAP, an audio-text alignment model; (2) HeartTranscriptor, a robust lyric recognition model optimized for real-world music scenarios; and (3) HeartCodec, a low-frame-rate (12.5 Hz) yet high-fidelity music codec tokenizer that captures long-range musical structure while preserving fine-grained acoustic details and enabling efficient autoregressive modeling; (4) HeartMuLa, an LLM-based song generation model capable of synthesizing high-fidelity music under rich, user-controllable conditions (e.g., textual style descriptions, lyrics, and reference audio). In addition, it provides two specialized modes: (i) fine-grained musical attribute control, which allows users to specify the style of different song sections (e.g., intro, verse, chorus) using natural language prompts; and (ii) short, engaging music generation, which is suitable as background music for short videos. Lastly, HeartMuLa improves significantly when scaled to 7B parameters. For the first time, we show that a Suno-level, commercial-grade system can be reproduced using academic-scale data and GPU resources. We expect these foundation models to serve as strong baselines for future research and to facilitate practical applications in multimodal content production.

cs.SD

STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-Identification

Multi-modal object Re-Identification (ReID) aims to exploit complementary information from different modalities to retrieve specific objects. However, existing methods often rely on hard token filtering or simple fusion strategies, which can lead to the loss of discriminative cues and increased background interference. To address these challenges, we propose STMI, a novel multi-modal learning framework consisting of three key components: (1) Segmentation-Guided Feature Modulation (SFM) module leverages SAM-generated masks to enhance foreground representations and suppress background noise through learnable attention modulation; (2) Semantic Token Reallocation (STR) module employs learnable query tokens and an adaptive reallocation mechanism to extract compact and informative representations without discarding any tokens; (3) Cross-Modal Hypergraph Interaction (CHI) module constructs a unified hypergraph across modalities to capture high-order semantic relationships. Extensive experiments on public benchmarks (i.e., RGBNT201, RGBNT100, and MSVR310) demonstrate the effectiveness and robustness of our proposed STMI framework in multi-modal ReID scenarios.

cs.CV

Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents

The paper introduces GUI-Owl-1.5, the latest native GUI agent model that features instruct/thinking variants in multiple sizes (2B/4B/8B/32B/235B) and supports a range of platforms (desktop, mobile, browser, and more) to enable cloud-edge collaboration and real-time interaction. GUI-Owl-1.5 achieves state-of-the-art results on more than 20+ GUI benchmarks on open-source models: (1) on GUI automation tasks, it obtains 56.5 on OSWorld, 71.6 on AndroidWorld, and 48.4 on WebArena; (2) on grounding tasks, it obtains 80.3 on ScreenSpotPro; (3) on tool-calling tasks, it obtains 47.6 on OSWorld-MCP, and 46.8 on MobileWorld; (4) on memory and knowledge tasks, it obtains 75.5 on GUI-Knowledge Bench. GUI-Owl-1.5 incorporates several key innovations: (1) Hybird Data Flywheel: we construct the data pipeline for UI understanding and trajectory generation based on a combination of simulated environments and cloud-based sandbox environments, in order to improve the efficiency and quality of data collection. (2) Unified Enhancement of Agent Capabilities: we use a unified thought-synthesis pipeline to enhance the model's reasoning capabilities, while placing particular emphasis on improving key agent abilities, including Tool/MCP use, memory and multi-agent adaptation; (3) Multi-platform Environment RL Scaling: We propose a new environment RL algorithm, MRPO, to address the challenges of multi-platform conflicts and the low training efficiency of long-horizon tasks. The GUI-Owl-1.5 models are open-sourced, and an online cloud-sandbox demo is available at https://github.com/X-PLUG/MobileAgent.

cs.AI

High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer

Recent progress in diffusion models has significantly advanced the field of human image animation. While existing methods can generate temporally consistent results for short or regular motions, significant challenges remain, particularly in generating long-duration videos. Furthermore, the synthesis of fine-grained facial and hand details remains under-explored, limiting the applicability of current approaches in real-world, high-quality applications. To address these limitations, we propose a diffusion transformer (DiT)-based framework which focuses on generating high-fidelity and long-duration human animation videos. First, we design a set of hybrid implicit guidance signals and a sharpness guidance factor, enabling our framework to additionally incorporate detailed facial and hand features as guidance. Next, we incorporate the time-aware position shift fusion module, modify the input format within the DiT backbone, and refer to this mechanism as the Position Shift Adaptive Module, which enables video generation of arbitrary length. Finally, we introduce a novel data augmentation strategy and a skeleton alignment model to reduce the impact of human shape variations across different identities. Experimental results demonstrate that our method outperforms existing state-of-the-art approaches, achieving superior performance in both high-fidelity and long-duration human image animation.

cs.CV

Natural Realization of Tens-of-GeV Dark Matter in the GNMSSM

This study presents a comparative analysis of the Minimal Supersymmetric Standard Model (MSSM), the $Z_3$-symmetric Next-to-Minimal Supersymmetric Standard Model ($Z_3$-NMSSM), and the General Next-to-Minimal Supersymmetric Standard Model (GNMSSM), incorporating constraints from dark matter (DM) relic density, the LUX-ZEPLIN 2024 experiment (LZ 2024), Higgs data, and the Large Hadron Collider (LHC). The results suggest that, among the three frameworks, only GNMSSM can naturally accommodate for light DM with a mass below $100~{\rm GeV}$. As such, the viable supersymmetry candidate is primarily Singlino-like. One key advantage of the GNMSSM is the effective decoupling between interactions that establish the relic density and those that control direct detection, allowing the model to satisfy all current experimental bounds simultaneously. We further explore two characteristic mass hierarchies in the GNMSSM parameter space, each exhibiting distinct phenomenological behaviors. The first hierarchy, $\tilde{S} < \tilde{B} < \tilde{H}$ (Singlino--Bino--Higgsino), involves a relatively light Bino and allows the Higgsino mass parameter, $μ_{\rm tot}$, to be as low as about $200~{\rm GeV}$, naturally yielding light DM at tens of GeV. The dominant annihilation channels are then $\tildeχ_1^0\tildeχ_1^0 \to A_sA_s$ in the $h_1$ scenario and $\tildeχ_1^0\tildeχ_1^0 \to h_sA_s$ in the $h_2$ scenario, where $h_s$ and $A_s$ denote singlet-dominated CP-even and CP-odd Higgs bosons, respectively. The second hierarchy, $\tilde{S} < \tilde{H} < \tilde{B}$, corresponds to a heavy Bino. In this case, although the DM phenomenology remains qualitatively similar, LHC constraints require $μ_{\rm tot} \gtrsim 900~{\rm GeV}$, implying a significant degree of fine-tuning in reproducing the $Z$-boson mass.

hep-ph

Attractive features of Higgsino Dark Matter in the Next-to-Minimal Supersymmetric Standard Model

In the Higgsino dark matter (DM) scenario of the Minimal Supersymmetric Model (MSSM), the mixing of Gaugino and Higgsino influences the mass splitting between neutralinos predominantly composed of Higgsino and introduces coupling between the DM and Higgs bosons. These effects modify the DM-nucleon scattering cross-section, causing conflicts with the latest direct detection results from LZ experiments for both substantial and minute mixings. Consequently, the experimental measurement of DM relic density necessitates the Higgsino DM mass to be approximately 1.1 TeV. We discovered that in the Higgsino DM scenario of the Next-to-Minimal Supersymmetric Model (NMSSM), the mixing of Higgsino and Singlino introduces analogous effects, with a crucial distinction being that the current LZ experiment permits significant mixing between Singlino and Higgsino. This pronounced mixing effect effectively attenuates the interactions between Higgsino-dominated neutralinos and standard model particles, enabling DM masses exceeding roughly 670 GeV to achieve the correct relic abundance. Through analytical formulas and numerical results, we elucidated these characteristics which were not observed before. Our research reveals that in the NMSSM, when comprehensively examining the mixing effects of Higgsino, Gaugino, and Singlino, the properties of Higgsino DM become markedly more intricate compared to the MSSM predictions.

hep-ph

LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation

Legal consultation is essential for safeguarding individual rights and ensuring access to justice, yet remains costly and inaccessible to many individuals due to the shortage of professionals. While recent advances in Large Language Models (LLMs) offer a promising path toward scalable, low-cost legal assistance, current systems fall short in handling the interactive and knowledge-intensive nature of real-world consultations. To address these challenges, we introduce LeCoDe, a real-world multi-turn benchmark dataset comprising 3,696 legal consultation dialogues with 110,008 dialogue turns, designed to evaluate and improve LLMs' legal consultation capability. With LeCoDe, we innovatively collect live-streamed consultations from short-video platforms, providing authentic multi-turn legal consultation dialogues. The rigorous annotation by legal experts further enhances the dataset with professional insights and expertise. Furthermore, we propose a comprehensive evaluation framework that assesses LLMs' consultation capabilities in terms of (1) clarification capability and (2) professional advice quality. This unified framework incorporates 12 metrics across two dimensions. Through extensive experiments on various general and domain-specific LLMs, our results reveal significant challenges in this task, with even state-of-the-art models like GPT-4 achieving only 39.8% recall for clarification and 59% overall score for advice quality, highlighting the complexity of professional consultation scenarios. Based on these findings, we further explore several strategies to enhance LLMs' legal consultation abilities. Our benchmark contributes to advancing research in legal domain dialogue systems, particularly in simulating more real-world user-expert interactions.

cs.CL

New Physics Search at the CEPC: a General Perspective

The Circular Electron-Positron Collider (CEPC), a proposed next-generation Higgs factory, provides new opportunities to explore physics beyond the Standard Model (SM). With its clean electron-positron collision environment and the ability to collect large samples of Higgs, W, and Z bosons, the CEPC enables precision measurements and searches for new physics. This white paper outlines the CEPC's discovery potential, including studies of exotic decays of the Higgs, Z, and top quarks, dark matter and dark sector phenomena, long-lived particles, supersymmetry, and neutrino-related signatures. Advanced detector technologies and reconstruction techniques, such as one-to-one correspondence reconstruction and jet origin identification, significantly improve sensitivity to rare and weakly interacting processes. The CEPC is particularly well suited to probe the electroweak phase transition and test models of electroweak baryogenesis and dark sector interactions. In addition, global fit analyses highlight the CEPC's complementary role in constraining a wide range of new physics scenarios. These features position the CEPC as a powerful tool for exploring the next frontier in fundamental particle physics in the post-Higgs discovery era.

hep-ex

Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling

Existing Large Language Model (LLM) based autoregressive (AR) text-to-speech (TTS) systems, while achieving state-of-the-art quality, still face critical challenges. The foundation of this LLM-based paradigm is the discretization of the continuous speech waveform into a sequence of discrete tokens by neural audio codec. However, single codebook modeling is well suited to text LLMs, but suffers from significant information loss; hierarchical acoustic tokens, typically generated via Residual Vector Quantization (RVQ), often lack explicit semantic structure, placing a heavy learning burden on the model. Furthermore, the autoregressive process is inherently susceptible to error accumulation, which can degrade generation stability. To address these limitations, we propose CaT-TTS, a novel framework for robust and semantically-grounded zero-shot synthesis. First, we introduce S3Codec, a split RVQ codec that injects explicit linguistic features into its primary codebook via semantic distillation from a state-of-the-art ASR model, providing a structured representation that simplifies the learning task. Second, we propose an ``Understand-then-Generate'' dual-Transformer architecture that decouples comprehension from rendering. An initial ``Understanding'' Transformer models the cross-modal relationship between text and the audio's semantic tokens to form a high-level utterance plan. A subsequent ``Generation'' Transformer then executes this plan, autoregressively synthesizing hierarchical acoustic tokens. Finally, to enhance generation stability, we introduce Masked Audio Parallel Inference (MAPI), a nearly parameter-free inference strategy that dynamically guides the decoding process to mitigate local errors.

cs.SD

MBCodec:Thorough disentangle for high-fidelity audio compression

High-fidelity neural audio codecs in Text-to-speech (TTS) aim to compress speech signals into discrete representations for faithful reconstruction. However, prior approaches faced challenges in effectively disentangling acoustic and semantic information within tokens, leading to a lack of fine-grained details in synthesized speech. In this study, we propose MBCodec, a novel multi-codebook audio codec based on Residual Vector Quantization (RVQ) that learns a hierarchically structured representation. MBCodec leverages self-supervised semantic tokenization and audio subband features from the raw signals to construct a functionally-disentangled latent space. In order to encourage comprehensive learning across various layers of the codec embedding space, we introduce adaptive dropout depths to differentially train codebooks across layers, and employ a multi-channel pseudo-quadrature mirror filter (PQMF) during training. By thoroughly decoupling semantic and acoustic features, our method not only achieves near-lossless speech reconstruction but also enables a remarkable 170x compression of 24 kHz audio, resulting in a low bit rate of just 2.2 kbps. Experimental evaluations confirm its consistent and substantial outperformance of baselines across all evaluations.

cs.SD

Mobile-Agent-v3: Fundamental Agents for GUI Automation

This paper introduces GUI-Owl, a foundational GUI agent model that achieves state-of-the-art performance among open-source end-to-end models on ten GUI benchmarks across desktop and mobile environments, covering grounding, question answering, planning, decision-making, and procedural knowledge. GUI-Owl-7B achieves 66.4 on AndroidWorld and 29.4 on OSWorld. Building on this, we propose Mobile-Agent-v3, a general-purpose GUI agent framework that further improves performance to 73.3 on AndroidWorld and 37.7 on OSWorld, setting a new state-of-the-art for open-source GUI agent frameworks. GUI-Owl incorporates three key innovations: (1) Large-scale Environment Infrastructure: a cloud-based virtual environment spanning Android, Ubuntu, macOS, and Windows, enabling our Self-Evolving GUI Trajectory Production framework. This generates high-quality interaction data via automated query generation and correctness validation, leveraging GUI-Owl to refine trajectories iteratively, forming a self-improving loop. It supports diverse data pipelines and reduces manual annotation. (2) Diverse Foundational Agent Capabilities: by integrating UI grounding, planning, action semantics, and reasoning patterns, GUI-Owl supports end-to-end decision-making and can act as a modular component in multi-agent systems. (3) Scalable Environment RL: we develop a scalable reinforcement learning framework with fully asynchronous training for real-world alignment. We also introduce Trajectory-aware Relative Policy Optimization (TRPO) for online RL, achieving 34.9 on OSWorld. GUI-Owl and Mobile-Agent-v3 are open-sourced at https://github.com/X-PLUG/MobileAgent.

cs.AI

Adaptive Duration Model for Text Speech Alignment

Speech-to-text alignment is a critical component of neural text to speech (TTS) models. Autoregressive TTS models typically use an attention mechanism to learn these alignments on-line, while non-autoregressive end to end TTS models rely on durations extracted from external sources. In this paper, we propose a novel duration prediction framework that can give promising phoneme-level duration distribution with given text. In our experiments, the proposed duration model has more precise prediction and adaptation ability to conditions, compared to previous baseline models. Specifically, it makes a considerable improvement on phoneme-level alignment accuracy and makes the performance of zero-shot TTS models more robust to the mismatch between prompt audio and input audio.

cs.SD

VideoGuard: Protecting Video Content from Unauthorized Editing

With the rapid development of generative technology, current generative models can generate high-fidelity digital content and edit it in a controlled manner. However, there is a risk that malicious individuals might misuse these capabilities for misleading activities. Although existing research has attempted to shield photographic images from being manipulated by generative models, there remains a significant disparity in the protection offered to video content editing. To bridge the gap, we propose a protection method named VideoGuard, which can effectively protect videos from unauthorized malicious editing. This protection is achieved through the subtle introduction of nearly unnoticeable perturbations that interfere with the functioning of the intended generative diffusion models. Due to the redundancy between video frames, and inter-frame attention mechanism in video diffusion models, simply applying image-based protection methods separately to every video frame can not shield video from unauthorized editing. To tackle the above challenge, we adopt joint frame optimization, treating all video frames as an optimization entity. Furthermore, we extract video motion information and fuse it into optimization objectives. Thus, these alterations can effectively force the models to produce outputs that are implausible and inconsistent. We provide a pipeline to optimize this perturbation. Finally, we use both objective metrics and subjective metrics to demonstrate the efficacy of our method, and the results show that the protection performance of VideoGuard is superior to all the baseline methods.

cs.CV

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations

Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, which suffer from significant information loss. Even with post-hoc refinement techniques such as flow matching, these methods fail to recover fine-grained details (e.g., prosodic nuances, speaker-specific timbres), especially in challenging scenarios like singing voice or music synthesis. We propose QTTS, a novel TTS framework built upon our new audio codec, QDAC. The core innovation of QDAC lies in its end-to-end training of an ASR-based auto-regressive network with a GAN, which achieves superior semantic feature disentanglement for scalable, near-lossless compression. QTTS models these discrete codes using two innovative strategies: the Hierarchical Parallel architecture, which uses a dual-AR structure to model inter-codebook dependencies for higher-quality synthesis, and the Delay Multihead approach, which employs parallelized prediction with a fixed delay to accelerate inference speed. Our experiments demonstrate that the proposed framework achieves higher synthesis quality and better preserves expressive content compared to baseline. This suggests that scaling up compression via multi-codebook modeling is a promising direction for high-fidelity, general-purpose speech and audio generation.

cs.SD

Towards Stepwise Domain Knowledge-Driven Reasoning Optimization and Reflection Improvement

Recently, stepwise supervision on Chain of Thoughts (CoTs) presents an enhancement on the logical reasoning tasks such as coding and math, with the help of Monte Carlo Tree Search (MCTS). However, its contribution to tasks requiring domain-specific expertise and knowledge remains unexplored. Motivated by the interest, we identify several potential challenges of vanilla MCTS within this context, and propose the framework of Stepwise Domain Knowledge-Driven Reasoning Optimization, employing the MCTS algorithm to develop step-level supervision for problems that require essential comprehension, reasoning, and specialized knowledge. Additionally, we also introduce the Preference Optimization towards Reflection Paths, which iteratively learns self-reflection on the reasoning thoughts from better perspectives. We have conducted extensive experiments to evaluate the advantage of the methodologies. Empirical results demonstrate the effectiveness on various legal-domain problems. We also report a diverse set of valuable findings, hoping to encourage the enthusiasm to the research of domain-specific LLMs and MCTS.

cs.AI