SearcharxivSearch

arXiv subjects

Zhen Ouyang

Publications and source records attributed to Zhen Ouyang.

8 recordsLinked to original sources

RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs

Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alignment interacts with routed expert specialization remains underexplored. A common intuition is that safety behavior may be controlled by routing harmful requests to distinct refusal-oriented experts. In this work, we provide empirical evidence for a different picture: routing patterns in aligned MoE LLMs are largely topic-driven, while safety behavior can be altered with little change to the model's intrinsic routing path. Motivated by this observation, we present RASET (Router-Agnostic Safety-Critical Expert Tuning), a red-teaming framework that probes safety enforcement that is localized in a small subset of experts while preserving the model's intrinsic routing behavior. RASET identifies safety-critical experts via a contrastive routing-sensitivity criterion and applies parameter-efficient tuning only to the selected experts, minimizing semantic disruption relative to router-steering interventions. Across five open-weight MoE backbones, RASET achieves high-quality safety-bypass yield (50.5% average $ASR_{hq}$, +37.6 points over the strongest baseline). These results reveal a distinct MoE safety risk, highlighting the need for expert-aware alignment mechanisms.

cs.CL

RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models

Representation engineering offers a lightweight means of controlling language-model behavior by modifying intermediate hidden states, but its direct application to Mixture-of-Experts (MoE) models introduces a structural mismatch. We first verify this failure mode through a series of empirical studies and find that preserving clean routing substantially recovers steering performance and that routing is more sensitive to semantic content than to behavioral changes under controlled content. Motivated by these findings, we introduce RARE, a router-agnostic representation engineering framework for MoE language models. RARE projects arbitrary behavioral perturbations onto the null space of the router matrix, thereby removing router-visible components, and further corrects routing drift propagated to selected downstream layers. To decide the best perturbation estimator in this framework, we evaluate five estimators on six heterogeneous open-weight MoE models across three steering scenarios: harmfulness, truthfulness, and factual editing. On harmfulness steering, RARE reaches an average attack success rate of 53.3% while retaining 67.8% MMLU accuracy, yielding a stronger aggregate effectiveness--utility trade-off than baselines. It further improves average TruthfulQA MC1 accuracy from 41.0% to 58.6% and CounterFact efficacy from 16.8% to 96.3%. These results support routing consistency as an important architectural consideration for adapting representation engineering to MoE models.

cs.CL

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

Agent Skills package reusable natural language procedures with executable resources, enabling software agents to acquire task specific capabilities without model adaptation. Automatically generating such Skills can improve task performance, yet evaluating a candidate solely from its artifact or final task outcome leaves unresolved which actions the equipped agent will perform and which side effects those actions will produce. We present TRUSS, an evidence guided framework for generating functionally effective and safety reliable Agent Skills. TRUSS first inspects functional claims against source and domain evidence while evaluating the complete artifact under nine predefined safety properties. Candidates admitted by this static gate are loaded by a shadow agent inside a Controllable Execution Environment, where brokered tools expose requested actions to policy enforcement and record their results as provenance preserving execution traces. Functional failures and property violations are linked back to the responsible Skill content and used to guide iterative refinement. We evaluate TRUSS on 168 SkillInject artifacts, 155 SkillSafetyBench cases, and all 187 tasks in SkillGenBench. TRUSS achieves 100.00\% precision and recall in vulnerability detection. Repair reduces attack success from 38.71\% to 19.35\% with GPT 5.5 and from 46.45\% to 29.68\% with GPT 5.4, with zero attack regression. For Skill generation, TRUSS raises task effectiveness from 17.11\% without Skills to 52.94\%, while increasing the benchmark Security rate from 50.80\% to 100.00\%. These results show that execution evidence can expose behavioral failures missed by artifact inspection and can guide Skill generation toward jointly verified functional and safety outcomes.

cs.AI

Zenith: Scaling up Ranking Models for Billion-scale Livestreaming Recommendation

Accurately capturing feature interactions is essential in recommender systems, and recent trends show that scaling up model capacity could be a key driver for next-level predictive performance. While prior work has explored various model architectures to capture multi-granularity feature interactions, relatively little attention has been paid to efficient feature handling and scaling model capacity without incurring excessive inference latency. In this paper, we address this by presenting Zenith, a scalable and efficient ranking architecture that learns complex feature interactions with minimal runtime overhead. Zenith is designed to handle a few high-dimensional Prime Tokens with Token Fusion and Token Boost modules, which exhibits superior scaling laws compared to other state-of-the-art ranking methods, thanks to its improved token heterogeneity. Its real-world effectiveness is demonstrated by deploying the architecture to TikTok Live, a leading online livestreaming platform that attracts billions of users globally. Our A/B test shows that Zenith achieves +1.05%/-1.10% in online CTR AUC and Logloss, and realizes +9.93% gains in Quality Watch Session / User and +8.11% in Quality Watch Duration / User.

cs.LG

Role of the hidden charm $N^*_{c\bar{c}}(4261)$ resonance in the $π^- p \to J/ψn$ reaction

We employ an effective Lagrangian approach and isobar model to investigate the role of the $N^*_{c\bar{c}}(4261)$ resonance with hidden charm in the $π^- p \to J/ψn$ reaction. The total and differential cross sections of this reaction are predicted by including contributions from both $N^*_{c\bar{c}}(4261)$ and nucleon pole. It is found that the maximal value of the total cross section can exceed 11 $μb$ and a clear $N^*_{c\bar{c}}(4261)$ peak is visible there, well distinguished from background. As center-of-mass energy increases to about 4.8 GeV, the background makes the differential cross section for backward angles rather different. Our theoretical results would provide valuable information for looking for the $N^*_{c\bar{c}}(4261)$ resonance in future experiments.

hep-ph

Production of charmed baryon $Λ_c(2940)^+$ at PANDA

In this work we evaluate the production rate of the charmed baryon $Λ_c(2940)^+$ at PANDA. For possible assignments of $Λ_c(2940)^+$: $J^P=1/2^\pm$, $3/2^\pm$ and $5/2^\pm$, the total cross section of $p\bar{p}\to \barΛ_c Λ_c(2940)^+$ is estimated, which may exceed 1 nb. With the designed luminosity ($2\times10^{-32}$cm$^{-2}$/s) of PANDA, our estimate indicates that ten thousand events per day if $Λ_c(2940)^+$ is of $J^P=1/2^+$ or $10^8$ per day if it is of $J^P=5/2^+$ can be expected. Those values actually set the lower and upper limits of the $Λ_c(2940)^+$ production. In addition, we present the Dalitz plot and carry out a rough background analysis of the $Λ_c(2940)^+$ production in the $p\bar{p}\to D^0 p\barΛ_c$ and $p\bar{p}\to Σ_c^{0,++}π^{+,-}\barΛ_c$ processes, which would provide valuable information for accurate determination of the $Λ_c(2940)^+$ identity.

hep-ph

Proposal for Studying $N^*$ Resonances with $\bar{p}p \to \bar{p}n π^+ $ Reaction

A theoretical study of $\bar{p}p \to \bar{p}n π^+ $ reaction for anti-proton beam energy from 1 to 4 GeV is made by including contributions from various known $N^*$ and $Δ^*$ resonances. It is found that for the beam energy around 1.5 GeV, the contribution of the Roper resonance $N^*_{(1440)}$ produced by the t-channel $σ$ exchange dominates over all other contributions. Since such reaction can be studied in the forthcoming $\bar{P}$ANDA experiment at Facility of Antiproton and Ion Research (FAIR), the reaction will be realistically the cleanest place for studying the properties of the Roper resonance and the best place for looking for other "missing" $N^*$ resonances with large coupling to $Nσ$.

hep-ph

Role of the N*(1440) resonance in the pp->pnpi+ reaction

New measurement by CELSIUS-WASA Collaboration on the pp->pn pi+ reaction reveals clear evidence for the presence of the Roper resonance N*(1440) which has been ignored in previous theoretical calculations. In this article, based on an effective Lagrangian approach and available knowledge on the Roper resonance, we investigate the role of the Roper resonance for the pp->pn pi+ reaction. It is found that the contribution from the Roper resonance N*(1440) becomes significant for kinetic energy above 1.1 GeV, consistent with the new experimental observation. The t-channel sigma-meson exchange is dominant for the production of the Roper resonance.

nucl-th