SearcharxivSearch

arXiv subjects

Zhenhao Zhu

Publications and source records attributed to Zhenhao Zhu.

4 recordsLinked to original sources

GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio

We present GuardReasoner-Omni, a reasoning-based guardrail model designed to moderate text, image, video, and audio data. First, we construct a comprehensive training corpus comprising 181k samples spanning these four modalities. Our training pipeline follows a two-stage paradigm to incentivize the model to deliberate before making decisions: (1) conducting SFT to cold-start the model with explicit reasoning capabilities and structural adherence; and (2) performing RL with a concise correctness reward to preserve accurate reasoning while suppressing redundant generation. We release a suite of models scaled at 3B and 7B parameters. Extensive experiments demonstrate that GuardReasoner-Omni achieves superior performance compared to existing state-of-the-art baselines across various guardrail benchmarks.

cs.CR

ExtendAttack: Attacking Servers of LRMs via Extending Reasoning

Large Reasoning Models (LRMs) have demonstrated promising performance in complex tasks. However, the resource-consuming reasoning processes may be exploited by attackers to maliciously occupy the resources of the servers, leading to a crash, like the DDoS attack in cyber. To this end, we propose a novel attack method on LRMs termed ExtendAttack to maliciously occupy the resources of servers by stealthily extending the reasoning processes of LRMs. Concretely, we systematically obfuscate characters within a benign prompt, transforming them into a complex, poly-base ASCII representation. This compels the model to perform a series of computationally intensive decoding sub-tasks that are deeply embedded within the semantic structure of the query itself. Extensive experiments demonstrate the effectiveness of our proposed ExtendAttack. Remarkably, it significantly increases response length and latency, with the former increasing by over 2.7 times for the o3 model on the HumanEval benchmark. Besides, it preserves the original meaning of the query and achieves comparable answer accuracy, showing the stealthiness.

cs.CR

Option-ID Based Elimination For Multiple Choice Questions

Multiple choice questions (MCQs) are a popular and important task for evaluating large language models (LLMs). Based on common strategies people use when answering MCQs, the process of elimination (PoE) has been proposed as an effective problem-solving method. Existing PoE methods typically either have LLMs directly identify incorrect options or score options and replace lower-scoring ones with [MASK]. However, both methods suffer from inapplicability or suboptimal performance. To address these issues, this paper proposes a novel option-ID based PoE ($\text{PoE}_{\text{ID}}$). $\text{PoE}_{\text{ID}}$ critically incorporates a debiasing technique to counteract LLMs token bias, enhancing robustness over naive ID-based elimination. It features two strategies: $\text{PoE}_{\text{ID}}^{\text{log}}$, which eliminates options whose IDs have log probabilities below the average threshold, and $\text{PoE}_{\text{ID}}^{\text{seq}}$, which iteratively removes the option with the lowest ID probability. We conduct extensive experiments with 6 different LLMs on 4 diverse datasets. The results demonstrate that $\text{PoE}_{\text{ID}}$, especially $\text{PoE}_{\text{ID}}^{\text{log}}$, significantly improves zero-shot and few-shot MCQs performance, particularly in datasets with more options. Our analyses demonstrate that $\text{PoE}_{\text{ID}}^{\text{log}}$ enhances the LLMs' confidence in selecting the correct option, and the option elimination strategy outperforms methods relying on [MASK] replacement. We further investigate the limitations of LLMs in directly identifying incorrect options, which stem from their inherent deficiencies.

cs.CL

Ram-pressure stripped radio tail and two ULXs in the spiral galaxy HCG 97b

We report LOFAR and VLA detections of extended radio emission in the spiral galaxy HCG 97b, hosted by an X-ray bright galaxy group. The extended radio emission detected at 144 MHz, 1.4 GHz and 4.86 GHz is elongated along the optical disk and has a tail that extends 27 kpc in projection towards the centre of the group at GHz frequencies or 60 kpc at 144 MHz. Chandra X-ray data show two off-nuclear ultra-luminous X-ray sources (ULXs), with the farther one being a plausible candidate for an accreting intermediate-mass black hole (IMBH). The asymmetry observed in both CO emission morphology and kinematics indicates that HCG 97b is undergoing ram-pressure stripping, with the leading side at the southeastern edge of the disk. Moreover, the VLA 4.86 GHz image reveals two bright radio blobs near one ULX, aligning with the disk and tail, respectively. The spectral indices in the disk and tail are comparable and flat ($α> -1$), suggesting the presence of recent outflows potentially linked to ULX feedback. This hypothesis gains support from estimates showing that the bulk velocity of the relativistic electrons needed for transport from the disk to the tail is approximately $\sim 1300$ $\rm km~s^{-1}$. This velocity is much higher than those observed in ram-pressure stripped galaxies ($100-600$ $\rm km~s^{-1}$), implying an alternative mechanism aiding the stripping process. Therefore, we conclude that HCG 97b is subject to ram pressure, with the formation of its stripped radio tail likely influenced by the putative IMBH activities.

astro-ph.GA