SearcharxivSearch

arXiv subjects

Xinpeng Wang

Publications and source records attributed to Xinpeng Wang.

At least 19 recordsLinked to original sources

Detecting and Suppressing Reward Hacking with Gradient Fingerprints

Reinforcement learning with verifiable rewards (RLVR) typically optimizes for outcome rewards without imposing constraints on intermediate reasoning. This leaves training susceptible to reward hacking, where models exploit loopholes (e.g., spurious patterns in training data) in the reward function to achieve high scores without solving the intended task. These reward-hacking behaviors are often implicit, as the intermediate chain-of-thought (CoT) may appear plausible on the surface, limiting the effectiveness of purely text-based monitoring. We propose Gradient Fingerprint (GRIFT), a method for detecting reward hacking using models' internal computations. Given a prompt and a model-generated CoT, GRIFT computes gradients of the CoT conditioned on the prompt and compresses them into a compact representation, which is then used to assess whether the CoT reflects reward hacking behavior. Across verifiable reasoning benchmarks spanning math, code, and logical reasoning, GRIFT substantially outperforms strong baselines, including CoT Monitor and TRACE, achieving over 25% relative improvement in detecting reward hacking behavior. Moreover, integrating GRIFT into the rejection fine-tuning pipeline for reasoning tasks reduces reward hacking and improves performance on the true task objective. Our results highlight a promising direction of leveraging gradient level representations for assessing the quality of CoT reasoning traces. Our code is available at: https://github.com/songtao-x/reward_hack.

cs.LG

An Extensive Empirical Study on Code Translation Technique

Automated code translation is increasingly important for software evolution, yet the relative strengths and limitations of learning-based and large language model (LLM)-based techniques remain insufficiently understood. To address this gap, we conduct a large-scale empirical study comparing representative code translation techniques across methodological paradigms and translation granularities. We evaluate learning-based methods, LLM-based methods, and general-purpose LLMs on multilingual method-level and class-level benchmarks involving multiple programming languages. Our analysis considers executable correctness, code similarity, translation direction, translation granularity, and failure patterns. The results show that LLMs and LLM-based methods generally outperform learning-based methods in method-level correctness, although similarity metrics alone do not reliably reflect functional correctness. Translation direction substantially affects performance, particularly when translating between languages with different type-system characteristics. Class-level translation remains considerably more difficult than method-level translation because it requires preserving global semantics, interfaces, member relationships, and cross-method dependencies. Our error analysis further shows that static semantic errors and logical errors are the primary challenges in existing code translation systems. These findings provide empirical evidence and practical guidance for developing more robust, type-aware, structure-aware, and context-aware code translation techniques.

cs.SE

Reconstruction of Primordial Power Spectrum from Gravitational Waves of High-Redshift Black Hole Binaries

High-redshift binary black hole (BBH) events are promising candidates for primordial black holes (PBHs) detectable by next-generation gravitational wave (GW) detectors. A redshifted mass distribution of detected PBH candidates can be obtained from GW observations, from which the underlying PBH mass function can be reconstructed. In this work, we develop a framework that applies the gradient-descent method to the observed redshifted mass distribution and reconstructs the PBH mass function and, subsequently, the primordial power spectrum (PPS) on small scales. As an illustrative application, we analyze BBH events in the LIGO--Virgo--KAGRA (LVK) catalogs under a specified PBH selection criterion. We find a regularization-stable candidate bump-like enhancement of order $\mathcal{O}(10^{-2})$ in the reconstructed PPS, centered around $k_{\mathrm{peak}}\simeq 5.7\times 10^5~\mathrm{Mpc}^{-1}$ under the adopted assumptions. Our results demonstrate the feasibility of reconstructing the small-scale PPS from high-redshift BBH observations with next-generation GW detectors.

astro-ph.CO

Automatic Layer Selection for Hallucination Detection

Recent studies on hallucination detection have shown that hallucination-related signals are more strongly encoded in intermediate layers than in the final layer of large language models (LLMs). Although a growing body of work has sought to exploit this property for hallucination detection, how to automate the selection of high-performing layers remains underexplored, and principled methods for this purpose are still lacking. To address this gap, we first propose several hypotheses for why such signals emerge in intermediate layers and evaluate corresponding criteria for automatic layer selection across diverse LLM architectures, scales, and tasks, covering both question answering and summarization hallucination detection benchmarks. However, we find that none of these criteria consistently delivers satisfactory performance. We therefore propose a new selection criterion, First Effective Peak of Intrinsic Dimension (FEPoID), which consistently identify optimal or near-optimal layers and outperforms both the aforementioned criteria and existing hallucination detection baselines. FEPoID is training-free and incurs negligible computational overhead. In addition, we study the generation behaviors of LLMs and introduce a simple yet effective truncation strategy, which further amplifies hallucination-related signals and substantially improves overall detection performance. Code is publicly available at https://github.com/DesoloYw/Automatic-Layer-Selection-for-Hallucination-Detection.git

cs.AI

Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization

Chinese demonstrates high semantic compactness and rich metaphorical expressiveness, enabling limited text to convey dense meanings while increasing the difficulty of generation and verification, particularly in short-form creative natural language generation (CNLG). In the real world, users often require personalized, fine-grained creative constraints, making reliable verification critical to guiding optimization. According to Brunswik's Lens Model from psychology, constraints' achievement can be inferred from sufficient observable cues. Existing studies are mainly outcome-oriented, implicitly assuming that the outcome itself provides adequate cues for verification. However, this assumption breaks down in Chinese short-form CNLG (e.g., naming or advertising) with diverse personalized constraints, where extremely brief outcomes inherently offer limited information. Explanations can naturally serve as extra cues. Nevertheless, under complex constraints, LLMs' explanations may suffer from hallucination, incompleteness, or ambiguity. To address these, we novelly formalize the Chinese short-form CNLG task as a heterogeneous multi-objective optimization (HMO) issue that needs to jointly optimize multiple personalized constraints and explanation reliability. We further propose MAGIC-HMO, a training-free multi-agent framework that optimizes these objectives through iterative generation and verification under an explanation-oriented multi-objective strategy. Experiments on \emph{Chinese Baby Naming}, a challenging benchmark, demonstrate that MAGIC-HMO significantly outperforms six strong baselines across various LLM backbones. Relevant data and codes are available at https://github.com/foolfun/MAGIC_HMO.

cs.CL

NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics

The ability to derive precise spatial and physical insights is a cornerstone of vision-language models (VLMs), yet their poor performances in related spatial intelligence tasks such as physical reasoning remain a fundamental barrier. The community critically lacks a scientific analysis revealing whether VLMs faithfully reach answers or plausibly make guesses. This work aims to provide a fundamental understanding of how VLMs perceive the physical world, and utilize physical laws, while assessing the reliability of model confidence. We propose NICE and FACT, a dual-diagnostic paradigm that explicitly decomposes quantitative reasoning for kinematic physics: FACT diagnoses visual fidelity, physical law comprehension, and temporal grounding. NICE studies our novel neighborhood-informed calibration method and novel metrics to evaluate and calibrate confidence reliability. Evaluated across 6 latest state-of-the-art VLMs, we uncover that models fail to identify visual preconditions or utilize necessary physical laws to reach answers. This work highlights and establishes a standardized diagnostic paradigm to guide the development of faithful, physically-grounded VLMs.

cs.CV

Composite Hybrid Inflation : Primordial Black Holes and Stochastic Gravitational Waves

We investigate the production of primordial black holes and gravitational waves in composite hybrid inflation. Starting from an effective chiral Lagrangian with a dilaton and pions, we identify inflation occurring due to the walking dynamics of the theory. A $\mathbb{Z}_2$ symmetry-breaking term in the pion sector induces a shift in the inflaton's trajectory, which leads to a tachyonic instability phase. Curvature perturbations grow exponentially, producing copious primordial black holes and a stochastic gravitational wave background. We show that the primordial black hole mass and the gravitational wave frequency are strongly restricted by the anomalous dimensions of the pion operators, with larger anomalous dimensions giving lighter primordial black holes and higher frequency gravitational waves. In both cases, the associated signatures lie within reach of future gravitational wave observatories.

hep-ph

Primordial black holes save $R^2$ inflation

In light of the latest Planck and Atacama Cosmology Telescope (P-ACT) joint results on the primordial scalar power spectrum, we show that the $R^2$ inflation model extended with a non-minimally coupled scalar field $χ$--namely the $χ$-extended $R^2$ inflation model--can naturally accommodate a larger spectral index $n_s$ and a small positive running $α_s$ at cosmic microwave background (CMB) scales, both of which are consistent with the latest P-ACT constraints. This is because the $χ$ field contributes a blue-tilted component to the primordial power spectrum, which both modifies the large-scale power and, as a result, significantly enhances power on small scales. The deviation of the $n_s$ and $α_s$ from the single field $R^2$ inflation is related to the non-minimal coupling constant $ξ$. The consequent enhancement in the primordial power spectrum can be large enough to lead to the formation of primordial black holes (PBHs) of mass $\lesssim 10^{20}\mathrm{g}$ as dark matter candidates. Furthermore, future observations of the small-scale power spectrum, CMB spectral distortions, and stochastic gravitational waves will provide decisive tests of this model and its predictions for PBHs. We stress its strong connection to the seesaw mechanism for the generation of the observed small masses.

astro-ph.CO

Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort

Reward hacking, where a reasoning model exploits loopholes in a reward function to achieve high rewards without solving the intended task, poses a significant threat. This behavior may be explicit, i.e. verbalized in the model's chain-of-thought (CoT), or implicit, where the CoT appears benign thus bypasses CoT monitors. To detect implicit reward hacking, we propose TRACE (Truncated Reasoning AUC Evaluation). Our key observation is that hacking occurs when exploiting the loophole is easier than solving the actual task. This means that the model is using less 'effort' than required to achieve high reward. TRACE quantifies effort by measuring how early a model's reasoning becomes sufficient to obtain the reward. We progressively truncate a model's CoT at various lengths, force the model to answer, and estimate the expected reward at each cutoff. A hacking model, which takes a shortcut, will achieve a high expected reward with only a small fraction of its CoT, yielding a large area under the accuracy-vs-length curve. TRACE achieves over 65% gains over our strongest 72B CoT monitor in math reasoning, and over 30% gains over a 32B monitor in coding. We further show that TRACE can discover unknown loopholes during training. Overall, TRACE offers a scalable unsupervised approach for oversight where current monitoring methods prove ineffective.

cs.AI

Refusal Direction is Universal Across Safety-Aligned Languages

Refusal mechanisms in large language models (LLMs) are essential for ensuring safety. Recent research has revealed that refusal behavior can be mediated by a single direction in activation space, enabling targeted interventions to bypass refusals. While this is primarily demonstrated in an English-centric context, appropriate refusal behavior is important for any language, but poorly understood. In this paper, we investigate the refusal behavior in LLMs across 14 languages using PolyRefuse, a multilingual safety dataset created by translating malicious and benign English prompts into these languages. We uncover the surprising cross-lingual universality of the refusal direction: a vector extracted from English can bypass refusals in other languages with near-perfect effectiveness, without any additional fine-tuning. Even more remarkably, refusal directions derived from any safety-aligned language transfer seamlessly to others. We attribute this transferability to the parallelism of refusal vectors across languages in the embedding space and identify the underlying mechanism behind cross-lingual jailbreaks. These findings provide actionable insights for building more robust multilingual safety defenses and pave the way for a deeper mechanistic understanding of cross-lingual vulnerabilities in LLMs.

cs.CL

DAMRO: Dive into the Attention Mechanism of LVLM to Reduce Object Hallucination

Despite the great success of Large Vision-Language Models (LVLMs), they inevitably suffer from hallucination. As we know, both the visual encoder and the Large Language Model (LLM) decoder in LVLMs are Transformer-based, allowing the model to extract visual information and generate text outputs via attention mechanisms. We find that the attention distribution of LLM decoder on image tokens is highly consistent with the visual encoder and both distributions tend to focus on particular background tokens rather than the referred objects in the image. We attribute to the unexpected attention distribution to an inherent flaw in the visual encoder itself, which misguides LLMs to over emphasize the redundant information and generate object hallucination. To address the issue, we propose DAMRO, a novel training-free strategy that $D$ive into $A$ttention $M$echanism of LVLM to $R$educe $O$bject Hallucination. Specifically, our approach employs classification token (CLS) of ViT to filter out high-attention outlier tokens scattered in the background and then eliminate their influence during decoding stage. We evaluate our method on LVLMs including LLaVA-1.5, LLaVA-NeXT and InstructBLIP, using various benchmarks such as POPE, CHAIR, MME and GPT-4V Aided Evaluation. The results demonstrate that our approach significantly reduces the impact of these outlier tokens, thus effectively alleviating the hallucination of LVLMs. The code is released at https://github.com/coder-gx/DAMRO.

cs.CL

When Tiny Halos Stir Spacetime: Gravitational Waves from Fifth-Force Mergers

Dark matter fermions interacting via attractive fifth forces mediated by a light mediator can form dark matter halos in the very early universe. We show that bound systems composed of these halos are capable of generating gravitational wave (GW) signals detectable today, even when the individual halos are very light. The Yukawa force dominates the dynamics of these halo binaries, rather than gravity. As a result, large GW signals can be produced at initially extremely high frequencies, which are then redshifted to frequency bands accessible to current or future GW observatories. In addition, the resulting GW signals carry distinctive features that enable future observations to distinguish them from conventional ones. Notably, even if only a tiny fraction of dark matter experiences strong fifth-force interactions, such effects provide a new avenue to discover self-interacting dark matter through GW observations.

astro-ph.CO

Enhancement of primordial curvature perturbations in $R^3$-corrected Starobinsky-Higgs inflation

We provide a systematic study of the Starobinsky-Higgs inflation model in the presence of an additional cubic term of the Ricci scalar. We investigate, in particular, the effects of the cubic term on the spectral index $n_s$ and the tensor-to-scalar ratio $r$. Through both analytical and numerical analyses, we show that the $R^3$-corrected Starobinsky-Higgs model can achieve compatibility with cosmic microwave background observations while producing distinct observational signatures with different frequency ranges. In addition, we discuss the complementarity between different observational probes, including the scalar-induced gravitational waves and spectral distortions, offering an independent probe of the enhanced curvature perturbations. Detection prospects are also discussed.

astro-ph.CO

Algorithmic Fidelity of Large Language Models in Generating Synthetic German Public Opinions: A Case Study

In recent research, large language models (LLMs) have been increasingly used to investigate public opinions. This study investigates the algorithmic fidelity of LLMs, i.e., the ability to replicate the socio-cultural context and nuanced opinions of human participants. Using open-ended survey data from the German Longitudinal Election Studies (GLES), we prompt different LLMs to generate synthetic public opinions reflective of German subpopulations by incorporating demographic features into the persona prompts. Our results show that Llama performs better than other LLMs at representing subpopulations, particularly when there is lower opinion diversity within those groups. Our findings further reveal that the LLM performs better for supporters of left-leaning parties like The Greens and The Left compared to other parties, and matches the least with the right-party AfD. Additionally, the inclusion or exclusion of specific variables in the prompts can significantly impact the models' predictions. These findings underscore the importance of aligning LLMs to more effectively model diverse public opinions while minimizing political biases and enhancing robustness in representativeness.

cs.CL

The Dual Primordial Black Hole Formation Scenario

We report a novel mechanism where two families of primordial black holes (PBHs) may form at nearly the same comoving scales but at two different epochs. It is realized in two-stage inflation where a non-inflationary stage is sandwiched by the two inflationary stages. In this case, smaller PBHs form when the comoving scale of interest re-enters the horizon during the break period, and larger PBHs form when the scale re-enters the horizon after inflation. This mechanism may realize both reheating of the universe through the evaporation of ultralight PBHs formed during the break stage and the dark matter by those formed after inflation. We show that this scenario may give rise to a distinctive signature in the stochastic gravitational wave background that can be tested by the near-future gravitational wave observatories such as LISA and DECIGO. Our work thus provides a unified observational window into the physics of inflation, reheating, and dark matter.

astro-ph.CO

Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior

Recent advancements in large language models (LLMs) have demonstrated that fine-tuning and human alignment can render LLMs harmless. In practice, such "harmlessness" behavior is mainly achieved by training models to reject harmful requests, such as "Explain how to burn down my neighbor's house", where the model appropriately declines to respond. However, this approach can inadvertently result in false refusal, where models reject benign queries as well, such as "Tell me how to kill a Python process". In this work, we demonstrate that prompting safety reflection before generating a response can mitigate false refusal behavior. Building on this finding, we introduce the Think-Before-Refusal (TBR) schema and conduct safety-aware instruction fine-tuning incorporating safety reflection. In an ablation study across 15 pre-trained models, we show that models fine-tuned with safety reflection significantly reduce false refusal behavior while maintaining safety and overall performance compared to those fine-tuned without safety reflection.

cs.CL

Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation

Training a language model to be both helpful and harmless requires careful calibration of refusal behaviours: Models should refuse to follow malicious instructions or give harmful advice (e.g."how do I kill someone?"), but they should not refuse safe requests, even if they superficially resemble unsafe ones (e.g. "how do I kill a Python process?"). Avoiding such false refusal, as prior work has shown, is challenging even for highly-capable language models. In this paper, we propose a simple and surgical method for mitigating false refusal in language models via single vector ablation. For a given model, we extract a false refusal vector and show that ablating this vector reduces false refusal rate while preserving the model's safety and general capabilities. We also show that our approach can be used for fine-grained calibration of model safety. Our approach is training-free and model-agnostic, making it useful for mitigating the problem of false refusal in current and future language models.

cs.CL

ClusMFL: A Cluster-Enhanced Framework for Modality-Incomplete Multimodal Federated Learning in Brain Imaging Analysis

Multimodal Federated Learning (MFL) has emerged as a promising approach for collaboratively training multimodal models across distributed clients, particularly in healthcare domains. In the context of brain imaging analysis, modality incompleteness presents a significant challenge, where some institutions may lack specific imaging modalities (e.g., PET, MRI, or CT) due to privacy concerns, device limitations, or data availability issues. While existing work typically assumes modality completeness or oversimplifies missing-modality scenarios, we simulate a more realistic setting by considering both client-level and instance-level modality incompleteness in this study. Building on this realistic simulation, we propose ClusMFL, a novel MFL framework that leverages feature clustering for cross-institutional brain imaging analysis under modality incompleteness. Specifically, ClusMFL utilizes the FINCH algorithm to construct a pool of cluster centers for the feature embeddings of each modality-label pair, effectively capturing fine-grained data distributions. These cluster centers are then used for feature alignment within each modality through supervised contrastive learning, while also acting as proxies for missing modalities, allowing cross-modal knowledge transfer. Furthermore, ClusMFL employs a modality-aware aggregation strategy, further enhancing the model's performance in scenarios with severe modality incompleteness. We evaluate the proposed framework on the ADNI dataset, utilizing structural MRI and PET scans. Extensive experimental results demonstrate that ClusMFL achieves state-of-the-art performance compared to various baseline methods across varying levels of modality incompleteness, providing a scalable solution for cross-institutional brain imaging analysis.

eess.IV