SearcharxivSearch

arXiv subjects

Jingyu Peng

Publications and source records attributed to Jingyu Peng.

10 recordsLinked to original sources

MOSAIC: Composable Safety Alignment with Modular Control Tokens

Safety alignment in large language models (LLMs) is commonly implemented as a single static policy embedded in model parameters. However, real-world deployments often require context-dependent safety rules that vary across users, regions, and applications. Existing approaches struggle to provide such conditional control: parameter-level alignment entangles safety behaviors with general capabilities, while prompt-based methods rely on natural language instructions that provide weak enforcement. We propose MOSAIC, a modular framework that enables compositional safety alignment through learnable control tokens optimized over a frozen backbone model. Each token represents a safety constraint and can be flexibly activated and composed at inference time. To train compositional tokens efficiently, we introduce order-based task sampling and a distribution-level alignment objective that mitigates over-refusal. Experiments show that MOSAIC achieves strong defense performance with substantially lower over-refusal while preserving model utility.

cs.AI

How to Utilize Complementary Vision-Text Information for 2D Structure Understanding

LLMs typically linearize 2D tables into 1D sequences to fit their autoregressive architecture, which weakens row-column adjacency and other layout cues. In contrast, purely visual encoders can capture spatial cues, yet often struggle to preserve exact cell text. Our analysis reveals that these two modalities provide highly distinct information to LLMs and exhibit strong complementarity. However, direct concatenation and other fusion methods yield limited gains and frequently introduce cross-modal interference. To address this issue, we propose DiVA-Former, a lightweight architecture designed to effectively integrate vision and text information. DiVA-Former leverages visual tokens as dynamic queries to distill long textual sequences into digest vectors, thereby effectively exploiting complementary vision--text information. Evaluated across 13 table benchmarks, DiVA-Former improves upon the pure-text baseline by 23.9\% and achieves consistent gains over existing baselines using visual inputs, textual inputs, or a combination of both.

cs.CV

Generation of proton beams at switchback boundary-like rotational discontinuities in the solar wind

Alfv\'enic rotational discontinuities (RDs) are abundant in the inner heliosphere and can be used to model the boundary of switchbacks, i.e. Alfv\'enic magnetic kinks. To investigate the effects of RDs on proton kinetics, we model a pair of switchback-boundary-like RDs with a hybrid Particle-In-Cell (PIC) approach in a 2D system. We find that, at one of the boundary RDs, a significant population of protons remains trapped over long times, creating a secondary beam-like component with temperature anisotropy $T_\perp/T_\|\gtrsim4$ in the proton velocity distribution function that excites ion cyclotron waves within the downstream portion of the transition layer. Further analysis suggests that the static electric field in the vicinity of the RD is the key factor in trapping the protons. This work indicates that switchback boundaries could represent a viable environment for the creation of proton beams in the heliosphere; it also highlights the need to investigate RD sub-structures, especially the embedded current systems of interplanetary RDs. Finally, this paper underscores the importance of high-resolution observations of the solar wind velocity distributions around RDs.

physics.space-ph

Learning a Single Token to Replace Long System Prompts in LLMs

Long system prompts are widely used to steer Large Language Models (LLMs), but repeatedly processing them at inference time is inefficient and consumes valuable context budget. This motivates a central question: can the behavioral effect of a long system prompt be retained using only a minimal learned representation? To enable this, we propose a lightweight training framework that learns a single Behavior-Equivalent Token ([BE]). The framework first trains [BE] to encode the semantic content of the original system prompt via reconstruction, and then distills the prompt's downstream behavior into this single token. Importantly, our method requires no update to the pretrained LLM weights, no auxiliary compression models, and no labeled responses. Empirical evaluations on three datasets show that replacing long prompts with a single [BE] token yields up to a $3000\times$ prompt compression ratio, while retaining about 98% of the downstream performance of the original system prompts. This substantially reduces inference cost and frees nearly the entire context window for user inputs and model outputs.

cs.CL

AdaSwitch: Balancing Exploration and Guidance in Knowledge Distillation via Adaptive Switching

Small language models (SLMs) are crucial for applications with strict latency and computational constraints, yet achieving high performance remains challenging. Knowledge distillation (KD) can transfer capabilities from large teacher models, but existing methods face a dilemma: off-policy distillation provides high-quality supervision but suffers from exposure bias (training inference mismatch), while on-policy approaches ensure consistency but are limited by the low quality of student-generated outputs. To address these issues, we propose AdaSwitch, a novel approach that dynamically combines on-policy and off-policy generation via an adaptive switching mechanism. AdaSwitch allows the student to explore its predictions within its capability and selectively integrates teacher guidance only when divergence exceeds a context-aware threshold. This paradigm preserves generation consistency while ensuring high-quality supervision. Experiments on three datasets demonstrate that AdaSwitch consistently improves accuracy and reasoning capability with moderate overhead.

cs.CL

Ion Stochastic Heating by Low-frequency Alfv\'en Wave Spectrum

Finite-amplitude low-frequency Alfv\'en waves are commonly found in plasma environments, such as space plasmas, and play a crucial role in ion heating. The nonlinear interaction between oblique Alfv\'en wave spectra and ions has been studied. As the number of wave modes increases, ions are more likely to exhibit chaotic motion and experience stochastic heating. The stochastic heating threshold in the parameter space can be characterized by a single parameter, the effective relative curvature radius $P_{{eff.}}$. The results show excellent agreement with the chaotic regions identified through test particle simulations. The anisotropic characteristics of stochastic heating are explained using a uniform solid angle distribution model. The stochastic heating rate $Q=\dot{T}$ is calculated, and its relationship with wave conditions is expressed as $Q/(\Omega_i m_i v_A^2) = H(\alpha) \tilde{v}^3 \tilde{B}_w^2 \tilde{\omega}_1$, where $\alpha$ is propagating angle, $\Omega_i$ is the gyrofrequency, $m_i$ is the ion mass, $v_A$ is the Alfv\'en speed, $\tilde{v}$ is the dimensionless speed, $\tilde{B}_w$ is the dimensionless wave amplitude, and $\tilde{\omega}_1$ is the lowest dimensionless wave frequency.

physics.plasm-ph

Chaotic Motion of Ions In Finite-amplitude Low-frequency Alfv\'en Waves

Finite-amplitude low-frequency Alfv\'en waves (AWs) are ubiquitous in space plasmas, where they play a key role in the transport and dissipation of energy, particularly in the heating of ions in the solar corona and solar wind. In this study, we investigate the nonlinear interaction between ions and obliquely propagating AWs. When the wave amplitude and propagation angle lie within specific ranges, ion motion becomes chaotic. We quantify this behavior using the maximum Lyapunov exponent ($\lambda_{\mathrm{m}}$) and define a new parameter, the Chaos Ratio (CR), to describe the fraction of chaotic particles across different initial states. The global chaos threshold is determined as the contour CR = 0.01. Analysis of magnetic moment variations reveals that the physical origin of chaos is pitch-angle scattering induced by \textit{wave-driven field-line curvature} (WFLC), which disrupts adiabatic invariance and leads to stochastic ion energization. The onset condition for chaos can be expressed by an effective relative curvature radius, $P_{eff.} < 25$. This analytical criterion delineates the boundary of the chaotic region in the ($k_x$, $k_z$, $B_w$) parameter space and agrees well with numerical results. The identified WFLC mechanism provides a new physical pathway for converting macroscale Alfv\'enic disturbances into microscopic ion heating. \textbf{This analysis offers a simplified model that illustrates a plausible ion energization mechanism in Alfv\'enic turbulent plasmas}, including those associated with solar wind switchbacks and coronal fluctuations. These results highlight a universal chaotic process that may underlie stochastic heating in heliospheric and astrophysical plasmas.

physics.plasm-ph

Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression

Despite substantial advancements in aligning large language models (LLMs) with human values, current safety mechanisms remain susceptible to jailbreak attacks. We hypothesize that this vulnerability stems from distributional discrepancies between alignment-oriented prompts and malicious prompts. To investigate this, we introduce LogiBreak, a novel and universal black-box jailbreak method that leverages logical expression translation to circumvent LLM safety systems. By converting harmful natural language prompts into formal logical expressions, LogiBreak exploits the distributional gap between alignment data and logic-based inputs, preserving the underlying semantic intent and readability while evading safety constraints. We evaluate LogiBreak on a multilingual jailbreak dataset spanning three languages, demonstrating its effectiveness across various evaluation settings and linguistic contexts.

cs.CL

Stepwise Reasoning Error Disruption Attack of LLMs

Large language models (LLMs) have made remarkable strides in complex reasoning tasks, but their safety and robustness in reasoning processes remain underexplored. Existing attacks on LLM reasoning are constrained by specific settings or lack of imperceptibility, limiting their feasibility and generalizability. To address these challenges, we propose the Stepwise rEasoning Error Disruption (SEED) attack, which subtly injects errors into prior reasoning steps to mislead the model into producing incorrect subsequent reasoning and final answers. Unlike previous methods, SEED is compatible with zero-shot and few-shot settings, maintains the natural reasoning flow, and ensures covert execution without modifying the instruction. Extensive experiments on four datasets across four different models demonstrate SEED's effectiveness, revealing the vulnerabilities of LLMs to disruptions in reasoning processes. These findings underscore the need for greater attention to the robustness of LLM reasoning to ensure safety in practical applications. Our code is available at: https://github.com/Applied-Machine-Learning-Lab/SEED-Attack.

cs.AI

Towards Few-shot Self-explaining Graph Neural Networks

Recent advancements in Graph Neural Networks (GNNs) have spurred an upsurge of research dedicated to enhancing the explainability of GNNs, particularly in critical domains such as medicine. A promising approach is the self-explaining method, which outputs explanations along with predictions. However, existing self-explaining models require a large amount of training data, rendering them unavailable in few-shot scenarios. To address this challenge, in this paper, we propose a Meta-learned Self-Explaining GNN (MSE-GNN), a novel framework that generates explanations to support predictions in few-shot settings. MSE-GNN adopts a two-stage self-explaining structure, consisting of an explainer and a predictor. Specifically, the explainer first imitates the attention mechanism of humans to select the explanation subgraph, whereby attention is naturally paid to regions containing important characteristics. Subsequently, the predictor mimics the decision-making process, which makes predictions based on the generated explanation. Moreover, with a novel meta-training process and a designed mechanism that exploits task information, MSE-GNN can achieve remarkable performance on new few-shot tasks. Extensive experimental results on four datasets demonstrate that MSE-GNN can achieve superior performance on prediction tasks while generating high-quality explanations compared with existing methods. The code is publicly available at https://github.com/jypeng28/MSE-GNN.

cs.LG