SearcharxivSearch

arXiv subjects

Jing-Jing Li

Publications and source records attributed to Jing-Jing Li.

9 recordsLinked to original sources

PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm

Current AI safety frameworks, which often treat harmfulness as binary, lack the flexibility to handle borderline cases where humans meaningfully disagree. To build more pluralistic systems, it is essential to move beyond consensus and instead understand where and why disagreements arise. We introduce PluriHarms, a benchmark designed to systematically study human harm judgments across two key dimensions -- the harm axis (benign to harmful) and the agreement axis (agreement to disagreement). Our scalable framework generates prompts that capture diverse AI harms and human values while targeting cases with high disagreement rates, validated by human data. The benchmark includes 150 prompts with 15,000 ratings from 100 human annotators, enriched with demographic and psychological traits and prompt-level features of harmful actions, effects, and values. Our analyses show that prompts that relate to imminent risks and tangible harms amplify perceived harmfulness, while annotator traits (e.g., toxicity experience, education) and their interactions with prompt content explain systematic disagreement. We benchmark AI safety models and alignment methods on PluriHarms, finding that while personalization significantly improves prediction of human harm judgments, considerable room remains for future progress. By explicitly targeting value diversity and disagreement, our work provides a principled benchmark for moving beyond "one-size-fits-all" safety toward pluralistically safe AI.

cs.CY

STAC: When Innocent Tools Form Dangerous Chains for LLM Agents

As LLMs advance into autonomous agents with tool-use capabilities, they introduce security challenges that extend beyond traditional content-based LLM safety concerns. This paper introduces Sequential Tool Attack Chaining (\STAC), a novel multi-turn attack framework that exploits agent tool use. \STAC chains together tool calls that each appear harmless in isolation but, when combined, collectively enable harmful operations that only become apparent at the final execution step. At the core of \STAC is an automated, closed-loop pipeline that synthesizes executable multi-step tool chains, validates them through in-environment execution, and reverse-engineers stealthy multi-turn prompts that reliably induce agents to execute the verified malicious sequence. Using this framework, we generate and systematically evaluate 483 \STAC cases, featuring 1,352 sets of user-agent-environment interactions and spanning diverse domains, tasks, agent types, and 10 failure modes. Our evaluations show that state-of-the-art LLM agents are highly vulnerable to \STAC, with an average final attack success rate (ASR) of 91.2\% -- exceeding 90\% for all but one of the eight agents evaluated. We further perform defense analysis and find that existing prompt-based defenses provide limited protection. To address this gap, we propose a new reasoning-driven defense prompt that achieves the strongest initial-turn protection, cutting ASR by up to 28.8\%; however, this advantage erodes sharply under adaptive attacks, and an experience-based defense (ToolShield) proves more durable over sustained multi-turn interactions. These results highlight a crucial gap: defending tool-enabled agents requires reasoning over entire action sequences and their cumulative effects, rather than evaluating isolated prompts or responses.

cs.CR

SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior

The ideal AI safety moderation system would be both structurally interpretable (so its decisions can be reliably explained) and steerable (to align to safety standards and reflect a community's values), which current systems fall short on. To address this gap, we present SafetyAnalyst, a novel AI safety moderation framework. Given an AI behavior, SafetyAnalyst uses chain-of-thought reasoning to analyze its potential consequences by creating a structured "harm-benefit tree," which enumerates harmful and beneficial actions and effects the AI behavior may lead to, along with likelihood, severity, and immediacy labels that describe potential impacts on stakeholders. SafetyAnalyst then aggregates all effects into a harmfulness score using 28 fully interpretable weight parameters, which can be aligned to particular safety preferences. We applied this framework to develop an open-source LLM prompt safety classification system, distilled from 18.5 million harm-benefit features generated by frontier LLMs on 19k prompts. On comprehensive benchmarks, we show that SafetyAnalyst (average F1=0.81) outperforms existing moderation systems (average F1$<$0.72) on prompt safety classification, while offering the additional advantages of interpretability, transparency, and steerability.

cs.CL

Latent Variable Sequence Identification for Cognitive Models with Neural Network Estimators

Extracting time-varying latent variables from computational cognitive models is a key step in model-based neural analysis, which aims to understand the neural correlates of cognitive processes. However, existing methods only allow researchers to infer latent variables that explain subjects' behavior in a relatively small class of cognitive models. For example, a broad class of relevant cognitive models with analytically intractable likelihood is currently out of reach from standard techniques, based on Maximum a Posteriori parameter estimation. Here, we present an approach that extends neural Bayes estimation to learn a direct mapping between experimental data and the targeted latent variable space using recurrent neural networks and simulated datasets. We show that our approach achieves competitive performance in inferring latent variable sequences in both tractable and intractable models. Furthermore, the approach is generalizable across different computational models and is adaptable for both continuous and discrete latent spaces. We then demonstrate its applicability in real world datasets. Our work underscores that combining recurrent neural networks and simulation-based inference to identify latent variable sequences can enable researchers to access a wider class of cognitive models for model-based neural analyses, and thus test a broader set of theories.

cs.LG

Enhancement of electron-positron pairs in combined potential wells with linear chirp frequency

The effect of linear chirp frequency on the process of electron-positron pairs production from vacuum in the combined potential wells is investigated by computational quantum field theory. Numerical results of electron number and energy spectrum under different frequency modulation parameters are obtained. By comparing with the fixed frequency, it is found that frequency modulation has a significant enhancement effect on the number of electrons. Especially when the frequency is small, appropriate frequency modulation enhances multiphoton processes in pair creation, thus promoting the pair creation. However, the number of electrons created by high frequency oscillating combined potential wells decreases after frequency modulation due to the phenomenon of high frequency suppression. The contours of the number of electrons varying with frequency and frequency modulation parameters are given, which may provide theoretical reference for possible experiments.

quant-ph

A catalogue of 74 new open clusters found in Gaia Data-Release 2

Based on astrometric data from Gaia DR2, we employ an unsupervised machine learning method to blindly search for open star clusters in the Milky Way within the Galactic latitude range of |b| < 20 degrees. In addition to 2,080 known clusters, 74 new open cluster candidates are found. In this work, we present the positions, apparent radii, parallaxes, proper motions and member stars of these candidates (https://cdsarc.u-strasbg.fr/ftp/vizier.submit//new_OC/). Meanwhile, to obtain the physical parameters of each candidate cluster, stellar isochrones are fit to the photometric data. The results show that the apparent radii and the observed proper motion dispersions of these new candidates are consistent with those of open clusters previously identified in Gaia DR2.

astro-ph.GA

Spin determination from the in-plane angular correlation analysis for various coordinate systems

In a reaction to excite the resonant state followed by the sequential cluster-decay, the in-plane angular correlation method is usually applied to determine the spin of the mother nucleus. However, the correlation pattern exhibited in a two-dimensional angular-correlation spectrum depends on the selected coordinate system. Particularly the parity-symmetric and the axial-symmetric processes should be presented in a way to enhance the correlation pattern whereas the non-symmetric process should be plotted elsewhere in order to reduce the correlation background. In this article, three possible coordinate systems, which were previously adopted in the literature, are described and compared to each other. The consistency of these systems is evaluated based on the real experimental data analysis for the 10.29-MeV state in $^{18}$O. A spin-parity of 4$^+$ is obtained for all three coordinate systems.

physics.data-an

A multiwavelength study of massive star-forming region IRAS 22506+5944

We present a multi-line study of the massive star-forming region IRAS 22506+5944. A new 6.7 GHz methanol maser was detected. 12CO, 13CO, C18O and HCO+ J = 1-0 transition observations reveal a star formation complex consisting mainly of two cores. The dominant core has a mass of more than 200 solar mass, while another one only about 35 solar mass. Both cores are obviously at different evolutionary stages. A 12CO energetic bipolar outflow was detected with an outflow mass of about 15 solar mass.

astro-ph.SR