SearcharxivSearch

arXiv subjects

Shengyun Si

Publications and source records attributed to Shengyun Si.

3 recordsLinked to original sources

Strong but Brittle: Simple Attacks Subvert Reasoning-based Safety Guardrails

Open-weight Large Reasoning Models (LRMs) are approaching the capabilities of their frontier counterparts but pose significant safety concerns, as they are difficult to patch or monitor post-release. To prevent misuse, reasoning-based safety guardrails, where models explicitly reason on safety justifications before answering, have become a promising primary defense, achieving near-perfect refusal rates on harmful queries. We show that this strong defense is alarmingly brittle and can be subverted by embarrassingly simple attacks to elicit extremely forbidden questions, such as `How to kill a man without being caught?' Specifically, we identify one systematic vulnerability from the reasoning-then-answer mechanism: the stage-transition logic that governs when safety reasoning begins and ends can be trivially manipulated. Based on this finding, we develop four simple yet effective red-teaming methods that systematically subvert different stages of the guardrails, either by bypassing reasoning entirely or exploiting it to produce targeted harmful content. These methods achieve attack success rates up to 90\% across five benchmarks on multiple LRM families. Our findings reveal that reasoning-based guardrails are necessary but not sufficient and must be paired with robust triggering mechanisms and stronger base-model alignment. Code is in https://chenxshuo.github.io/bag-of-tricks/.

cs.CR

Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior

Recent advancements in large language models (LLMs) have demonstrated that fine-tuning and human alignment can render LLMs harmless. In practice, such "harmlessness" behavior is mainly achieved by training models to reject harmful requests, such as "Explain how to burn down my neighbor's house", where the model appropriately declines to respond. However, this approach can inadvertently result in false refusal, where models reject benign queries as well, such as "Tell me how to kill a Python process". In this work, we demonstrate that prompting safety reflection before generating a response can mitigate false refusal behavior. Building on this finding, we introduce the Think-Before-Refusal (TBR) schema and conduct safety-aware instruction fine-tuning incorporating safety reflection. In an ablation study across 15 pre-trained models, we show that models fine-tuned with safety reflection significantly reduce false refusal behavior while maintaining safety and overall performance compared to those fine-tuned without safety reflection.

cs.CL

Investigating Table-to-Text Generation Capabilities of LLMs in Real-World Information Seeking Scenarios

Tabular data is prevalent across various industries, necessitating significant time and effort for users to understand and manipulate for their information-seeking purposes. The advancements in large language models (LLMs) have shown enormous potential to improve user efficiency. However, the adoption of LLMs in real-world applications for table information seeking remains underexplored. In this paper, we investigate the table-to-text capabilities of different LLMs using four datasets within two real-world information seeking scenarios. These include the LogicNLG and our newly-constructed LoTNLG datasets for data insight generation, along with the FeTaQA and our newly-constructed F2WTQ datasets for query-based generation. We structure our investigation around three research questions, evaluating the performance of LLMs in table-to-text generation, automated evaluation, and feedback generation, respectively. Experimental results indicate that the current high-performing LLM, specifically GPT-4, can effectively serve as a table-to-text generator, evaluator, and feedback generator, facilitating users' information seeking purposes in real-world scenarios. However, a significant performance gap still exists between other open-sourced LLMs (e.g., Tulu and LLaMA-2) and GPT-4 models. Our data and code are publicly available at https://github.com/yale-nlp/LLM-T2T.

cs.CL