arXiv · 2609.22228
Assessing Adversarial Robustness of Latent Reasoning Models
Abstract
Large language models increasingly rely on long chain-of-thought (CoT) trajectories for complex reasoning, but autoregressive generation brings substantial memory and inference costs. Latent reasoning models (LRMs) offer a more efficient alternative by compressing intermediate reasoning into a small number of continuous latent vectors. Despite their efficiency, however, the adversarial robustness of LRMs remains largely underexplored. In this work, we systematically evaluate the robustness of latent reasoning across textual and multimodal settings, covering eight models and six benchmarks. We find that, across our evaluated settings, LRMs are generally less robust than explicit CoT baselines under adversarial perturbations, with particularly severe degradation under white-box attacks. Further analysis reveals distinct failure modes across modalities: textual latent states exhibit brittle dynamics and high sensitivity to specific input patterns, while latent states in multimodal models can remain largely invariant to input perturbations and have limited influence on final predictions. These findings expose robustness limitations of current latent reasoning approaches and highlight the need to jointly consider efficiency and robustness when designing implicit reasoning systems. We have open-sourced our code to facilitate reproduction of our research https://github.com/PKU-ML/latent-reasoning-model-assessment.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shaolong Chen, Ang Li, Mingjie Li, Yisen Wang. 2026-09-03. Assessing Adversarial Robustness of Latent Reasoning Models. https://arxiv.org/abs/2609.22228
Cite the original work for its findings. Save a collection to share your selection of sources.