arXiv · 2609.30849
Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength
Abstract
Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to measure perturbation strength in a unified manner across input and CoT perturbations. We then evaluate the self-consistency in explanations generated from various LLMs under controlled strength conditions, ensuring a fair comparison across perturbation types. Experiments show that our proposed LLM-based perturbation strength measure outperforms other embedding- and probability-based approaches and that input perturbations generally affect LLMs more strongly than CoT perturbations. Our work suggests that judgments about a model's self-consistency is fair only within the same perturbation type.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Phuong Q. Le, Kemal Kurniawan, Jey Han Lau. 2026-09-25. Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength. https://arxiv.org/abs/2609.30849
Cite the original work for its findings. Save a collection to share your selection of sources.