arXiv · 2510.04031
Does Using Counterfactual Help LLMs Explain Textual Importance in Classification?
Abstract
Large language models (LLMs) are becoming useful in many domains due to their impressive abilities that arise from large training datasets and large model sizes. More recently, they have been shown to be very effective in textual classification tasks, motivating the need to explain the LLMs' decisions. Motivated by practical constrains where LLMs are black-boxed and LLM calls are expensive, we study how incorporating counterfactuals into LLM reasoning can affect the LLM's ability to identify the top words that have contributed to its classification decision. To this end, we introduce a framework called the decision changing rate that helps us quantify the importance of the top words in classification. Our experimental results show that using counterfactuals can be helpful.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Nelvin Tan, James Asikin Cheung, Yu-Ching Shih, Dong Yang, Amol Salunkhe. 2025-10-05. Does Using Counterfactual Help LLMs Explain Textual Importance in Classification?. https://arxiv.org/abs/2510.04031
Cite the original work for its findings. Save a collection to share your selection of sources.