arXiv · 2511.07974
Towards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency Partition
Abstract
Attribution-based explanation techniques capture key patterns to enhance visual interpretability; however, these patterns often lack the granularity needed for insight in fine-grained tasks, particularly in cases of model misclassification, where explanations may be insufficiently detailed. To address this limitation, we propose a fine-grained counterfactual explanation framework that generates both object-level and part-level interpretability, addressing two fundamental questions: (1) which fine-grained features contribute to model misclassification, and (2) where dominant local features influence counterfactual adjustments. Our approach yields explainable counterfactuals in a non-generative manner by quantifying similarity and weighting component contributions within regions of interest between correctly classified and misclassified samples. Furthermore, we introduce a saliency partition module grounded in Shapley value contributions, isolating features with region-specific relevance. Extensive experiments demonstrate the superiority of our approach in capturing more granular, intuitively meaningful regions, surpassing fine-grained methods.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lintong Zhang, Kang Yin, Seong-Whan Lee. 2025-11-11. Towards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency Partition. https://arxiv.org/abs/2511.07974
Cite the original work for its findings. Save a collection to share your selection of sources.