arXiv · 2211.03656
Towards learning to explain with concept bottleneck models: mitigating information leakage
Abstract
Concept bottleneck models perform classification by first predicting which of a list of human provided concepts are true about a datapoint. Then a downstream model uses these predicted concept labels to predict the target label. The predicted concepts act as a rationale for the target prediction. Model trust issues emerge in this paradigm when soft concept labels are used: it has previously been observed that extra information about the data distribution leaks into the concept predictions. In this work we show how Monte-Carlo Dropout can be used to attain soft concept predictions that do not contain leaked information.
Explore related subjects
Keep this discovery
Joshua Lockhart, Nicolas Marchesotti, Daniele Magazzeni, Manuela Veloso. 2022-11-07. Towards learning to explain with concept bottleneck models: mitigating information leakage. https://arxiv.org/abs/2211.03656
Cite the original work for its findings. Save a collection to share your selection of sources.