arXiv · 2007.06898
Our Evaluation Metric Needs an Update to Encourage Generalization
Abstract
Models that surpass human performance on several popular benchmarks display significant degradation in performance on exposure to Out of Distribution (OOD) data. Recent research has shown that models overfit to spurious biases and `hack' datasets, in lieu of learning generalizable features like humans. In order to stop the inflation in model performance -- and thus overestimation in AI systems' capabilities -- we propose a simple and novel evaluation metric, WOOD Score, that encourages generalization during evaluation.
Explore related subjects
Keep this discovery
Swaroop Mishra, Anjana Arunkumar, Chris Bryan, Chitta Baral. 2020-07-14. Our Evaluation Metric Needs an Update to Encourage Generalization. https://arxiv.org/abs/2007.06898
Cite the original work for its findings. Save a collection to share your selection of sources.