The (Un)reliability of Saliency methods – Google Research
In the exploration of deep model interpretation, saliency methods emerge as a popular technique for evaluating feature importance. They assign importance scores to input features, indicating their utility in model performance. High scores suggest significant performance degradation in their absence. However, investigations, such as those by Google Research, reveal the inherent unreliability of these methods. The crux of the issue lies in their sensitivity to non-influential factors and failure to maintain input invariance, leading to potentially misleading attributions. This challenges the effectiveness of saliency methods in providing accurate explanations of deep learning behaviors.