Towards Debiasing DNN Models from Spurious Feature Influence

Mengnan Du; Ruixiang Tang; Weijie Fu; Xia Hu

doi:10.1609/aaai.v36i9.21185

Authors

Mengnan Du Texas A&M University
Ruixiang Tang Rice University
Weijie Fu Hefei University of Technology
Xia Hu Rice University

DOI:

https://doi.org/10.1609/aaai.v36i9.21185

Keywords:

Philosophy And Ethics Of AI (PEAI)

Abstract

Recent studies indicate that deep neural networks (DNNs) are prone to show discrimination towards certain demographic groups. We observe that algorithmic discrimination can be explained by the high reliance of the models on fairness sensitive features. Motivated by this observation, we propose to achieve fairness by suppressing the DNN models from capturing the spurious correlation between those fairness sensitive features with the underlying task. Specifically, we firstly train a bias-only teacher model which is explicitly encouraged to maximally employ fairness sensitive features for prediction. The teacher model then counter-teaches a debiased student model so that the interpretation of the student model is orthogonal to the interpretation of the teacher model. The key idea is that since the teacher model relies explicitly on fairness sensitive features for prediction, the orthogonal interpretation loss enforces the student network to reduce its reliance on sensitive features and instead capture more task relevant features for prediction. Experimental analysis indicates that our framework substantially reduces the model's attention on fairness sensitive features. Experimental results on four datasets further validate that our framework has consistently improved the fairness with respect to three group fairness metrics, with a comparable or even better accuracy.

Towards Debiasing DNN Models from Spurious Feature Influence

Authors

DOI:

Keywords:

Abstract

Downloads

Published

How to Cite

Issue

Section

Information