UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers
UNMASK automates the hunt for shortcut features, then tests whether models actually rely on them.
The paper says the pipeline finds surface patterns in unlabeled training data, validates them statistically, and checks causal dependence with counterfactual interventions. On MNLI classifiers, it rediscovers lexical-overlap and negation biases, verifying 9 of 10 features for BERT and 6 for RoBERTa. Using those confirmed features for Deep Feature Reweighting improved HANS accuracy by up to 12.58 points. On CivilComments-WILDS, its programmatic groups matched hand-labeled DFR worst-group accuracy without demographic annotation. HF Daily Papers' note
The paper says the pipeline finds surface patterns in unlabeled training data, validates them statistically, and checks causal dependence with counterfactual interventions. On MNLI classifiers, it rediscovers lexical-overlap and negation biases, verifying 9 of 10 features for BERT and 6 for RoBERTa. Using those confirmed features for Deep Feature Reweighting improved HANS accuracy by up to 12.58 points. On CivilComments-WILDS, its programmatic groups matched hand-labeled DFR worst-group accuracy without demographic annotation. HF Daily Papers' note
score 4