SAGE: Subpopulation-Aware Generative Enhancement for Mitigating Spurious Correlations
SAGE uses synthetic augmentation to balance hidden subpopulations when spurious labels are not available.
The paper proposes a two-stage framework that clusters training data into sub-labels, then fine-tunes a conditional generative model and text encoder to create targeted synthetic examples. Those generated samples are used to fill underrepresented regions and build a balanced validation set for last-layer reweighting. The authors report worst-group accuracy of 89.5% on Waterbirds, 85.7% on CelebA, and 79.1% on MetaShift, beating group-label-free baselines by up to 7.7 points. ArXiv · AI/CL/LG's note
The paper proposes a two-stage framework that clusters training data into sub-labels, then fine-tunes a conditional generative model and text encoder to create targeted synthetic examples. Those generated samples are used to fill underrepresented regions and build a balanced validation set for last-layer reweighting. The authors report worst-group accuracy of 89.5% on Waterbirds, 85.7% on CelebA, and 79.1% on MetaShift, beating group-label-free baselines by up to 7.7 points. ArXiv · AI/CL/LG's note
score 4