Megadose Built for builders and researchers.

What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Document

· HF Daily Papers ·
Gradients raised exact 32-token document recovery from 13.71% to 37.77% in the paper’s GPT-2 split-learning attack.

The authors say an observer at the split can reconstruct most client text from activations alone, then recover more when gradients are also visible. In their GPT-2 test, token recovery rose from 94.20% to 97.38% with gradients. They stress that document-level counting makes the added risk look much larger, because one wrong token fails the whole document. Their defense test found secret mixup nearly blocked exact document recovery while still leaving 83-91% of tokens exposed. HF Daily Papers' note

score 5

Categories: Research