Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?
Untying embeddings beat weight tying under DP-SGD in the paper’s GPT-2 and DistilGPT2 tests.
The authors report gains of up to 4.74 percentage points in accuracy on SST-2, QNLI, and QQP. They also say untied embeddings make memory-efficient ghost clipping usable for DP-SGD. Weight tying, by contrast, creates shared-parameter interactions that complicate ghost norm computation and reduce its advantage. The paper says untied models used over 60% less memory while preserving ghost clipping benefits. ArXiv · AI/CL/LG's note
The authors report gains of up to 4.74 percentage points in accuracy on SST-2, QNLI, and QQP. They also say untied embeddings make memory-efficient ghost clipping usable for DP-SGD. Weight tying, by contrast, creates shared-parameter interactions that complicate ghost norm computation and reduce its advantage. The paper says untied models used over 60% less memory while preserving ghost clipping benefits. ArXiv · AI/CL/LG's note
score 3