LESSER: Post-Training Data Selection with Output-Layer Gradients
LESSER cuts data-selection feature extraction cost by using output-layer gradients instead of full-parameter gradients.
The paper presents LESSER as a drop-in wrapper for post-training data selection methods. It reports a 9.7x FLOP reduction for supervised fine-tuning benchmarks and 3.0x for RL benchmarks. The authors say downstream performance tracks full-gradient selection even when individual sample rankings differ. ArXiv · AI/CL/LG's note
The paper presents LESSER as a drop-in wrapper for post-training data selection methods. It reports a 9.7x FLOP reduction for supervised fine-tuning benchmarks and 3.0x for RL benchmarks. The authors say downstream performance tracks full-gradient selection even when individual sample rankings differ. ArXiv · AI/CL/LG's note
score 5