Megadose Built for builders and researchers.

LESSER: Post-Training Data Selection with Output-Layer Gradients

· ArXiv · AI/CL/LG ·
LESSER cuts data-selection feature extraction cost by using output-layer gradients instead of full-parameter gradients.

The paper presents LESSER as a drop-in wrapper for post-training data selection methods. It reports a 9.7x FLOP reduction for supervised fine-tuning benchmarks and 3.0x for RL benchmarks. The authors say downstream performance tracks full-gradient selection even when individual sample rankings differ. ArXiv · AI/CL/LG's note

score 5

Categories: Research