Megadose Built for builders and researchers.

Sharpening Tax in Post-Training

· HF Daily Papers ·
Post-training may make agent models more reliable on first tries while reducing what they can solve with repeated attempts.

The paper says base LLMs, when given a light inference harness, can sometimes beat post-trained versions on pass@K despite lower pass@1. Its proposed “Sharpening Tax” measures that loss of test-time scalability after post-training. Across 14 model pairs and three agentic benchmarks, the authors report the tax in most settings. They also introduce PTGS, a sampler meant to reduce the tax during RL training while improving single-shot accuracy. HF Daily Papers' note

score 5

Categories: Research