Megadose AI progress, ranked and analyzed.

Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training

· HF Daily Papers ·
The paper claims online draft co-training can speed RL rollout generation at 122B scale without derailing the policy baseline.

The system targets speculative decoding during long-context RL post-training, where rollout generation is the main cost. It adds branch attention support to context-parallel attention and uses TapChannel to move target features across pipeline stages without changing the pipeline schedule. The authors report strong scaling at 256K tokens, memory savings over prior work, and modest overhead from the feature transport path. HF Daily Papers' note

score 5

Categories: Research