TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
TRACE trains FP4 rollout paths to match the quantized training path, instead of tuning each side separately.
The paper targets the rollout cost of RL post-training for MoE language models. Its method uses rollout-side quantization results to guide FP4 rounding decisions during training, reducing train-rollout mismatch. It also caches selected mantissa and scale information from deeper layers to limit the overhead of that guidance. Across four large-scale MoE models, the authors report FP4 weight/activation and KV-cache rollout performance comparable to BF16, with up to 5.4x rollout speedup. HF Daily Papers' note
The paper targets the rollout cost of RL post-training for MoE language models. Its method uses rollout-side quantization results to guide FP4 rounding decisions during training, reducing train-rollout mismatch. It also caches selected mantissa and scale information from deeper layers to limit the overhead of that guidance. Across four large-scale MoE models, the authors report FP4 weight/activation and KV-cache rollout performance comparable to BF16, with up to 5.4x rollout speedup. HF Daily Papers' note
score 5