TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning
TRIAGE keeps native 4-bit RL training stable while matching full-precision results in the paper’s math benchmarks.
The paper says learner-sampler mismatch in NVFP4 can push policy-gradient updates in destabilizing directions before the problem becomes global. TRIAGE diagnoses that mismatch at the response-segment level, rebalances selected updates, and applies bounded repair where severe mismatch remains. The authors report stable runs on Qwen3-4B and Qwen3-30B-A3B, with full-precision-level performance across five mathematical reasoning benchmarks. Native NVFP4 with TRIAGE is reported to deliver up to 2.3x higher rollout throughput than BF16. HF Daily Papers' note
The paper says learner-sampler mismatch in NVFP4 can push policy-gradient updates in destabilizing directions before the problem becomes global. TRIAGE diagnoses that mismatch at the response-segment level, rebalances selected updates, and applies bounded repair where severe mismatch remains. The authors report stable runs on Qwen3-4B and Qwen3-30B-A3B, with full-precision-level performance across five mathematical reasoning benchmarks. Native NVFP4 with TRIAGE is reported to deliver up to 2.3x higher rollout throughput than BF16. HF Daily Papers' note
score 5