OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher
A render-free RL teacher is used to fine-tune camera driving policies in closed loop without paying the usual simulation cost.
OPTED trains a privileged teacher on vectorized HD-map and bounding-box inputs, then uses that teacher to supervise pre-trained camera-based students during closed-loop post-training. The paper applies the method to TransFuser and VaVAM in AlpaSim, using 3D Gaussian Splatting reconstructions of real driving logs. Reported driving scores rise by 1.6x and 9.5x, and controlled tests match direct RL closed-loop performance with roughly 1,000x fewer simulator interactions. ArXiv · AI/CL/LG's note
OPTED trains a privileged teacher on vectorized HD-map and bounding-box inputs, then uses that teacher to supervise pre-trained camera-based students during closed-loop post-training. The paper applies the method to TransFuser and VaVAM in AlpaSim, using 3D Gaussian Splatting reconstructions of real driving logs. Reported driving scores rise by 1.6x and 9.5x, and controlled tests match direct RL closed-loop performance with roughly 1,000x fewer simulator interactions. ArXiv · AI/CL/LG's note
score 5