Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation
The paper cuts VLA denoising from 10 steps to 2, dropping model inference time from 61.557 ms to 21.956 ms.
The authors argue that real-time robot use is limited by both model latency and robot-side delays. Their analysis finds early flow-matching steps are comparatively stable, while later steps handle sharper directional correction, leading them to use non-uniform two-stage denoising. In a physical garment-folding evaluation with π0.5 as baseline, Legato led training-based methods and Temporal Smoothing led training-free methods. Pairing the two-step denoising with execution methods reduced inference cost with only a small task-performance loss. HF Daily Papers' note
The authors argue that real-time robot use is limited by both model latency and robot-side delays. Their analysis finds early flow-matching steps are comparatively stable, while later steps handle sharper directional correction, leading them to use non-uniform two-stage denoising. In a physical garment-folding evaluation with π0.5 as baseline, Legato led training-based methods and Temporal Smoothing led training-free methods. Pairing the two-step denoising with execution methods reduced inference cost with only a small task-performance loss. HF Daily Papers' note
score 5