Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation
ReflexBench tests robot models on manipulation tasks where reaction time matters.
The paper says current VLA benchmarks lean on static tasks and miss dynamic interaction settings. Its ReflexBench benchmark has six dynamic tasks and can vary latency under synchronous or asynchronous inference. The authors also introduce ReflexVLA, using future prediction, temporal fusion, batched visual encoding, and CUDA Graph replay to cut deployment delay. Experiments report better dynamic manipulation performance, competitive static-benchmark accuracy, and real-world validation. ArXiv · AI/CL/LG's note
The paper says current VLA benchmarks lean on static tasks and miss dynamic interaction settings. Its ReflexBench benchmark has six dynamic tasks and can vary latency under synchronous or asynchronous inference. The authors also introduce ReflexVLA, using future prediction, temporal fusion, batched visual encoding, and CUDA Graph replay to cut deployment delay. Experiments report better dynamic manipulation performance, competitive static-benchmark accuracy, and real-world validation. ArXiv · AI/CL/LG's note
score 6