Megadose AI progress, ranked and analyzed.

FoldQuantVLA: Native Low-Bit Quantization of Vision-Language-Action Models via Consistent Folding

· ArXiv · AI/CL/LG ·
The paper reports four-bit robot-policy inference with modest speedups and a measured recovery in real-robot success when selected layers stay at eight bits.

FoldQuantVLA is a post-training quantization method for vision-language-action models, keeping a consistent activation format through calibration, rounding, and integer execution. The authors say W4A4 runs 1.20–1.33x faster than floating-point TensorRT on Jetson AGX Orin and 1.25–1.52x faster on desktop. Holding language attention-output and feed-forward down projections at W8A8 improved held-out action fidelity across four checkpoints. In four real-robot tasks, that mixed setting lifted GR00T N1.7 success from 80.0% under uniform W4A4 to 92.5%, with 1 ms more Orin latency. ArXiv · AI/CL/LG's note

score 5

Categories: Research