3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
3DZip cuts 3D VLM scene tokens to 128 while keeping 94.7% of original QA performance.
The paper targets the token overload created when 3D vision-language models project 2D features into world coordinates. Its compression pipeline uses voxelization, feature-diversity anchor selection, and spatially constrained merging to preserve geometric structure. Across three 3D question-answering benchmarks, it beats existing compression methods and reports 1.92x faster inference. Accepted to ECCV 2026. HF Daily Papers' note
The paper targets the token overload created when 3D vision-language models project 2D features into world coordinates. Its compression pipeline uses voxelization, feature-diversity anchor selection, and spatially constrained merging to preserve geometric structure. Across three 3D question-answering benchmarks, it beats existing compression methods and reports 1.92x faster inference. Accepted to ECCV 2026. HF Daily Papers' note
score 4