JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
The benchmark finds a sharp failure point for VLMs when jigsaw geometry scales beyond small grids.
JigShape uses interlocking tab-and-blank pieces to make puzzle placement less ambiguous than rectangular-cut benchmarks. Across 95K puzzle instances, zero-shot VLMs mostly perform at chance, with only GPT-5.5 beating random on 4x4 puzzles. Fine-tuning reaches above 97% on 4x4, but performance falls apart on denser grids, including below 5% for fine-tuned models on 12x12. The authors frame this as evidence that current systems struggle to maintain geometric constraint satisfaction as puzzle size grows. HF Daily Papers' note
JigShape uses interlocking tab-and-blank pieces to make puzzle placement less ambiguous than rectangular-cut benchmarks. Across 95K puzzle instances, zero-shot VLMs mostly perform at chance, with only GPT-5.5 beating random on 4x4 puzzles. Fine-tuning reaches above 97% on 4x4, but performance falls apart on denser grids, including below 5% for fine-tuned models on 12x12. The authors frame this as evidence that current systems struggle to maintain geometric constraint satisfaction as puzzle size grows. HF Daily Papers' note
score 5