Megadose AI progress, ranked and analyzed.

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

· HF Daily Papers ·
The benchmark finds a sharp failure point for VLMs when jigsaw geometry scales beyond small grids.

JigShape uses interlocking tab-and-blank pieces to make puzzle placement less ambiguous than rectangular-cut benchmarks. Across 95K puzzle instances, zero-shot VLMs mostly perform at chance, with only GPT-5.5 beating random on 4x4 puzzles. Fine-tuning reaches above 97% on 4x4, but performance falls apart on denser grids, including below 5% for fine-tuned models on 12x12. The authors frame this as evidence that current systems struggle to maintain geometric constraint satisfaction as puzzle size grows. HF Daily Papers' note

score 5

Categories: Research