Megadose AI progress, ranked and analyzed.

Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models

· HF Daily Papers ·
The benchmark finds text-only LLMs can code given geometry, but diverge when they have to invent the layout.

AM-Bench splits the problem into translating a specified picture description into code and composing a picture from a looser prompt. Across eight open-weight text-and-code-only models, translation was reliable, while layout quality varied enough to suggest code generation was not the limiting factor. The paper also reports that raw SVG improved layout scores over procedural code for every model tested. Activation probes suggest models form a coarse prompt-driven layout before generation, then track geometry as output unfolds. HF Daily Papers' note

score 4

Categories: Research