BrickBench: Evaluating Agentic Brick Design
Agents can make LEGO-like builds that pass checks, but not yet match human design quality.
BrickBench tests text-conditioned agents on assemblies that must satisfy the prompt and remain physically buildable. The benchmark scores validity, alignment, and design across three settings with different scale and part-library limits. The authors also provide BrickAgent, an environment for coding agents to construct, inspect, and validate designs. Leading agents meet many verifiable physical and semantic requirements, but still trail human designs. HF Daily Papers' note
BrickBench tests text-conditioned agents on assemblies that must satisfy the prompt and remain physically buildable. The benchmark scores validity, alignment, and design across three settings with different scale and part-library limits. The authors also provide BrickAgent, an environment for coding agents to construct, inspect, and validate designs. Leading agents meet many verifiable physical and semantic requirements, but still trail human designs. HF Daily Papers' note
score 4