LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures
LMBuild tests whether agent-designed 3D objects can actually be assembled and used.
The benchmark represents objects as parts, joints, materials, and construction sequences, not just finished geometry. Its evaluation covers structural soundness, functional affordance, design quality, and physical realization. Across 30 systems, the authors found frontier closed-source models had moved past basic soundness and alignment as the main failure points, while function and operability remained harder. Functional specifications improved part completeness, kinematics, and physical operability. Source: HF Daily Papers' note
The benchmark represents objects as parts, joints, materials, and construction sequences, not just finished geometry. Its evaluation covers structural soundness, functional affordance, design quality, and physical realization. Across 30 systems, the authors found frontier closed-source models had moved past basic soundness and alignment as the main failure points, while function and operability remained harder. Functional specifications improved part completeness, kinematics, and physical operability. Source: HF Daily Papers' note
score 4