Megadose Built for builders and researchers.

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

· ArXiv · AI/CL/LG ·
The benchmark turns multi-reference image generation into operator-level tests, then shows where top models break.

TRACE-Bench defines four operations: Anchor, Disentangle, Apply, and Compose. Its roughly 1,600 cases are built from 631 formula templates and about 4,000 reference images, with complexity measured by operator slots from 1 to 8. In tests of nine leading models, the weaker points were disentanglement and attribute binding, not scene-level composition. Even the best model reached only 0.74 on attribute fidelity. ArXiv · AI/CL/LG's note

score 5

Categories: Research