Megadose Built for builders and researchers.

TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

· HF Daily Papers ·
TRACE-Bench tests multi-reference image generation by breaking prompts into four atomic operations and scoring where models fail.

The benchmark covers about 1,600 cases built from 631 formula templates and roughly 4,000 reference images. Its operators are Anchor, Disentangle, Apply, and Compose, with prompt complexity measured by operator slots. In tests of nine leading models, the paper says the main weaknesses were disentanglement and attribute binding, not scene-level composition. Even the best model scored 0.74 on attribute fidelity. HF Daily Papers' note

score 5

Categories: Research