WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
WithEveryone tackles the failure point where group generation loses track of who is supposed to be who.
The paper describes a framework for generating images with up to ten reference identities by assigning each person an addressed token and a planned identity-layout position. Its training loss uses annotated face regions to supervise the intended identity directly, instead of relying on unstable face matching after prediction. On the authors’ identity-disjoint benchmark, it reports higher face similarity than GPT-Image-2, fewer copy-paste artifacts, 97.3% identity coverage, and a 2.8% duplicate rate. Code is listed as forthcoming. HF Daily Papers' note
The paper describes a framework for generating images with up to ten reference identities by assigning each person an addressed token and a planned identity-layout position. Its training loss uses annotated face regions to supervise the intended identity directly, instead of relying on unstable face matching after prediction. On the authors’ identity-disjoint benchmark, it reports higher face similarity than GPT-Image-2, fewer copy-paste artifacts, 97.3% identity coverage, and a 2.8% duplicate rate. Code is listed as forthcoming. HF Daily Papers' note
score 5