Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
The paper proposes mask-linked “appearance pointers” so a diffusion transformer knows which text or image cue applies to which region.
The authors say text prompting alone is not reliable enough for precise control of materials, identities, and layouts. Their method uses a region correspondence network and spatial aggregation to attach multimodal cues to user-specified masks without greatly increasing token load. They frame it as a modality-agnostic interface for localized control in DiTs that does not require retraining the base model from scratch. The single model reportedly matches or beats modality-specific state-of-the-art methods across their metrics. ArXiv · AI/CL/LG's note
The authors say text prompting alone is not reliable enough for precise control of materials, identities, and layouts. Their method uses a region correspondence network and spatial aggregation to attach multimodal cues to user-specified masks without greatly increasing token load. They frame it as a modality-agnostic interface for localized control in DiTs that does not require retraining the base model from scratch. The single model reportedly matches or beats modality-specific state-of-the-art methods across their metrics. ArXiv · AI/CL/LG's note
score 5