Learning to Read the Contextual Tokens in Diffusion Transformers
The paper claims MM-DiT text tokens carry readable information about the image as it forms, even without a prompt.
The authors train a small bottleneck network to map intermediate contextual tokens into a frozen LLM’s input space, so the LLM can answer questions about the emerging image. They report that broad scene semantics appear early in denoising, while finer details become readable later. The same kind of image-specific information remains decodable when the diffusion model is given an empty prompt. They also introduce Contextual Alignment, which they say improves generation quality and distributional coverage by reinforcing that visual-semantic signal. ArXiv · AI/CL/LG's note
The authors train a small bottleneck network to map intermediate contextual tokens into a frozen LLM’s input space, so the LLM can answer questions about the emerging image. They report that broad scene semantics appear early in denoising, while finer details become readable later. The same kind of image-specific information remains decodable when the diffusion model is given an empty prompt. They also introduce Contextual Alignment, which they say improves generation quality and distributional coverage by reinforcing that visual-semantic signal. ArXiv · AI/CL/LG's note
score 5