Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation
The paper claims an image-generation agent’s gains can be partly baked back into the diffusion model itself.
D-OPCD uses the agent-improved prompt as privileged context during distillation, so the model can recover some harness benefit from the original query alone. In the authors’ tests, direct-generation score rose from 60.52 to 65.09 across four benchmarks. After that transfer, their Auto Skill Evolver ran again on the updated generator and added 1.83 points over a skill-free harness. Source: HF Daily Papers' note.
D-OPCD uses the agent-improved prompt as privileged context during distillation, so the model can recover some harness benefit from the original query alone. In the authors’ tests, direct-generation score rose from 60.52 to 65.09 across four benchmarks. After that transfer, their Auto Skill Evolver ran again on the updated generator and added 1.83 points over a skill-free harness. Source: HF Daily Papers' note.
score 4