VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following
VISA turns failed synthetic training samples into feedback for the next round.
The paper proposes a multimodal instruction-data pipeline that keeps memory across synthesis rounds instead of generating and filtering once. It analyzes images for verifiable constraints, generates candidate instructions, checks them with tools and LLM judges, and uses failures to guide recovery. Accepted samples are tested against the target model so later rounds can focus on weaknesses and avoid repeated templates. The authors report gains on MM-IFEval while maintaining performance across seven public multimodal benchmarks. ArXiv · AI/CL/LG's note
The paper proposes a multimodal instruction-data pipeline that keeps memory across synthesis rounds instead of generating and filtering once. It analyzes images for verifiable constraints, generates candidate instructions, checks them with tools and LLM judges, and uses failures to guide recovery. Accepted samples are tested against the target model so later rounds can focus on weaknesses and avoid repeated templates. The authors report gains on MM-IFEval while maintaining performance across seven public multimodal benchmarks. ArXiv · AI/CL/LG's note
score 5