Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation
The paper claims GUI agents can keep improving after deployment without human-labeled answers.
The proposed loop has the agent explore an unseen interface, have an MLLM-based reflector judge the result, then fold that reflection back into the model through on-policy self-distillation. The authors add contrastive calibration to keep failed explorations from poisoning the supervision signal. Across six benchmarks, they report a 7.4% average accuracy gain over the base model. Code is planned for release. ArXiv · AI/CL/LG's note
The proposed loop has the agent explore an unseen interface, have an MLLM-based reflector judge the result, then fold that reflection back into the model through on-policy self-distillation. The authors add contrastive calibration to keep failed explorations from poisoning the supervision signal. Across six benchmarks, they report a 7.4% average accuracy gain over the base model. Code is planned for release. ArXiv · AI/CL/LG's note
score 5