ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents
ComputerSD trains computer-use agents with feedback from each executed GUI step, not just the final task result.
The paper says a fine-tuned GUI analyzer turns real-time screen transitions into guidance and a step-level value score after every action. That signal is used to regulate token-level self-distillation while the system also optimizes trajectory-level GRPO. On OSWorld-Verified, the method beats outcome-only GRPO by 1.9 points on Qwen3-VL-8B-Thinking and 4.1 points on EvoCUA-8B. The authors also report out-of-distribution evaluations as supporting generalization. ArXiv · AI/CL/LG's note
The paper says a fine-tuned GUI analyzer turns real-time screen transitions into guidance and a step-level value score after every action. That signal is used to regulate token-level self-distillation while the system also optimizes trajectory-level GRPO. On OSWorld-Verified, the method beats outcome-only GRPO by 1.9 points on Qwen3-VL-8B-Thinking and 4.1 points on EvoCUA-8B. The authors also report out-of-distribution evaluations as supporting generalization. ArXiv · AI/CL/LG's note
score 5