LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
LLaDA-UI applies block-wise diffusion decoding to a 16.7B-parameter vision-language agent for GUI control.
The paper frames GUI agents as a latency-sensitive test for diffusion LLMs because they need to read screens and produce grounded actions repeatedly. LLaDA-UI pairs a native-resolution vision encoder with the LLaDA2.0-mini-base diffusion language backbone, then fine-tunes on mobile, desktop, web, and grounding data. The authors report that it beats Qwen2.5-VL-7B and surpasses Qwen3-VL-8B on four of six GUI benchmarks. Source: HF Daily Papers' note.
The paper frames GUI agents as a latency-sensitive test for diffusion LLMs because they need to read screens and produce grounded actions repeatedly. LLaDA-UI pairs a native-resolution vision encoder with the LLaDA2.0-mini-base diffusion language backbone, then fine-tunes on mobile, desktop, web, and grounding data. The authors report that it beats Qwen2.5-VL-7B and surpasses Qwen3-VL-8B on four of six GUI benchmarks. Source: HF Daily Papers' note.
score 6