GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments
GUI-CC tests whether generated mobile UI states can hold together across multi-step agent use.
The benchmark shifts evaluation from one-screen prediction to world models used as agent environments. It includes an offline track over real GUIOdyssey trajectories and an online loop with probing agents across 30 mobile apps. The paper reports that current models can make plausible-looking screens while losing task-relevant context or failing to support executable rollouts. HF Daily Papers' note
The benchmark shifts evaluation from one-screen prediction to world models used as agent environments. It includes an offline track over real GUIOdyssey trajectories and an online loop with probing agents across 30 mobile apps. The paper reports that current models can make plausible-looking screens while losing task-relevant context or failing to support executable rollouts. HF Daily Papers' note
score 5