Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation
A simulated persona panel matched human GenUI ratings far better than a single LLM judge.
The paper introduces ESPP, where psychologically diverse, evidence-grounded personas rate a generated UI screenshot, exchange opinions, and are socially weighted into one judgment. In the authors’ results, ESPP raised Pearson correlation with human judgment from `0.716` to `0.922`. A prompt-ensemble control recovered only about a third of that gap, which the paper treats as evidence that persona and grounding drive most of the gain. Keeping individual persona scores also exposed subgroup disagreements that a single judge would hide. ArXiv · AI/CL/LG's note
The paper introduces ESPP, where psychologically diverse, evidence-grounded personas rate a generated UI screenshot, exchange opinions, and are socially weighted into one judgment. In the authors’ results, ESPP raised Pearson correlation with human judgment from `0.716` to `0.922`. A prompt-ensemble control recovered only about a third of that gap, which the paper treats as evidence that persona and grounding drive most of the gain. Keeping individual persona scores also exposed subgroup disagreements that a single judge would hide. ArXiv · AI/CL/LG's note
score 4