Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
PALATE tests role-playing agents with simulated users and personalized satisfaction rubrics instead of fixed dialogue continuations.
The paper argues that borrowed conversation histories distort evaluation because the agent is judged on a trajectory it did not help create. PALATE uses five per-user simulators across 300 character profiles to run free-form multi-turn sessions with candidate agents. Its personalized rubrics agreed more closely with human judgments than a general quality rubric on held-out annotations. The benchmark evaluated 16 candidates and reports results by user-agent pair, not as a single user-independent ranking. HF Daily Papers' note
The paper argues that borrowed conversation histories distort evaluation because the agent is judged on a trajectory it did not help create. PALATE uses five per-user simulators across 300 character profiles to run free-form multi-turn sessions with candidate agents. Its personalized rubrics agreed more closely with human judgments than a general quality rubric on held-out annotations. The benchmark evaluated 16 candidates and reports results by user-agent pair, not as a single user-independent ranking. HF Daily Papers' note
score 4