Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
Strong LLMs still lose narrative consistency when users push against the plot.
The paper introduces NCP-Bench, a 100-environment benchmark built from movie synopses to test whether narrator and player agents preserve stated narrative commitments over long interactions. Its checks track trajectories, commitments, and initial facts during play. In the reported experiments, fluent output did not mean logical consistency: even GPT-5.2 reached only a 42% survival rate after 20 turns. Fact conflicts ran from 40% to 68% across models, and fully satisfying all achievement commitments within 100 turns was rare. HF Daily Papers' note
The paper introduces NCP-Bench, a 100-environment benchmark built from movie synopses to test whether narrator and player agents preserve stated narrative commitments over long interactions. Its checks track trajectories, commitments, and initial facts during play. In the reported experiments, fluent output did not mean logical consistency: even GPT-5.2 reached only a 42% survival rate after 20 turns. Fact conflicts ran from 40% to 68% across models, and fully satisfying all achievement commitments within 100 turns was rare. HF Daily Papers' note
score 4