Same Feedback, Different Answer: Measuring Run-to-Run Instability in Frontier-Model Customer Feedback Analysis
Taxonomy grounding sharply reduced repeat-run drift in customer-feedback analysis.
The paper tests eight frontier models on recurring feedback-analysis tasks and measures whether categories and counts change when the evidence stays fixed. Its framework tracks “theme churn” and “volume disagreement” across repeated runs. With Claude Opus 4.8 on a fixed 1,000-record corpus, the taxonomy-grounded agent cut theme churn by 86-88% versus raw generation and hierarchical decomposition, with zero disagreement on matched-theme volumes. The authors say the same repeatability problem applies beyond feedback work to repeated synthesis of unstructured corpora.
ArXiv · AI/CL/LG's note
The paper tests eight frontier models on recurring feedback-analysis tasks and measures whether categories and counts change when the evidence stays fixed. Its framework tracks “theme churn” and “volume disagreement” across repeated runs. With Claude Opus 4.8 on a fixed 1,000-record corpus, the taxonomy-grounded agent cut theme churn by 86-88% versus raw generation and hierarchical decomposition, with zero disagreement on matched-theme volumes. The authors say the same repeatability problem applies beyond feedback work to repeated synthesis of unstructured corpora.
ArXiv · AI/CL/LG's note
score 5