Which Values Do LLMs Confuse? A Schwartz-Based Recognition Study
The models usually found the right value neighborhood, but often swapped Schwartz-adjacent categories.
The study tests 21 instruction-tuned LLM runs on 1,000 Russian situational texts labeled across Schwartz’s ten basic values. Across the reliable runs, top-1 accuracy was 0.683, while top-3 accuracy reached 0.892. More than half of semantic errors involved adjacent values, with recurring confusions such as Universalism to Benevolence, Tradition to Conformity, and Security to Power. The authors argue that value evaluations should track not just exact hits, but ranked recovery and directed error patterns. ArXiv · AI/CL/LG's note
The study tests 21 instruction-tuned LLM runs on 1,000 Russian situational texts labeled across Schwartz’s ten basic values. Across the reliable runs, top-1 accuracy was 0.683, while top-3 accuracy reached 0.892. More than half of semantic errors involved adjacent values, with recurring confusions such as Universalism to Benevolence, Tradition to Conformity, and Security to Power. The authors argue that value evaluations should track not just exact hits, but ranked recovery and directed error patterns. ArXiv · AI/CL/LG's note
score 4