Labels Override Definitions in Jev-Style Typed Decision Models
In these typed classifiers, the label often beats the written rule.
The paper tests open-weight Jev-style decision models and finds that predictions mostly follow option labels, not the definitions meant to control them. Removing definitions barely changes one model’s accuracy, while neutral renaming to A/B improves it. The authors trace the failure to prompt rendering: writing options as “label: definition” creates the bias, while using only the definition avoids it. One model that rendered only definitions was unaffected until the label format was introduced. HF Daily Papers' note
The paper tests open-weight Jev-style decision models and finds that predictions mostly follow option labels, not the definitions meant to control them. Removing definitions barely changes one model’s accuracy, while neutral renaming to A/B improves it. The authors trace the failure to prompt rendering: writing options as “label: definition” creates the bias, while using only the definition avoids it. One model that rendered only definitions was unaffected until the label format was introduced. HF Daily Papers' note
score 4