Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It
Changing option names alone made typed decision models reverse their rankings while still producing valid outputs.
The paper tests Jev and related open-weight models by keeping the task, rubrics, and option set fixed, then swapping which names attach to which rubrics. Moving from neutral labels to semantically loaded `no/yes` labels caused large answer flips, including an AUC drop from .94 to .23 on 1,200 workflow decisions. Random character labels returned the models to the neutral-control regime, pointing to the polarity of the option names rather than renaming itself. The type-error rate stayed at 0%, even when accuracy degraded sharply. ArXiv · AI/CL/LG's note
The paper tests Jev and related open-weight models by keeping the task, rubrics, and option set fixed, then swapping which names attach to which rubrics. Moving from neutral labels to semantically loaded `no/yes` labels caused large answer flips, including an AUC drop from .94 to .23 on 1,200 workflow decisions. Random character labels returned the models to the neutral-control regime, pointing to the polarity of the option names rather than renaming itself. The type-error rate stayed at 0%, even when accuracy degraded sharply. ArXiv · AI/CL/LG's note
score 4