Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes
Learned soft prefixes can push fixed models away from correct syllogistic judgments at high rates.
The paper tests opaque continuous prefixes on exactly labeled syllogistic reasoning tasks while leaving the underlying models unchanged. Across Qwen3.6-35B-A3B MoE, Qwen3-8B, and Gemma 4 31B, those prefixes redirected many correct answers and generalized across unseen forms and prompt/interface changes. In repeated tests, learned prefixes beat paired random controls in all 16 model-direction-split comparisons by 37 to 99 points. The author argues the main effect is a broad preference for one answer meaning, not a reliable transferable logical operation. ArXiv · AI/CL/LG's note
The paper tests opaque continuous prefixes on exactly labeled syllogistic reasoning tasks while leaving the underlying models unchanged. Across Qwen3.6-35B-A3B MoE, Qwen3-8B, and Gemma 4 31B, those prefixes redirected many correct answers and generalized across unseen forms and prompt/interface changes. In repeated tests, learned prefixes beat paired random controls in all 16 model-direction-split comparisons by 37 to 99 points. The author argues the main effect is a broad preference for one answer meaning, not a reliable transferable logical operation. ArXiv · AI/CL/LG's note
score 4