Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment
The paper says post-training can determine whether a model follows its own moral judgment under pressure.
The study tests 248 scenarios by asking each model both what it would do as an agent and which option it judges right. OLMo-3-7B-Instruct chose actions it had judged wrong in about one-fifth of pressuring cases, more than when the pressure was removed. The effect appeared in OLMo-3 and Llama-3.1-8B-Instruct, but not across the full panels for Tulu 3 or Qwen2.5-7B-Instruct. The author argues the gap is a measurable post-training target, not a fixed property of the base weights. ArXiv · AI/CL/LG's note
The study tests 248 scenarios by asking each model both what it would do as an agent and which option it judges right. OLMo-3-7B-Instruct chose actions it had judged wrong in about one-fifth of pressuring cases, more than when the pressure was removed. The effect appeared in OLMo-3 and Llama-3.1-8B-Instruct, but not across the full panels for Tulu 3 or Qwen2.5-7B-Instruct. The author argues the gap is a measurable post-training target, not a fixed property of the base weights. ArXiv · AI/CL/LG's note
score 4