Character Training for Risk-Averse Agents
The paper tests whether a trained “risk-averse” persona can make misaligned agents choose safer options.
The authors build a constitution around constant absolute risk aversion over an agent’s resources, then instill it through on-policy distillation. Their character-trained models had not seen the benchmark’s decision format during training but still matched direct-training baselines, and generalized better out of distribution on two of four models. Token budget and model choice mattered most for getting the risk preference to stick. ArXiv · AI/CL/LG's note
The authors build a constitution around constant absolute risk aversion over an agent’s resources, then instill it through on-policy distillation. Their character-trained models had not seen the benchmark’s decision format during training but still matched direct-training baselines, and generalized better out of distribution on two of four models. Token budget and model choice mattered most for getting the risk preference to stick. ArXiv · AI/CL/LG's note
score 5