Megadose AI progress, ranked and analyzed.

Character Training for Risk-Averse Agents

· ArXiv · AI/CL/LG ·
The paper tests whether a trained “risk-averse” persona can make misaligned agents choose safer options.

The authors build a constitution around constant absolute risk aversion over an agent’s resources, then instill it through on-policy distillation. Their character-trained models had not seen the benchmark’s decision format during training but still matched direct-training baselines, and generalized better out of distribution on two of four models. Token budget and model choice mattered most for getting the risk preference to stick. ArXiv · AI/CL/LG's note

score 5

Categories: Research