Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition
Qwen3-32B’s time preferences can be pushed toward shorter or longer horizons by adding a learned activation direction.
The paper finds a short-term versus long-term direction in the model’s residual stream using contrastive linear probes. Applying that direction changes answers on held-out temporal-choice tests and on a monetary delay task. The steering moves the model’s indifference point between smaller-sooner and larger-later rewards in both directions. Moderate steering also improves a planning-related score on TravelPlanner. ArXiv · AI/CL/LG's note
The paper finds a short-term versus long-term direction in the model’s residual stream using contrastive linear probes. Applying that direction changes answers on held-out temporal-choice tests and on a monetary delay task. The steering moves the model’s indifference point between smaller-sooner and larger-later rewards in both directions. Moderate steering also improves a planning-related score on TravelPlanner. ArXiv · AI/CL/LG's note
score 5