Megadose Built for builders and researchers.

Predicting Alignment Generalization with Value Representations

· ArXiv · AI/CL/LG ·
Activation patterns predicted how value tuning spills over better than value descriptions did.

The paper defines “alignment generalization prediction”: estimating how fine-tuning for one value changes behavior on other held-out values. Across 66 values from modern alignment targets, activation-based representations reached a 0.45 correlation with the authors’ generalization matrix, versus 0.05 for description-based baselines. The authors also use these representations to compare values inside multi-value targets, finding that similarity is significantly correlated with model robustness. They report early evidence of a shared, model-independent value space and build a taxonomy from observed generalization dynamics. ArXiv · AI/CL/LG's note

score 5

Categories: Research