Rethinking Personalized Generation: Test-Time Alignment via Factorized Ranking Models
A small personalized ranker beats much larger reward models by treating generation as candidate selection.
The paper argues personalized generation has unused headroom at test time, especially with Best-of-N sampling. Its proposed MLP ranking model reuses the base generator’s embeddings to score large candidate pools with low overhead. Across nine datasets, the authors report it outperforms billion-parameter general reward models while using under 0.4% of their parameters and far lower scoring latency. ArXiv · AI/CL/LG's note
The paper argues personalized generation has unused headroom at test time, especially with Best-of-N sampling. Its proposed MLP ranking model reuses the base generator’s embeddings to score large candidate pools with low overhead. Across nine datasets, the authors report it outperforms billion-parameter general reward models while using under 0.4% of their parameters and far lower scoring latency. ArXiv · AI/CL/LG's note
score 5