Collaborative Personalized Preference Alignment for LLMs under Data Deficiency
APO trains shared aligner starting points that can adapt to different user preferences from 20 local examples.
The paper targets personalization when each user has little feedback and users disagree across multiple response objectives. Its method groups users with compatible updates, then balances descent and controlled ascent inside each group to reduce gradient conflict. The resulting initialization is refined through few-shot local adaptation. Experiments on Fed-ChatbotPA and UltraFeedback report consistent gains over existing methods. HF Daily Papers' note
The paper targets personalization when each user has little feedback and users disagree across multiple response objectives. Its method groups users with compatible updates, then balances descent and controlled ascent inside each group to reduce gradient conflict. The resulting initialization is refined through few-shot local adaptation. Experiments on Fed-ChatbotPA and UltraFeedback report consistent gains over existing methods. HF Daily Papers' note
score 4