Megadose Built for builders and researchers.

From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

· HF Daily Papers ·
PersonTTS searches for controllers that fit a user’s combined accuracy, latency, and cost requirements, then reuses past search work to make that discovery cheaper.

The paper frames personalized test-time scaling as a joint satisfaction problem, not a single accuracy-cost or accuracy-latency tradeoff. Its proposed framework initializes new controller searches from requirement-matched prior experience and adds distilled procedural guidance, while still evaluating candidates on the target profile. On AIME and HMMT, the authors report better satisfaction of unseen user profiles and held-out problems than strong TTS baselines. They also say cross-user reuse improves policy quality under the same candidate-evaluation budget while reducing discovery-agent time and cost. HF Daily Papers' note

score 4

Categories: Research