Megadose Built for builders and researchers.

Reinforcement Learning with Conformal Action Sets: An Application to Sequential Recommendation

· ArXiv · AI/CL/LG ·
The paper proposes RLCP, a recommender method that changes how many actions it keeps during a session instead of using a fixed slate.

It prunes actions with critic scores and an online threshold updated from binary feedback about whether the retained set still covers a proxy target. The authors prove a bound on the observed proxy miss rate and decompose reward loss into filtering and selection losses. On KuaiRand-Pure and MovieLens 1M, at least one RLCP variant produced the highest catalog diversity in all 19 configurations, while keeping session depth competitive and retained sets no larger. ArXiv · AI/CL/LG's note

score 4

Categories: Research