Recursive Video In-Context Learning for Agentic Robot
The paper’s key move is to make a robot agent fetch only the demonstration video detail it needs at each step.
RV-ICL converts one task demo into a hierarchy of sub-events, from coarse keyframes down to short clips around actions like grasps and releases. The agent reads broad structure before planning, then re-enters the hierarchy during execution when a sub-goal needs finer visual detail. The authors report gains over RPent on LIBERO-PRO and LIBERO-Plus without training the method. ArXiv · AI/CL/LG's note
RV-ICL converts one task demo into a hierarchy of sub-events, from coarse keyframes down to short clips around actions like grasps and releases. The agent reads broad structure before planning, then re-enters the hierarchy during execution when a sub-goal needs finer visual detail. The authors report gains over RPent on LIBERO-PRO and LIBERO-Plus without training the method. ArXiv · AI/CL/LG's note
score 5