EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning
EgoLAP turns human first-person activity into language-level motion intent for robot control.
The paper frames raw human trajectories as a bad direct target because human and robot bodies act differently. Its pretraining setup uses structured language actions and motion-level reasoning tied to geometry, physics, and object affordances. In the reported real-world and simulated tests, EgoLAP reached 80.1% mean real-world task progress, a 2.3x gain over alternative action representations. ArXiv · AI/CL/LG's note
The paper frames raw human trajectories as a bad direct target because human and robot bodies act differently. Its pretraining setup uses structured language actions and motion-level reasoning tied to geometry, physics, and object affordances. In the reported real-world and simulated tests, EgoLAP reached 80.1% mean real-world task progress, a 2.3x gain over alternative action representations. ArXiv · AI/CL/LG's note
score 6