Systematically Exploring the Capabilities of GPT-6 Astra as Embodied Policies
Astra looks useful for robot task decisions, but not yet reliable as a low-level controller.
The paper evaluates GPT-6 Astra across manipulation, navigation, locomotion, and humanoid tasks as an embodied policy. Hybrid setups did better than direct control in several cases, including 48% success on a RoboDojo subset and 38.7% on RoboCasa365. Navigation was stronger, with 92% success on RxR instruction following and 82% on HM3D object search, though search paths took substantial detours. Locomotion remained weak: five sequential attempts on one obstacle course failed to reach the goal, and latency made practical control difficult. HF Daily Papers' note
The paper evaluates GPT-6 Astra across manipulation, navigation, locomotion, and humanoid tasks as an embodied policy. Hybrid setups did better than direct control in several cases, including 48% success on a RoboDojo subset and 38.7% on RoboCasa365. Navigation was stronger, with 92% success on RxR instruction following and 82% on HM3D object search, though search paths took substantial detours. Locomotion remained weak: five sequential attempts on one obstacle course failed to reach the goal, and latency made practical control difficult. HF Daily Papers' note
score 5