Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
The paper adds object-level geometric attention to DP3 so a robot policy can stay locked on the target in cluttered 3D scenes.
Attention-DP3 uses open-vocabulary 2D segmentation on RGB images, lifts the target masks into 3D with calibrated camera geometry, and feeds those object-centric cues into the existing DP3 diffusion backbone. Its conditioning separates targetness, within-target saliency, and background suppression. The authors report consistent gains on Adroit, DexArt, MetaWorld, and a real-world SO101 setup, with up to a 31% advantage over DP3 under heavy clutter. HF Daily Papers' note
Attention-DP3 uses open-vocabulary 2D segmentation on RGB images, lifts the target masks into 3D with calibrated camera geometry, and feeds those object-centric cues into the existing DP3 diffusion backbone. Its conditioning separates targetness, within-target saliency, and background suppression. The authors report consistent gains on Adroit, DexArt, MetaWorld, and a real-world SO101 setup, with up to a 31% advantage over DP3 under heavy clutter. HF Daily Papers' note
score 4