Megadose Built for builders and researchers.

Q-Learning with Scalar Adjoint Matching

· HF Daily Papers ·
SQAM replaces costly per-step adjoint updates with a scalar shortcut for fine-tuning flow policies.

The paper argues that pretrained flow policies have batch-averaged velocity Jacobians that mostly concentrate on the diagonal. From that, the authors derive a scalar adjoint that scales the final action’s value gradient by flow time, avoiding vector-Jacobian products at every flow step. They add a value penalty on policy-generated actions, saying critic control there is especially important under the scalar adjoint. SQAM beats the strongest baseline by 18 to 35 points on the four hardest OGBench domains and improves over supervised fine-tuning on three real bimanual robot tasks. HF Daily Papers' note

score 4

Categories: Research