Score-Calibrated Flow for Sampling from Unnormalized Densities with Applications to Generative Online Reinforcement Learning
The paper proposes a way to train flow-based RL policies from critic scores without importance sampling.
Score-Calibrated Flow uses self-consistency conditions to learn a sampler for an unnormalized target density when target samples are unavailable. The authors frame the velocity condition as a fixed-point problem and train it with a stop-gradient objective while keeping the sample-interpolate-regress structure of conditional flow matching. In online RL, the critic gradient supplies the target score at generated actions. The paper reports matching or improving generative-policy baselines on RL benchmarks while cutting training time. ArXiv · AI/CL/LG's note
Score-Calibrated Flow uses self-consistency conditions to learn a sampler for an unnormalized target density when target samples are unavailable. The authors frame the velocity condition as a fixed-point problem and train it with a stop-gradient objective while keeping the sample-interpolate-regress structure of conditional flow matching. In online RL, the critic gradient supplies the target score at generated actions. The paper reports matching or improving generative-policy baselines on RL benchmarks while cutting training time. ArXiv · AI/CL/LG's note
score 4