Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport
OptiFlow trains a one-step offline RL flow policy by using value-weighted optimal transport instead of direct critic maximization.
The paper frames one-step flow policy learning as a state-wise sample-allocation problem. Its method couples samples from a value-aware reference flow policy to a faster one-step policy, prioritizing critic-estimated high-value actions while preserving action-space compatibility. The authors argue this keeps the learned policy near high-value, dataset-supported modes and reduces mode collapse or out-of-distribution drift. They report strong results across diverse offline reinforcement learning benchmarks. ArXiv · AI/CL/LG's note
The paper frames one-step flow policy learning as a state-wise sample-allocation problem. Its method couples samples from a value-aware reference flow policy to a faster one-step policy, prioritizing critic-estimated high-value actions while preserving action-space compatibility. The authors argue this keeps the learned policy near high-value, dataset-supported modes and reduces mode collapse or out-of-distribution drift. They report strong results across diverse offline reinforcement learning benchmarks. ArXiv · AI/CL/LG's note
score 4