Discrete Beckmann Transport Models for One-Step Language Modeling and Reasoning
The paper claims a teacher-free transport map can drive token generation to a simplex fixed point in one step.
DBTM replaces distilled many-step sampling with a time-independent flow trained directly from data. The authors frame its fixed-point behavior through a conservation-equation residual, avoiding teacher flows and time conditioning. A partially trained map can be iterated as refinement, with extra evaluations improving outputs rather than serving as ODE steps. They report better one- and few-step language modeling and reasoning results than discrete diffusion and continuous flow baselines. ArXiv · AI/CL/LG's note
DBTM replaces distilled many-step sampling with a time-independent flow trained directly from data. The authors frame its fixed-point behavior through a conservation-equation residual, avoiding teacher flows and time conditioning. A partially trained map can be iterated as refinement, with extra evaluations improving outputs rather than serving as ODE steps. They report better one- and few-step language modeling and reasoning results than discrete diffusion and continuous flow baselines. ArXiv · AI/CL/LG's note
score 5