RetroThinker: Enabling Retrospective Thinking in Speech LLMs
RetroThinker lets a streaming speech model revise its own reasoning while the user is still speaking.
The paper introduces a post-training framework for Moshi that adds self-verification and forward correction to chain-of-thought steps during inference. It combines supervised fine-tuning on retrospective reasoning data with length-based DPO aimed at early, concurrent reasoning. On GSM8K, the authors report an 11% absolute accuracy gain at comparable latency versus non-retrospective baselines. Source: ArXiv · AI/CL/LG's note.
The paper introduces a post-training framework for Moshi that adds self-verification and forward correction to chain-of-thought steps during inference. It combines supervised fine-tuning on retrospective reasoning data with length-based DPO aimed at early, concurrent reasoning. On GSM8K, the authors report an 11% absolute accuracy gain at comparable latency versus non-retrospective baselines. Source: ArXiv · AI/CL/LG's note.
score 5