Chain-of-Experience for Continual LLM Improvement
The paper tests whether LLMs can improve during inference by carrying forward “experience” from iterative feedback.
The authors call the setup Chain-of-Experience, where models build traces from self-feedback or environmental signals such as correctness and coding test pass rates. Across math, coding, and knowledge tasks on eight LLMs, iterative experience beat feedback-free baselines. They report a 5.6% overall improvement and 19% lower API cost, with most gains arriving early in the iterations. HF Daily Papers' note
The authors call the setup Chain-of-Experience, where models build traces from self-feedback or environmental signals such as correctness and coding test pass rates. Across math, coding, and knowledge tasks on eight LLMs, iterative experience beat feedback-free baselines. They report a 5.6% overall improvement and 19% lower API cost, with most gains arriving early in the iterations. HF Daily Papers' note
score 5