S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
The benchmark tests whether agents can turn their own trial-and-error logs into better future play.
S3Gym frames self-improvement as self-testing, self-judging, and self-improvement across seven text-based games with executable verifiers. The paper finds gains are uneven: raw history, compressed summary memory, and parameter training each help in some settings and fail in others. Summaries work when experience becomes reusable strategy, while raw history can matter more for state-specific decisions. Training can improve performance sharply, but the authors also report instability and negative transfer. HF Daily Papers' note
S3Gym frames self-improvement as self-testing, self-judging, and self-improvement across seven text-based games with executable verifiers. The paper finds gains are uneven: raw history, compressed summary memory, and parameter training each help in some settings and fail in others. Summaries work when experience becomes reusable strategy, while raw history can matter more for state-specific decisions. Training can improve performance sharply, but the authors also report instability and negative transfer. HF Daily Papers' note
score 5