S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
S3Gym tests whether LLM agents can turn their own trial runs into better future play.
The benchmark splits exploration from held-out evaluation across seven text-based games with executable verifiers. The paper compares raw history in context, score-conditioned summary memory, and parameter training as ways to reuse experience. Results are mixed: summaries help when lessons compress into rules, while raw history can win when success depends on exact state details. Parameter training sometimes improves performance, but can also be unstable or transfer badly across tasks. ArXiv · AI/CL/LG's note
The benchmark splits exploration from held-out evaluation across seven text-based games with executable verifiers. The paper compares raw history in context, score-conditioned summary memory, and parameter training as ways to reuse experience. Results are mixed: summaries help when lessons compress into rules, while raw history can win when success depends on exact state details. Parameter training sometimes improves performance, but can also be unstable or transfer badly across tasks. ArXiv · AI/CL/LG's note
score 5