Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments
The paper’s strongest gain comes from checking and correcting actions while the agent is running.
The authors add adaptive episodic memory and bounded self-reflection to SwiftSage, then test the variants on ScienceWorld. The full system posts the best mean final score, success rate, and successful-step efficiency. Self-reflection is the strongest standalone module, suggesting failed or invalid execution is the main bottleneck here. Memory helps more once that runtime loop is steadier. ArXiv · AI/CL/LG's note
The authors add adaptive episodic memory and bounded self-reflection to SwiftSage, then test the variants on ScienceWorld. The full system posts the best mean final score, success rate, and successful-step efficiency. Self-reflection is the strongest standalone module, suggesting failed or invalid execution is the main bottleneck here. Memory helps more once that runtime loop is steadier. ArXiv · AI/CL/LG's note
score 4