Megadose AI progress, ranked and analyzed.

Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents

· ArXiv · AI/CL/LG ·
The paper argues that token budgets are a poor proxy for what a coding agent actually remembers.

Across 55 archived coding-agent trajectories, the authors found that instructions, artifacts, tool outputs, and agent-generated state behave differently under retention and compression. They tested object-aware compression and retrieval-based memory policies, but calibration gains did not reliably carry over to held-out tasks. Their replay in a real system also found serving limits that nominal context budgets missed. ArXiv · AI/CL/LG's note

score 5

Categories: Research