Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents
The paper compresses older tool observations into latent tokens while keeping recent observations in text, trading some SWE-bench accuracy for much smaller context.
LOHA keeps the agent’s own turns and the last K observations readable, while moving older tool output into soft-token form. ACD trains the agent to use that latent view while anchoring behavior against the base model on plain-text inputs. With K=3, context per call fell 43% for Qwen3-4B and 57% for SWE-Master-4B-RL, with lower resolve rates than the uncompressed bases. At K=8, performance recovered closer to baseline, and under a 32K-token limit the compressed Qwen3 setup beat the adapted full-text agent on a 199-instance subset. ArXiv · AI/CL/LG's note
LOHA keeps the agent’s own turns and the last K observations readable, while moving older tool output into soft-token form. ACD trains the agent to use that latent view while anchoring behavior against the base model on plain-text inputs. With K=3, context per call fell 43% for Qwen3-4B and 57% for SWE-Master-4B-RL, with lower resolve rates than the uncompressed bases. At K=8, performance recovered closer to baseline, and under a 32K-token limit the compressed Qwen3 setup beat the adapted full-text agent on a 199-instance subset. ArXiv · AI/CL/LG's note
score 5