Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer
Better compact docs made faithful summaries, but they did not make coding agents fix more issues.
The paper introduces a roundtrip benchmark that judges descriptions by whether regenerated code still passes the original tests. That signal produced a prompt that reached full fidelity and generalized to unseen files. But across two model families and ten repositories, neither static compact documentation nor retrieved context beat giving the agent the issue alone when source code was available. ArXiv · AI/CL/LG's note
The paper introduces a roundtrip benchmark that judges descriptions by whether regenerated code still passes the original tests. That signal produced a prompt that reached full fidelity and generalized to unseen files. But across two model families and ten repositories, neither static compact documentation nor retrieved context beat giving the agent the issue alone when source code was available. ArXiv · AI/CL/LG's note
score 5