Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair
The paper says coding agents can find a correct patch and then lose it in later revision.
In HumanEval repair runs, “ever-correct” improved across revisions, but correctness tied to the current code and current traces fell after forced revision. The authors report stale verifier traces caused substantially more harm from correct starts than current traces in a 14B replication. Their response is a typed loop contract that binds evidence to exact code states, preserves verified checkpoints, and records auditable admission receipts. They explicitly frame the implementation as a conformance artifact, not proof of better repair skill. ArXiv · AI/CL/LG's note
In HumanEval repair runs, “ever-correct” improved across revisions, but correctness tied to the current code and current traces fell after forced revision. The authors report stale verifier traces caused substantially more harm from correct starts than current traces in a 14B replication. Their response is a typed loop contract that binds evidence to exact code states, preserves verified checkpoints, and records auditable admission receipts. They explicitly frame the implementation as a conformance artifact, not proof of better repair skill. ArXiv · AI/CL/LG's note
score 5