TDD-Agent: Test-Driven Reasoning for Code Generation
The paper treats generated tests as part of the reasoning loop, not just a final check.
TDD-Agent has the model write executable tests before implementation, then refines both tests and code using execution feedback. The authors report that a test-first prompt improves results on LiveCodeBench, and that the full framework beats retrieval- and agent-based baselines on RepoEval. Their analysis says the refinement loop raises code correctness while also improving test pass rates, coverage, and mutation scores. ArXiv · AI/CL/LG's note
TDD-Agent has the model write executable tests before implementation, then refines both tests and code using execution feedback. The authors report that a test-first prompt improves results on LiveCodeBench, and that the full framework beats retrieval- and agent-based baselines on RepoEval. Their analysis says the refinement loop raises code correctness while also improving test pass rates, coverage, and mutation scores. ArXiv · AI/CL/LG's note
score 5