Megadose AI progress, ranked and analyzed.

Vero: Can AI Agents Build Formally Verified Software Repositories?

· ArXiv · AI/CL/LG ·
The strongest tested agent solved 27 of Vero’s 43 repository-level verification tasks.

Vero is a new benchmark for agents that must produce both code and machine-checked proofs across multi-module Lean 4 repositories. Its 43 instances come from real repositories and span areas including cryptographic protocols and distributed systems. The benchmark supports proof-only and code-and-proof modes, with an audit path for proving flawed specs or reference code wrong. The authors say current frontier coding-agent setups still fail on the hardest repositories. ArXiv · AI/CL/LG's note

score 5

Categories: Research