Megadose AI progress, ranked and analyzed.

ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders

· HF Daily Papers ·
The benchmark starts agents from fuzzy product requirements, then tests whether they can build the intended project through interaction.

ICAE-Bench uses an automated User Agent to reveal hidden constraints during the task. Its ambiguities are grounded in real open-source repositories with executable behavior, rather than unconstrained prompts. Evaluation combines black-box tests with diagnostics for correctness, API and semantic similarity, structure, design quality, and interaction quality. HF Daily Papers' note

score 5

Categories: Research