ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders
The benchmark starts agents from fuzzy product requirements, then tests whether they can build the intended project through interaction.
ICAE-Bench uses an automated User Agent to reveal hidden constraints during the task. Its ambiguities are grounded in real open-source repositories with executable behavior, rather than unconstrained prompts. Evaluation combines black-box tests with diagnostics for correctness, API and semantic similarity, structure, design quality, and interaction quality. HF Daily Papers' note
ICAE-Bench uses an automated User Agent to reveal hidden constraints during the task. Its ambiguities are grounded in real open-source repositories with executable behavior, rather than unconstrained prompts. Evaluation combines black-box tests with diagnostics for correctness, API and semantic similarity, structure, design quality, and interaction quality. HF Daily Papers' note
score 5