Megadose Today's AI, without the flood.

Opera: A Verbal Critic Framework for Long-horizon Coding Agents

· HF Daily Papers ·
Opera keeps critic feedback alive until the coding agent actually fixes the diagnosed problem.

The paper frames failed agent feedback as a tracking problem: critics often judge or correct a trajectory, then stop watching. Opera stores each correction as a persistent note, checks it through scheduled and event-driven reviews, and audits whether the feedback is supported by visible evidence before sending it. In reported tests, it raised non-critic agent resolve rates by up to 12.4 points on Terminal-Bench 2.1, 15.0 on a SWE-Bench Pro subset, and 8.9 on DeepSWE v1.1. The authors also say Opera-guided rollouts helped fine-tune Qwen3.5-9B, improving held-out SWE-Bench Pro resolve rate without using a critic at inference time. HF Daily Papers' note

score 5

Categories: Research