Megadose AI progress, ranked and analyzed.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

· HF Daily Papers ·
Tencent’s benchmark is built to stay searchable in public without making its task prompts easy to recover.

WorkBuddy Bench covers coding agents across Code, Web, Office, and Security tasks, each packaged for reproducible runs. Its tasks are reverse-engineered from real commits, pull requests, or business scenarios, then rewritten as short role-played requests. The release includes task directories, images, harnesses, tests, and reference solutions, so the authors lean on construction and versioning rather than secrecy for contamination resistance. The paper reports leaderboards across model families, but no overall average because each subset uses a different scoring method. HF Daily Papers' note

score 5

Categories: Research