Megadose AI progress, ranked and analyzed.

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

· HF Daily Papers ·
The benchmark tests whether self-evolving agents can repair working code after the physics changes under them.

PACE-Bench contains 144 source-to-target adaptation pairs across six physics domains, keeping the goal and interface fixed while mutating the environment. Agents get diagnostic sandbox feedback and a limited attempt budget to evolve source code that worked before but fails in the target. The reported results leave the benchmark unsaturated: Reflexion + Qwen3-14B solves 35.9% of full pairs, while GPT-5.5 reaches 66.7% on the Statics subset. The authors argue the hard part is redesigning mechanisms, not merely identifying changed parameters. Source: HF Daily Papers' note.

score 5

Categories: Research