Megadose AI progress, ranked and analyzed.

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

· ArXiv · AI/CL/LG ·
PACE-Bench tests whether self-evolving agents can repair code after the physics changes underneath it.

The benchmark contains 144 source-to-target adaptation pairs across six physics domains, keeping the goal and interface fixed while mutating the target environment. Agents get sandbox feedback and a limited attempt budget to turn a design that worked in the source into one that works in the target. The reported results leave the benchmark unsolved: Reflexion + Qwen3-14B reaches 35.9% on the full set, while GPT-5.5 solves 66.7% of the Statics subset under the full budget. The authors argue the bottleneck is mechanism redesign, not simply inferring changed parameters. ArXiv · AI/CL/LG's note

score 5

Categories: Research