Aspire: Can Models Self-Evolve from Vague Goals?
ASPIRE tests whether agents can turn a vague capability goal into their own training plan, then survive hidden evaluation.
The benchmark gives only a natural-language goal while keeping 520 expert-written downstream items hidden across six goals. Agents must choose data, update methods, validation signals, and when to evaluate, covering both weight updates and harness edits. The paper reports that current agents can run those loops, but weight-level improvements are sparse and unstable. Self-evaluations often overfit narrow checks, with mismatched training data and later search sometimes wiping out earlier gains. ArXiv · AI/CL/LG's note
The benchmark gives only a natural-language goal while keeping 520 expert-written downstream items hidden across six goals. Agents must choose data, update methods, validation signals, and when to evaluate, covering both weight updates and harness edits. The paper reports that current agents can run those loops, but weight-level improvements are sparse and unstable. Self-evaluations often overfit narrow checks, with mismatched training data and later search sometimes wiping out earlier gains. ArXiv · AI/CL/LG's note
score 5