Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
A 35B model trained on execution-grounded program-evolution operators beat its base model by a wide margin on MLE-Bench Lite.
The paper introduces OpenMLE, a full stack for studying recursive self-improvement in machine learning engineering. Frontis-MA1 is post-trained around four operators: Draft, Improve, Debug, and Crossover, then uses them in long-horizon search. On MLE-Bench Lite, it raises Medal Average from 39.39% to 60.61% over the base model, and to 71.21% with OpenMLE-Evo-Max. The authors say the gains also transfer to held-out NatureBench Lite and that they are releasing the model weights and stack. Source: ArXiv · AI/CL/LG's note.
The paper introduces OpenMLE, a full stack for studying recursive self-improvement in machine learning engineering. Frontis-MA1 is post-trained around four operators: Draft, Improve, Debug, and Crossover, then uses them in long-horizon search. On MLE-Bench Lite, it raises Medal Average from 39.39% to 60.61% over the base model, and to 71.21% with OpenMLE-Evo-Max. The authors say the gains also transfer to held-out NatureBench Lite and that they are releasing the model weights and stack. Source: ArXiv · AI/CL/LG's note.
score 6