Megadose Built for builders and researchers.

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

· HF Daily Papers ·
WEFT trains tool-using agents by evolving the whole interaction setup, not just adding more executable environments.

The paper says reliable gains require the environment, task, agent harness, and evaluator to work coherently. WEFT uses execution traces and state evidence to locate failures, revise the responsible component, and test changes through fresh rollouts. Its training setup adds prefix-preserving sampling, atomic-turn credit assignment, and MegaMCP for isolated, recoverable concurrent tool use. The authors report WEFT-8B and WEFT-14B beating matched-size environment-scaling baselines, with WEFT-14B ahead of Agent-World-14B by 6.41, 2.23, and 12.27 points on three named benchmarks. HF Daily Papers' note

score 5

Categories: Research