PhantomEnvironments: Training LLM Agents in Fictional Worlds
Rule-built fictional corpora trained search agents that transferred to real multi-hop benchmarks.
The paper says PhantomEnvironments uses templated fictional worlds, not LLM-generated facts, to create cheap multi-turn RL settings with verifiable rewards. Agents trained there improved on real-world multi-hop search tasks, in some cases beating real-world training data on newer benchmarks. The authors report generalization to unseen fictional universes and say Qwen models increased search budget roughly with question difficulty. Their ablation points to hop count as the main driver of transfer.
ArXiv · AI/CL/LG's note
The paper says PhantomEnvironments uses templated fictional worlds, not LLM-generated facts, to create cheap multi-turn RL settings with verifiable rewards. Agents trained there improved on real-world multi-hop search tasks, in some cases beating real-world training data on newer benchmarks. The authors report generalization to unseen fictional universes and say Qwen models increased search budget roughly with question difficulty. Their ablation points to hop count as the main driver of transfer.
ArXiv · AI/CL/LG's note
score 6