ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments
ScienceIDE turns scientific repositories into executable training environments, then uses verified agent trajectories to train new models.
The paper frames this as a “scientific experience bottleneck”: useful scientific code is hard for agents to learn from because tools, conventions, and correctness tests are scattered. ScienceIDE uses expert-defined cases and acceptance criteria to convert repositories into environments for task generation, execution, and verification. The authors trained PhAI-IDE models at 72B, 9B, and 4B parameters, reporting gains on held-out scientific-code repair and selected broader benchmarks. ArXiv · AI/CL/LG's note
The paper frames this as a “scientific experience bottleneck”: useful scientific code is hard for agents to learn from because tools, conventions, and correctness tests are scattered. ScienceIDE uses expert-defined cases and acceptance criteria to convert repositories into environments for task generation, execution, and verification. The authors trained PhAI-IDE models at 72B, 9B, and 4B parameters, reporting gains on held-out scientific-code repair and selected broader benchmarks. ArXiv · AI/CL/LG's note
score 6