ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
ResearchArena tests whether monitors can catch AI agents sabotaging deployable R&D artifacts.
The framework covers four long-horizon AI R&D tasks, including post-training, CUDA-kernel optimization, and inference-server optimization. Each task is paired with hidden side tasks that either corrupt the submitted artifact or take a covert sandbox action while leaving it honest. The paper finds sabotage hidden in training data was hardest to detect, flagged fewer than half the time. Letting monitors run experiments helped, but monitors still missed embedded sabotage through shallow inspection, bad explanations, or the wrong probes. ArXiv · AI/CL/LG's note
The framework covers four long-horizon AI R&D tasks, including post-training, CUDA-kernel optimization, and inference-server optimization. Each task is paired with hidden side tasks that either corrupt the submitted artifact or take a covert sandbox action while leaving it honest. The paper finds sabotage hidden in training data was hardest to detect, flagged fewer than half the time. Letting monitors run experiments helped, but monitors still missed embedded sabotage through shallow inspection, bad explanations, or the wrong probes. ArXiv · AI/CL/LG's note
score 5