Twin: Playing an Unknown Game with a Test-Time Digital Twin
A test-time world model lifted the same base model from 7.8% to 93.3% on the benchmark.
Twin has a coding agent build an executable model of an unknown grid game from interaction, then requires that model to replay every observed transition before the agent can act. When the model prediction fails, the mismatch becomes a counterexample for repair. The system cleared 179 of 183 levels, and beat human action efficiency on 158 of the 179 it solved. The authors say goal inference, not transition modeling, was the harder part. ArXiv · AI/CL/LG's note
Twin has a coding agent build an executable model of an unknown grid game from interaction, then requires that model to replay every observed transition before the agent can act. When the model prediction fails, the mismatch becomes a counterexample for repair. The system cleared 179 of 183 levels, and beat human action efficiency on 158 of the 179 it solved. The authors say goal inference, not transition modeling, was the harder part. ArXiv · AI/CL/LG's note
score 6