Attenuated in-context identification in time-series foundation models: diagnosis under counterfactual inputs and repair by synthetic forced-system fine-tuning
The paper finds that several time-series foundation models understate or miss dynamic input effects in what-if tests.
Using exact counterfactuals from forced engineering systems, the author reports that TimesFM-2.5 and TabPFN-TS behave as memoryless through their default covariate interfaces. Chronos-2 captures dynamics in context, but its estimated effects are attenuated and its impulse response shape is wrong. Adding inference-time context dither lowers what-if error across six synthetic classes, while a short synthetic fine-tune restores much of the response magnitude. Classical identification still beats the fine-tuned model on three of four measured plants, and the fine-tune costs some univariate forecasting skill. ArXiv · AI/CL/LG's note
Using exact counterfactuals from forced engineering systems, the author reports that TimesFM-2.5 and TabPFN-TS behave as memoryless through their default covariate interfaces. Chronos-2 captures dynamics in context, but its estimated effects are attenuated and its impulse response shape is wrong. Adding inference-time context dither lowers what-if error across six synthetic classes, while a short synthetic fine-tune restores much of the response magnitude. Classical identification still beats the fine-tuned model on three of four measured plants, and the fine-tune costs some univariate forecasting skill. ArXiv · AI/CL/LG's note
score 4