WaveTLM: Reliable Time-Series Language Modeling through Task Compilation
WaveTLM is pitched as a way to make time-series model outputs obey task contracts, not just sound plausible.
The paper separates “task-object reliability” from predictive quality, targeting failures like wrong tensor shape, scale, channel order, timing, or invalid labels. It introduces ExecTS-QA, a benchmark covering forecasting, imputation, classification, anomaly detection, and waveform analysis. A single WaveTLM checkpoint reports 99.40% contract-valid coverage on ExecTS-QA, versus 37.83% for the strongest evaluated string-first baseline. The authors say code, construction scripts, and the dataset will be released upon publication. ArXiv · AI/CL/LG's note
The paper separates “task-object reliability” from predictive quality, targeting failures like wrong tensor shape, scale, channel order, timing, or invalid labels. It introduces ExecTS-QA, a benchmark covering forecasting, imputation, classification, anomaly detection, and waveform analysis. A single WaveTLM checkpoint reports 99.40% contract-valid coverage on ExecTS-QA, versus 37.83% for the strongest evaluated string-first baseline. The authors say code, construction scripts, and the dataset will be released upon publication. ArXiv · AI/CL/LG's note
score 5