Megadose AI progress, ranked and analyzed.

Evaluating LLMs on Conversational Text-to-SQL under Chain Ambiguity and Intent Drift

· ArXiv · AI/CL/LG ·
TIDE-Bench tests whether models can track ambiguity and changed intent across multi-turn SQL requests.

The benchmark is built from 514 BIRD anchor SQLs and contains 1,542 samples. It targets two cases: chained clarifications with conditional dependencies, and user requests that retract and replace an earlier element. In tests of 12 advanced LLMs, the authors report a persistent bottleneck in chain identification, a gap between recognizing and resolving intent drift, and overlapping failures when both are present. The paper was accepted to EMNLP 2026, and the authors say the code is released for further research. ArXiv · AI/CL/LG's note

score 4

Categories: Research