Megadose Built for builders and researchers.

Understanding Errors in LLM-Based Question Answering over Imperfect Tables

· ArXiv · AI/CL/LG ·
Marking bad cells is not enough; the repaired table is what moves accuracy.

The paper tests LLM question answering on human-reviewed imperfect tables, varying row order and comparing original, error-marked, and repaired versions. Error discovery changed when rows were reordered, even though the table contents and gold answers stayed the same. Verified error locations still left a large gap versus tables with repairs already applied, where code-assisted QA rose by 39.0–59.1 points across three systems. The authors’ GBDI workflow combines shuffled views with explicit verification and handling guidance, improving observed QA accuracy by 3.8–18.5 points over a code-agent baseline on RADAR-T. ArXiv · AI/CL/LG's note

score 3

Categories: Research