LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering
LitTraceQA tests whether scientific QA systems can retrieve papers, ground evidence, and answer from that evidence as separate steps.
The benchmark asks systems to return canonical paper IDs, evidence locations, and answers in formats such as text, multiple choice, or tables. It covers evidence from tables, figures, text spans, equations or algorithms, and citation contexts. The public development split has 55 examples, while the larger annotation collection includes 4,978 unique-question records over 4,859 gold papers. ArXiv · AI/CL/LG's note
The benchmark asks systems to return canonical paper IDs, evidence locations, and answers in formats such as text, multiple choice, or tables. It covers evidence from tables, figures, text spans, equations or algorithms, and citation contexts. The public development split has 55 examples, while the larger annotation collection includes 4,978 unique-question records over 4,859 gold papers. ArXiv · AI/CL/LG's note
score 4