FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
The benchmark tests agents on 2,026 deep-research finance-document tasks drawn from 1,009 documents.
The authors built a Finance-LaTeX skill to synthesize complex financial documents, producing 2,000 documents and 6,000 QA pairs. FinanceComplexQA is open-ended, bilingual, and covers multiple finance scenarios, task types, layouts, and expert-level reasoning questions. They use it to compare leading RAG systems and agentic reasoning tools, then analyze failures in computation, multi-hop reasoning, summarization, and industry analysis. HF Daily Papers' note
The authors built a Finance-LaTeX skill to synthesize complex financial documents, producing 2,000 documents and 6,000 QA pairs. FinanceComplexQA is open-ended, bilingual, and covers multiple finance scenarios, task types, layouts, and expert-level reasoning questions. They use it to compare leading RAG systems and agentic reasoning tools, then analyze failures in computation, multi-hop reasoning, summarization, and industry analysis. HF Daily Papers' note
score 4