V-FiLLM: Verified Financial LLM Reasoning Benchmark
V-FiLLM builds finance QA items from executable table-based computation trees, so answers are verified rather than model-labeled.
The benchmark controls difficulty across reasoning depth, expression breadth, financial concept complexity, and context size. In tests on open-source models, accuracy dropped as much as 51% with deeper reasoning and 47 points under adversarial numerical perturbations. The authors also report that LoRA tuning on verified chain-of-thought traces raised held-out accuracy from 81.1% to 85.6%. ArXiv · AI/CL/LG's note
The benchmark controls difficulty across reasoning depth, expression breadth, financial concept complexity, and context size. In tests on open-source models, accuracy dropped as much as 51% with deeper reasoning and 47 points under adversarial numerical perturbations. The authors also report that LoRA tuning on verified chain-of-thought traces raised held-out accuracy from 81.1% to 85.6%. ArXiv · AI/CL/LG's note
score 5