Megadose AI progress, ranked and analyzed.

V-FiLLM: Verified Financial LLM Reasoning Benchmark

· ArXiv · AI/CL/LG ·
V-FiLLM builds finance QA items from executable table-based computation trees, so answers are verified rather than model-labeled.

The benchmark controls difficulty across reasoning depth, expression breadth, financial concept complexity, and context size. In tests on open-source models, accuracy dropped as much as 51% with deeper reasoning and 47 points under adversarial numerical perturbations. The authors also report that LoRA tuning on verified chain-of-thought traces raised held-out accuracy from 81.1% to 85.6%. ArXiv · AI/CL/LG's note

score 5

Categories: Research