Megadose AI progress, ranked and analyzed.

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

· HF Daily Papers ·
LLMs falter on full financial statements when formulas are not handed to them.

The paper introduces FinIndices, a benchmark built around uncropped financial statements of up to 32K tokens. Its tests target single-index computation and table-index tabulation, with adversarial traps for temporal and accounting-structure errors. The authors report sharp drops when formula hints are removed, including Gemini-3.1-Pro falling from 70.70% to 38.22% on table tasks. They also find that supervised fine-tuning improves zero-hint performance, but only partially restores structured reasoning. HF Daily Papers' note

score 4

Categories: Research