Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays
The paper says a common accept-or-review method for document fields can miss its promised error control on real receipt data.
On 13,859 Claude Sonnet 5 fields from 800 CORD receipts, the author reports three failure modes: document clustering, score-refit leakage, and threshold collapse from tie-heavy scores. A split fit/validation protocol brought expected selective risk under a 10% target, but individual resplits still exceeded that target 47.5% of the time. Stronger Mondrian Learn-then-Test certificates worked by group, though the document-level tier was described as nearly vacuous in this setting. The paper says its stricter provenance conditioning helped on the Sonnet CORD capture but did not replicate under Haiku or Qwen. HF Daily Papers' note
On 13,859 Claude Sonnet 5 fields from 800 CORD receipts, the author reports three failure modes: document clustering, score-refit leakage, and threshold collapse from tie-heavy scores. A split fit/validation protocol brought expected selective risk under a 10% target, but individual resplits still exceeded that target 47.5% of the time. Stronger Mondrian Learn-then-Test certificates worked by group, though the document-level tier was described as nearly vacuous in this setting. The paper says its stricter provenance conditioning helped on the Sonnet CORD capture but did not replicate under Haiku or Qwen. HF Daily Papers' note
score 4