CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions
The dataset trains models to use visual workspaces for reasoning problems that start as text.
CoVA-SFT contains 51.9K samples and more than 222K multimodal reasoning steps across 17 tasks. The authors also introduce CoVA-Bench, a 1,700-sample held-out benchmark for evaluation. Fine-tuned models beat interleaved chain-of-thought baselines by more than 2x on average, but still trail strong text-only CoT baselines. Source: HF Daily Papers' note
CoVA-SFT contains 51.9K samples and more than 222K multimodal reasoning steps across 17 tasks. The authors also introduce CoVA-Bench, a 1,700-sample held-out benchmark for evaluation. Fine-tuned models beat interleaved chain-of-thought baselines by more than 2x on average, but still trail strong text-only CoT baselines. Source: HF Daily Papers' note
score 5